Task scheduling method and device, data processing unit, program product and medium
By using the data processing unit to schedule tasks based on the deviation and trend data of hardware resources, the problem of insufficient matching between task allocation and actual operation status in the server cluster is solved, and more efficient task allocation and scheduling is achieved.
Patent Information
- Application Number
- CN202511278209.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-09-09
AI Technical Summary
In the prior art, the task scheduling effect within the server cluster does not match the actual operation of the server well, and is difficult to match with the load change trend of the server.
The data processing unit determines the target server object for task allocation based on the deviation and deviation trend data between the server object's hardware resource usage and the target usage, ensuring that the hardware resource usage is close to the target usage and taking into account the actual operation of the server object.
It improves the matching effect between task allocation and the actual operation status of server objects, and improves the efficiency and accuracy of task scheduling.
Smart Images

Figure CN120762868A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of server technology, and in particular to a task scheduling method, device, data processing unit, program product, and medium. Background Art
[0002] With the emergence of computing-intensive tasks such as machine learning and simulation, how to schedule server clusters to efficiently execute computing tasks has become a technical problem that needs to be solved urgently.
[0003] In related technologies, task scheduling within a cluster is generally performed by server processors, and the scheduling effect does not match the actual operation status of the server well. Summary of the Invention
[0004] The present invention provides a task scheduling method, device, data processing unit, program product, and medium. The data processing unit can allocate target server objects to tasks to be processed based on the deviation between the server object's usage of hardware resources and the target usage of the hardware resources, as well as deviation trend data. This can ensure that the target server object's usage of hardware resources is close to the target usage, and can ensure that task allocation is closer to the actual operation status of the server object.
[0005] To solve the above technical problems, the present invention provides a task scheduling method applied to a data processing unit, the method comprising: Receive tasks to be processed and determine the hardware resource requirements of the tasks to be processed; For at least two server objects in a server cluster, determining first and second usages of hardware resources by the server objects at a first moment and a second moment, respectively, determining an expected usage using the second usage and the demand, and determining first and second deviations between a target usage of the hardware resource and the first and expected usage, respectively; wherein the first moment is a historically designated moment and the second moment is a task issuance moment; determining deviation trend data using the first deviation amount and the second deviation amount, and determining an allocation priority of the server object using the second deviation amount and the deviation trend data; The target server object for executing the pending task is determined based on the allocation priority of each server object, and the pending task is sent to the target server object.
[0006] The present invention also provides a task scheduling device, which is applied to a data processing unit and includes: The task receiving module is used to receive tasks to be processed and determine the hardware resource requirements of the tasks to be processed; a resource usage determination module configured to determine, for at least two server objects in a server cluster, first and second usages of hardware resources by the server objects at a first moment and a second moment, respectively, determine an expected usage using the second usage and the demand, and determine a first deviation and a second deviation between a target usage of the hardware resource and the first usage and the expected usage, respectively; wherein the first moment is a historically designated moment and the second moment is a task issuance moment; an allocation priority determination module, configured to determine deviation trend data using the first deviation amount and the second deviation amount, and determine an allocation priority of the server object using the second deviation amount and the deviation trend data; The task allocation module is used to determine the target server object for executing the pending task according to the allocation priority of each server object, and send the pending task to the target server object.
[0007] The present invention also provides a data processing unit, comprising: memory for storing computer programs; The processor is used to implement the above-mentioned task scheduling method when executing a computer program.
[0008] The present invention also provides a computer program product, including a computer program or instructions, which implements the above-mentioned task scheduling method when executed by a processor.
[0009] The present invention also provides a non-volatile computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are loaded and executed by a processor, the above-mentioned task scheduling method is implemented.
[0010] The application has the beneficial effects that: the application can firstly perform scheduling by the data processing unit. In addition, when receiving a to-be-processed task, the data processing unit can determine the demand of the to-be-processed task for hardware resources, then for at least two server objects in the server cluster, the data processing unit can determine the first usage and the second usage of the hardware resources of the server objects at the first time and the second time respectively, determine the expected usage by using the second usage and the demand, determine the first deviation and the second deviation between the target usage of the hardware resources and the first usage and the expected usage respectively, that is, the deviation between the first usage of the hardware resources of the server object and the target usage, and the deviation between the expected usage of the hardware resources of the server object and the target usage when the to-be-processed task is assumed to be assigned to the server object. Then, the data processing unit can determine the deviation trend data by using the first deviation and the second deviation, and determine the assignment priority of the server object by using the second deviation and the deviation trend data, and then determine the target server object for executing the to-be-processed task according to the assignment priority of each server object, and assign the to-be-processed task to the target server object. It can be seen that the data processing unit can perform task scheduling based on the target usage of the hardware resources, and can ensure that the usage of the target server object for the hardware resources is close to the target usage after the task is assigned. In addition, when performing task scheduling, the data processing unit can consider the first deviation and the second deviation between the usage of the hardware resources of the server object and the target usage, and can determine the assignment priority of the server object based on the second deviation and the deviation trend data, which can improve the matching effect between task assignment and the actual operation of the server object, thereby improving the task scheduling effect.
[0011] The application also provides a task scheduling device, a data processing unit, a program product and a medium, which have the beneficial effects. BRIEF DESCRIPTION OF DRAWINGS
[0012] In order to more clearly illustrate the embodiments of the application, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.
[0013] Figure 1 The structural block diagram of the first task scheduling system provided by the embodiments of the application; Figure 2 The structural block diagram of the second task scheduling system provided by the embodiments of the application; Figure 3 The flowchart of the task scheduling method provided by the embodiments of the application; Figure 4 A structural block diagram of a third task scheduling system provided by an embodiment of the present invention; Figure 5 A structural block diagram of a task scheduling device provided by an embodiment of the present invention; Figure 6 This is a structural block diagram of a data processing unit provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0014] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0015] It should be noted that, in the description of the present invention, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. The terms "first," "second," etc., in the present invention are used to distinguish similar objects, and are not used to describe a particular order or precedence.
[0016] In order to enable those skilled in the art to better understand the solutions of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0017] With the emergence of computationally intensive tasks such as machine learning and simulation, how to schedule server clusters to efficiently execute computing tasks has become a technical problem that urgently needs to be solved. In related technologies, task scheduling within a cluster is generally performed by a server processor, and the scheduling effect is poorly matched with the actual operation of the server. For example, it is difficult to match the load change trend of the server. In view of this, in order to address the technical problem of how to improve the scheduling effect of server cluster tasks, the present invention can provide a task scheduling method, which can allocate target server objects to tasks to be processed by a data processing unit based on the deviation between the usage of hardware resources by the server object and the target usage of the hardware resources, and the deviation trend data, thereby ensuring that the usage of hardware resources by the target server object is close to the target usage, and ensuring that task allocation is closer to the actual operation of the server object.
[0018] For ease of understanding, the system environment applicable to this embodiment is first introduced. Figure 1 , Figure 1This is a structural block diagram of the first task scheduling system provided by an embodiment of the present invention. The system may include at least two server nodes 2, each server node 2 being an independent server device, and together forming a server cluster. In addition, each server node 2 has an attached data processing unit 1 (DPU, Data Processing Unit), which has computing power and can offload part of the computing tasks from the processor of the server node 2 to the data processing unit 1. The server node 2 and the data processing unit 1 can communicate with each other through a bus structure (such as PCIe bus, Peripheral Component Interconnect express, a high-speed serial computer expansion bus standard, the bus structure is not in Figure 1 In addition, the data processing units 1 may also be connected to each other, such as by establishing a network connection through a network port or a network cable, thereby achieving network communication.
[0019] It should be noted that the task scheduling method provided in this embodiment can be executed by a single data processing unit 1. Furthermore, to improve scheduling efficiency and avoid single points of failure, the task scheduling method can also be executed jointly by at least two data processing units 1, such as all data processing units 1 in a server cluster. In this case, multiple data processing units 1 can form a data processing unit cluster and jointly perform task scheduling within the server cluster.
[0020] Furthermore, the present embodiment will later introduce the concept of server objects, which are virtual scheduling objects that can be used by the data processing unit 1 for task scheduling. The server object can correspond to either the server node 2 or the virtual server in the server node 2. For example, multiple processors and multiple network card devices can be set in the server node 2, and a single processor and a single network card device can meet the data processing and data transmission requirements. Therefore, a virtual server can be created based on a single processor and a single network card device. At this time, multiple virtual servers can be created in a single server node 2, which can correspond to multiple server objects. For ease of understanding, please refer to Figure 2 , Figure 2 This is a structural block diagram of the second task scheduling system provided by an embodiment of the present invention, which shows that each server node 2 can correspond to multiple server objects 21, and the data processing unit 1 will perform task scheduling based on the server objects 21.
[0021] Based on the above system structure introduction, the task scheduling method provided by this embodiment is introduced below. Figure 3 , Figure 3 This is a flowchart of a task scheduling method provided by an embodiment of the present invention. This method is applied to a data processing unit and may include: S101: Receive a task to be processed and determine the hardware resource requirements of the task to be processed.
[0022] In this embodiment, the data processing unit serves as the entry point for receiving pending tasks and is responsible for providing task scheduling services for the server cluster. The data processing unit first establishes a main scheduling function (scheduler) and a task queue. The main scheduling function receives externally issued pending tasks, determines the hardware resource requirements of the server object (such as processor resources, storage resources, network resources, bandwidth resources, etc.) required by the pending tasks, and sequentially adds the pending tasks to the task queue. Furthermore, the data processing unit also establishes a task scheduling thread (scheduler_thread) to perform task scheduling. After writing a pending task to the task queue, the main scheduling function uses a semaphore (task_semaphore) to notify the scheduling thread to remove the pending task from the task queue. Furthermore, the task queue can be configured with a mutex lock. The main scheduling function locks the mutex before writing the pending task to the task queue and then unlocks the mutex. Similarly, the scheduling thread locks the mutex before removing the pending task from the task queue and then unlocks the mutex.
[0023] It should be noted that this embodiment does not limit how to determine the demand for hardware resources of the task to be processed. For example, a demand forecast can be performed. According to the task type / task content of the task to be processed, the first demand for hardware resources of the task at a specified historical moment can be determined, and then the demand for hardware resources of the task to be processed can be predicted based on the first demand for hardware resources of the task to be processed.
[0024] In one embodiment, determining the hardware resource requirements of the task to be processed may include: Step 11: predicting a second demand for hardware resources by the task to be processed based on the first demand for hardware resources by the task to be processed at a specified historical moment.
[0025] Furthermore, given that predicting hardware resource requirements consumes a significant amount of the data processing unit's hardware resources, when the data processing unit's remaining hardware resource usage is insufficient (e.g., when only 10% of the hardware resources remain), requirements can be randomly generated for pending tasks. Specifically, when the data processing unit's own hardware resource usage reaches a preset threshold, requirements can be randomly generated for pending tasks; otherwise, requirements can be predicted for pending tasks.
[0026] In another embodiment, the method may further include: Step 21: Determine whether the hardware resource usage reaches a preset threshold; if it reaches the preset threshold, proceed to step 22; if it does not reach the preset threshold, proceed to step 23.
[0027] Step 22: Randomly generate demand quantities for pending tasks.
[0028] Step 23: Entering the step of predicting the demand for hardware resources of the task to be processed based on the first demand for hardware resources of the task to be processed at the historical specified time.
[0029] S102. For at least two server objects in the server cluster, determine the first usage and the second usage of the hardware resources by the server objects at the first moment and the second moment respectively, determine the expected usage using the second usage and the demand, and determine the first deviation and the second deviation between the target usage of the hardware resources and the first usage and the expected usage respectively; wherein the first moment is the historical designated moment, and the second moment is the moment when the task is issued.
[0030] In this embodiment, the task issuance moment is the moment when the task to be processed is issued, and the historical designated moment is the designated moment before the task is issued. The data processing unit can set corresponding target usage for the hardware resources used by the server object. For example, a first target usage is set for the processor, a second target usage is set for the memory, and a third target usage is set for the network bandwidth. The target usage may be lower than the total amount of hardware resources, for example, the target usage is 80% of the total amount of hardware resources. The purpose of setting the target usage is to control the server object's usage of hardware resources to be maintained at around the target usage. On the one hand, it can ensure that the hardware resources are fully utilized, and on the other hand, it can avoid the overall load pressure of the server cluster being too high.
[0031] Furthermore, in order to ensure that task scheduling is closer to the actual operation of the server object, the scheduling thread in the data processing unit needs to analyze the changing trend of the hardware resource usage of each server object at the first moment and the second moment, and needs to focus on analyzing the deviation changing trend between the hardware resource usage of each server object and the target usage. Among them, the first moment can be a historically specified moment, and the second moment can be the moment when the task is issued (i.e., the current moment). To this end, for each server object in the server cluster, the data processing unit can obtain its first usage and second usage of hardware resources at the first moment and the second moment. Subsequently, the data processing unit can determine to issue the task to the server object at the second moment, and the expected usage of the hardware resources generated by the server object can be obtained by summing the second usage with the previously determined demand for hardware resources of the task to be processed. Finally, after obtaining the first usage and the expected usage, the data processing unit can determine the first deviation and the second deviation between the target usage of the hardware resource and the first usage and the expected usage, respectively. The deviation is calculated as follows: Deviation = target usage - actual usage; Among them, the actual usage is the first usage or the second usage. It can be seen that if the actual usage is less than the target usage, the deviation is a positive value. If the actual usage is greater than the target usage, the deviation is a negative value. It can be understood that when the deviation is a positive value, it means that the server object has remaining hardware resources and can be assigned tasks, and the larger the deviation is, the more suitable it is for assigning tasks. When the deviation is a negative value, it means that the server object's usage of hardware resources has exceeded the target usage, and assigning tasks to it may result in insufficient remaining hardware resources, so it is not suitable for assigning tasks, and the smaller the deviation is, the less suitable it is for assigning tasks. Furthermore, after obtaining the first deviation and the second deviation, the deviation trend data can be determined in the next step.
[0032] It should be noted that this embodiment does not limit the number of first moments in which the data processing unit needs to determine the first usage value of the server object. For example, it can be for all first moments from the server object startup moment to the second moment, or for the first moment in a preset time period before the second moment.
[0033] Furthermore, a server object can correspond to either a server node or a virtual server within a server node. To effectively improve resource utilization within a server node and enhance task scheduling effectiveness, this embodiment can create corresponding server objects based on each network card device and each processor within the server object, thereby mapping the hardware resource usage values of each network card device and each processor to the corresponding server object. Therefore, before performing task scheduling, the main scheduling function of the data processing unit can determine the number of server objects based on the number of server nodes, the number of network cards, and the number of processors, and create server objects based on the number of server objects so that the scheduling thread can perform task scheduling.
[0034] In one embodiment, the method may further include: Step 31: Obtain the number of server nodes in the server cluster, and obtain the number of network cards and the number of processors in the server nodes.
[0035] Step 32: Determine the number of server objects using the number of server nodes, the number of network cards, and the number of processors, and create server objects based on the number of server objects, so as to form server objects using each network card device and each processor.
[0036] S103: Determine deviation trend data using the first deviation amount and the second deviation amount, and determine an allocation priority of the server object using the second deviation amount and the deviation trend data.
[0037] In the embodiment, the data processing unit in the data processing unit can determine the deviation trend data by using the first deviation amount and the second deviation amount, so as to determine the allocation priority of the server object by using the second deviation amount and the deviation trend data. Of course, if it is determined in the previous step that the expected use amount of the server object exceeds the total amount of the hardware resources, it is obviously impossible to allocate tasks to the server object. Therefore, before step S103 is performed, the server object whose expected use amount does not exceed the total amount of the hardware resources can also be taken as a candidate server object, and step S103 can be performed only for the candidate server object.
[0038] In an implementation, before the deviation trend data is determined by using the first deviation amount and the second deviation amount, the following can also be included: Step 41: In the server objects, the server object whose expected use amount does not exceed the total amount of the hardware resources is taken as a candidate server object.
[0039] Step 42: For the candidate server object, the step of determining the deviation trend data by using the first deviation amount and the second deviation amount, and determining the allocation priority of the server object by using the second deviation amount and the deviation trend data is entered.
[0040] Of course, when the candidate server object is screened, the concurrency (i.e., the number of simultaneously executed tasks) of the candidate server object can also be considered, and only the server object whose concurrency is less than a preset number and whose expected use amount does not exceed the total amount of the hardware resources is taken as the candidate server object.
[0041] Further, the deviation trend data first needs to indicate the first deviation accumulation of the server object in the first time, such as a large deviation from the target use value, a small deviation, etc. For this, the embodiment can accumulate the deviation amounts of multiple times to obtain the accumulated deviation amount. It should be noted that the calculation of the accumulated deviation amount at least needs to use the second deviation amount of the second time and the first deviation amount of the previous first time of the second time, and of course can further use the first deviation amount of each first time in a preset time period before the second time, and of course can further use the first deviation amount of each first time from the starting time of the server object to the second time. In order to avoid excessive accumulation and reduce the reliability of the accumulated deviation amount, the accumulated deviation amount can be calculated by using the second deviation amount of the second time and the first deviation amount of the first time in a preset time period before the second time. It should be noted that the embodiment does not limit the length of the preset time period, which can be set according to actual application requirements.
[0042] Furthermore, the deviation trend data must also indicate fluctuations in the deviation at the second moment, such as whether the rate of change was large or small. To this end, this embodiment can also use the first and second deviations to determine the deviation change rate of the server object at at least two moments in time. Furthermore, this embodiment can use the cumulative deviation and the deviation change rate as deviation trend data.
[0043] In one embodiment, determining deviation trend data using the first deviation amount and the second deviation amount, and determining the allocation priority of the server object using the second deviation amount and the deviation trend data may include: Step 51: Using the first deviation and the second deviation, determine the cumulative deviation and the deviation change rate corresponding to the server object at at least two moments.
[0044] Step 52: Determine the allocation priority of each server object using the second deviation, the cumulative deviation, and the deviation change rate.
[0045] In a specific embodiment, determining the cumulative deviation and the deviation change rate corresponding to the server object at at least two moments using the first deviation and the second deviation may include: Step 61: Determine a cumulative deviation using the second deviation, the first deviation at a first moment in a preset time period before the second moment, and the time interval.
[0046] Step 62: Determine the deviation change rate using the second deviation, the first deviation at the first moment before the second moment, and the time interval.
[0047] Specifically, the cumulative deviation can be expressed as: ; Among them, T s It represents the time interval (such as 10ms), i represents the i-th moment, j-th moment is the second moment, and e(i) represents the deviation of the i-th moment.
[0048] The rate of change of deviation can be expressed as: ; Here, e(j) represents the second deviation, and e(j-1) represents the first deviation at the first moment before the second moment.
[0049] Furthermore, after obtaining the second deviation, cumulative deviation, and deviation change rate, the second deviation, cumulative deviation, and deviation change rate can be used to determine the allocation priority of each server object. The second deviation amount indicates whether the server object has allocable hardware resources at the second time. If the second deviation amount is positive, it indicates that the server object has remaining hardware resources, and the task allocation can be performed, and the greater the deviation amount, the more suitable for allocating the task. If the second deviation amount is negative, it indicates that the usage amount of hardware resources of the server object has exceeded the target usage amount, and further allocating the task can cause the server object to have insufficient remaining hardware resources, and thus is not suitable for allocating the task, and the smaller the deviation amount, the less suitable for allocating the task.
[0050] The cumulative deviation amount indicates whether the server object sufficiently uses the hardware resources at the first time. If the cumulative deviation amount is positive, it indicates that the usage amount of hardware resources of the server object at the first time does not reach the target usage amount, and the usage of hardware resources is insufficient, and the greater the cumulative deviation amount, the more insufficient the usage of hardware resources, and the more suitable for allocating the task. If the cumulative deviation amount is negative, it indicates that the usage amount of hardware resources of the server object at the first time exceeds the target usage amount, and the hardware resources are excessively used, and the smaller the cumulative deviation amount, the more excessive the usage of hardware resources, and the less suitable for allocating the task.
[0051] The deviation change rate indicates the load change of the server object at the second time. When the deviation change rate is positive, it indicates that the server load is decreasing, and the greater the deviation change rate, the faster the load decreases, and the more suitable for allocating the task. When the deviation change rate is negative, it indicates that the server load is increasing, and the smaller the deviation change rate, the faster the load increases, and the less suitable for allocating the task.
[0052] Therefore, based on the second deviation amount, the cumulative deviation amount, and the deviation change rate, the embodiment can comprehensively consider the remaining situation of hardware resources of the server object at the second time, the usage situation of hardware resources at the first time, and the load change situation at the second time, so as to more reliably determine the allocation priority.
[0053] It should be noted that the embodiment does not limit how to determine the allocation priority of each server object by using the second deviation amount, the cumulative deviation amount, and the deviation change rate. For example, the second deviation amount, the cumulative deviation amount, and the deviation change rate can be simply summed to obtain the allocation priority, or the second deviation amount, the cumulative deviation amount, and the deviation change rate can be weighted and summed to obtain the allocation priority. Considering that the remaining situation of hardware resources of the server object at the second time, the usage situation of hardware resources at the first time, and the load change situation at the second time have different influences on the allocation priority, the embodiment can adopt the weighted sum manner to determine the allocation priority.
[0054] In an implementation manner, determining the allocation priority of each server object by using the second deviation amount, the cumulative deviation amount, and the deviation change rate can include: Step 71: Using the first parameter, the second parameter, and the third parameter, respectively, weighted sum is performed on the second deviation, the cumulative deviation, and the deviation change rate to obtain the allocation priority of the server object.
[0055] Specifically, the allocation priority can be expressed as: ; Among them, k p 、k i 、k d They are the first parameter, the second parameter and the third parameter respectively.
[0056] Furthermore, to increase the likelihood of the server object being assigned when the second deviation of the server object is large, and to decrease the likelihood of the server object being assigned when the second deviation of the server object is small, this embodiment may also adjust the first parameter, the second parameter, and the third parameter based on the second deviation. Specifically, a preset parameter value adjustment relationship may be set, and this relationship information may record a mapping relationship between a preset second deviation interval and a preset parameter adjustment amount. Furthermore, the second deviation may be matched with the second deviation interval to determine the parameter adjustment amount corresponding to the second deviation, and the parameter adjustment amount may be used to adjust the first parameter, the second parameter, and the third parameter.
[0057] In one embodiment, the method may further include: Step 81: matching a parameter adjustment amount corresponding to the second deviation amount in a preset parameter value adjustment relationship, and adjusting the first parameter, the second parameter, and the third parameter using the parameter adjustment amount.
[0058] It should be noted that this embodiment does not limit the specific preset parameter value adjustment relationship, which can be set according to actual application requirements. After the parameter adjustment, the allocation priority can be expressed as: ; Among them, k p (j), k i (j), k d (j) are the first parameter, second parameter and third parameter at the jth moment, respectively, and are calculated by the following formulas: ; ; ; in, 、 、 Indicates the parameter adjustment amount at the jth moment.
[0059] Furthermore, considering that the server object can use multiple hardware resources and each hardware resource has a different impact on task execution, the scheduling thread in the data processing unit can calculate the sub-priority corresponding to each hardware resource for the server object, and perform weighted summation of each sub-priority according to the preset weight corresponding to each hardware resource to obtain the allocation priority of the server object, where the calculation method of the sub-priority is the same as the calculation method of the above-mentioned allocation priority.
[0060] In one embodiment, determining the allocation priority of the server object using the second deviation amount and the deviation trend data may include: Step 91: For at least two hardware resources, determine the sub-priority level corresponding to the hardware resource using the second deviation amount and deviation trend data corresponding to the hardware resource.
[0061] Step 92: performing weighted summation on each sub-priority based on the preset weight corresponding to each hardware resource to obtain the allocation priority of the server object.
[0062] S104: Determine the target server object for executing the pending task according to the allocation priority of each server object, and send the pending task to the target server object.
[0063] Specifically, a higher allocation priority indicates that the server object likely has more remaining hardware resources, is less likely to have fully utilized its hardware resources at the initial moment, or has experienced a greater decrease in load, making it more suitable for task allocation. Therefore, the server object with the highest allocation priority can be selected as the target server object.
[0064] In one embodiment, determining a target server object for executing a pending task according to the allocation priority of each server object includes: Step 1101: The server object with the highest allocation priority is used as the target server object.
[0065] Furthermore, after determining the target server object, the scheduling thread in the data processing unit can dispatch pending tasks to the corresponding target server object for processing. Subsequently, the data processing unit can update the target server object's resource usage, concurrency count, and other information. It can also record and predict the target server object's task response time and task completion time, and wait for the next round of scheduling.
[0066] Furthermore, a task completion thread (task_completion_thread) may be set in the data processing unit. The task completion thread may release the resources currently occupied by the scheduling thread after the scheduling thread completes a round of scheduling, so as to avoid excessive use of data processing unit resources.
[0067] Furthermore, since the above steps S102-S103 consume a large amount of hardware resources of the data processing unit for calculation, when the overall load of the server cluster is not high, it is easy to cause a waste of hardware resources. Therefore, in order to improve the targeted scheduling, the scheduling thread of the data processing unit executes steps S102-S104 only when it determines that the usage of any hardware resource by any server object has reached a preset threshold. If it is determined that the usage of any hardware resource by any server object has not reached the preset threshold, the scheduling thread of the data processing unit can send the pending task to a server object with idle hardware resources.
[0068] In one embodiment, before determining the hardware resource requirements of the task to be processed, the following steps may also be performed: Step 1201: When it is determined that the usage of any hardware resource by any server object reaches a preset threshold, the process proceeds to the step of determining the demand for hardware resources by the task to be processed.
[0069] Step 1202: When it is determined that the usage of any hardware resource by any server object does not reach a preset threshold, the to-be-processed task is sent to a server object with idle hardware resources.
[0070] It should be noted that this embodiment does not limit the specific value of the preset threshold, which can be set according to actual application requirements. For example, the preset threshold can be the target usage or can be lower than the target usage.
[0071] Based on the above embodiment, the present invention can first be scheduled by a data processing unit. In addition, when receiving a pending task, the data processing unit can determine the demand for hardware resources of the pending task, and then, for at least two server objects in the server cluster, can determine the first usage and second usage of the hardware resources of the server object at a first moment and a second moment, respectively, wherein the first moment is a historically designated moment and the second moment is the moment when the task is issued, and use the second usage and demand to determine the expected usage, and determine the first deviation and second deviation between the target usage of the hardware resource and the first usage and the expected usage, respectively, i.e., the deviation between the first usage of the hardware resource by the server object and the target usage, and assuming that the second pending task is issued to the server object, the deviation between the expected usage of the hardware resource by the server object and the target usage. Subsequently, the data processing unit can use the first deviation and the second deviation to determine deviation trend data, and use the second deviation and the deviation trend data to determine the allocation priority of the server object, and then determine the target server object for executing the pending task according to the allocation priority of each server object, and issue the pending task to the target server object. It can be seen that the data processing unit can schedule tasks based on the target usage of hardware resources, and can ensure that after the task is issued, the target server object's use of hardware resources is close to the target usage; in addition, when scheduling tasks, the data processing unit can consider the first deviation and second deviation between the server object's use of hardware resources and the target usage, and can determine the allocation priority of the server object based on the second deviation and the deviation trend data. It can determine the priority while comprehensively considering the deviation and the deviation trend, and can improve the matching effect between task allocation and the actual operation of the server object, thereby improving the task scheduling effect.
[0072] Based on the above embodiment, in this embodiment, at least two data processing units can form a data processing unit cluster to jointly dispatch tasks to be processed to the target server object. The working process of the data processing unit in the data processing unit cluster is introduced below.
[0073] In another embodiment, the method may further include: S201: Receive a task to be processed, and determine the hardware resource requirements of the task to be processed.
[0074] S202. For at least two server objects in the server node to which it belongs, determine the first usage and the second usage of the hardware resources by the server objects at the first moment and the second moment respectively, determine the expected usage using the second usage and the demand, and determine the first deviation and the second deviation between the target usage of the hardware resources and the first usage and the expected usage respectively; wherein the first moment is the historical designated moment, and the second moment is the moment when the task is issued.
[0075] S203: Determine deviation trend data using the first deviation amount and the second deviation amount, and determine the initial allocation priority of the server objects in the server node to which the server object belongs using the second deviation amount and the deviation trend data.
[0076] Different from the above embodiment, in steps S202 to S203, the data processing unit only calculates the allocation priority of the server objects in the server node to which it belongs, and obtains the initial allocation priority of the server objects in the server node to which it belongs.
[0077] S204: Establish a correspondence between the server object and the initial allocation priority, and write the correspondence into the shared memory.
[0078] S205 , reading the corresponding relationships written by itself and other data processing units from the shared memory, and sorting the server objects in each corresponding relationship according to the initial allocation priority in each corresponding relationship.
[0079] S206: Determine the target server object for executing the task to be processed according to the sorting result.
[0080] In steps S204 to S206, each data processing unit will synchronize the initial allocation priority of each server object in each server node to each data processing unit based on the data synchronization mechanism, so that each data processing unit can determine the target server object that is most suitable for the assigned task. Specifically, the data processing unit can establish a correspondence between the server object and the initial allocation priority, and write the correspondence into a shared memory, which is shared by all data processing units. Subsequently, after all data processing units have completed writing, each data processing unit can read the correspondence written by itself and other data processing units from the shared memory, and sort the server objects in each correspondence according to the initial allocation priority in each correspondence, thereby completing the determination of the target allocation priority.
[0081] It should be noted that the above-mentioned shared memory can be an independent shared memory device and can be accessed by each data processing unit. The above-mentioned shared memory can also be a logical shared memory, which is jointly formed by the memory space of all data processing units. Specifically, a shared memory area is set in the memory area of each data processing unit, and these shared memory areas can form the above-mentioned shared memory. At the same time, a network connection can be established between the data processing units. At this time, the writing, synchronization, and reading processes of the shared memory can be: 1. The data processing unit writes the pre-established correspondence to the shared memory area of its own memory space; 2. The data processing unit exchanges data in the shared memory area with other data processing units through the network, writes the correspondence in its own shared memory area to the shared memory area of other data processing units, and writes the correspondence in the shared memory area of other data processing units to its own shared memory area; 3. After completing the exchange, the data processing unit reads the full amount of correspondence from the shared memory area of its own memory space to determine the target allocation priority.
[0082] Based on this, a shared memory area is set in the memory area of each data processing unit; writing the corresponding relationship into the shared memory may include: Step 1301: Write the corresponding relationship into its own shared memory area.
[0083] Step 1302: Exchange data in the shared memory area with other data processing units through the network to write the corresponding relationship in its own shared memory area into the shared memory area of other data processing units, and write the corresponding relationship in the shared memory area of other data processing units into its own shared memory area.
[0084] Please refer to Figure 4 , Figure 4 This is a block diagram of the structure of the third task scheduling system provided by an embodiment of the present invention. It can be seen that each data processing unit 1 may include a memory 113, and the memory 113 has a shared memory area 1131. The data processing units 1 can synchronize the allocation of priorities by exchanging data in the shared memory area 1131.
[0085] It should be noted that this embodiment does not limit the method of data exchange. For example, the data processing unit can submit the data in the shared memory area to a master data processing unit, and then the master data processing unit distributes the data to all data processing units. For another example, the data processing units can exchange data in the shared memory area in pairs, such as pre-building a ring topology relationship between the data processing units, and then the data processing unit can send its own data to an adjacent data processing unit and receive data sent by another adjacent data processing unit, thereby completing synchronization. In order to achieve decentralization, the data processing units can adopt a method of exchanging data in the shared memory area in pairs to achieve shared memory synchronization.
[0086] S207, determine whether the target server object is located in the server node to which it belongs; if so, proceed to step S208; if not, proceed to step S209.
[0087] S208: Send the pending task to the target server object.
[0088] S209: Ignore pending tasks.
[0089] In steps S207 to S209, the data processing unit can determine whether the target server object is located in the server node to which it belongs. If so, it can be sent; otherwise, the pending task can be ignored, thereby achieving the nearest sending of tasks.
[0090] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0091] Please refer to Figure 5 , Figure 5 This is a structural block diagram of a task scheduling device provided by an embodiment of the present invention. The device is applied to a data processing unit and may include: The task receiving module 501 is used to receive tasks to be processed and determine the hardware resource requirements of the tasks to be processed; The resource usage determination module 502 is configured to determine, for at least two server objects in the server cluster, first and second usages of hardware resources by the server objects at a first moment and a second moment, respectively, determine an expected usage using the second usage and the demand, and determine a first deviation and a second deviation between a target usage of the hardware resource and the first usage and the expected usage, respectively, where the first moment is a historically designated moment and the second moment is a task issuance moment; an allocation priority determination module 503 for determining deviation trend data using the first deviation amount and the second deviation amount, and determining an allocation priority of the server object using the second deviation amount and the deviation trend data; The task allocation module 504 is configured to determine a target server object for executing the task to be processed according to the allocation priority of each server object, and to send the task to be processed to the target server object.
[0092] Optionally, the allocation priority determination module 503 may include: a deviation trend data determination submodule, configured to determine a cumulative deviation and a deviation change rate corresponding to the server object at at least two moments using the first deviation and the second deviation; The allocation priority determination submodule is used to determine the allocation priority of each server object by using the second deviation, the cumulative deviation and the deviation change rate.
[0093] Optionally, the deviation trend data determination submodule includes: a cumulative deviation calculation unit, configured to determine a cumulative deviation using the second deviation, the first deviation at a first moment in a preset time period before the second moment, and the time interval; The deviation change rate calculation unit is used to determine the deviation change rate by using the second deviation, the first deviation at a first moment before the second moment, and the time interval.
[0094] Optionally, the allocation priority determination submodule may include: The allocation priority determination unit is used to use the first parameter, the second parameter, and the third parameter to perform weighted summation on the second deviation, the cumulative deviation, and the deviation change rate to obtain the allocation priority of the server object.
[0095] Optionally, the allocation priority determination submodule may further include: The parameter adjustment unit is used to match the parameter adjustment amount corresponding to the second deviation amount in the preset parameter value adjustment relationship, and adjust the first parameter, the second parameter, and the third parameter using the parameter adjustment amount.
[0096] Optionally, the allocation priority determination module 503 includes: a sub-priority determination sub-module, configured to determine, for at least two hardware resources, the sub-priority corresponding to the hardware resource by using the second deviation amount and deviation trend data corresponding to the hardware resource; The sub-priority weighting sub-module is used to perform weighted summation on each sub-priority according to the preset weight corresponding to each hardware resource to obtain the allocation priority of the server object.
[0097] Optionally, the task allocation module 504 may be configured to: The server object with the maximum allocation priority is taken as a target server object.
[0098] Optionally, the apparatus can further include: a candidate server object screening module configured to screen, from the server objects, a server object whose expected usage amount does not exceed the total amount of hardware resources as a candidate server object; The allocation priority determination module 503 can be configured to, for the candidate server object, enter a step of determining deviation trend data by using the first deviation amount and the second deviation amount, and determining the allocation priority of the server object by using the second deviation amount and the deviation trend data.
[0099] Optionally, the task receiving module 501 can be configured to, when determining that the usage amount of any server object to any hardware resource reaches a preset threshold, enter a step of determining the demand amount of the to-be-processed task to the hardware resource; The apparatus can further include: a common scheduling module configured to, when determining that the usage amount of any server object to any hardware resource does not reach the preset threshold, issue the to-be-processed task to the server object with idle hardware resource.
[0100] Optionally, the apparatus can further include: a quantity obtaining module configured to obtain the number of server nodes of the server cluster, and obtain the number of network cards and the number of processors in the server nodes; a server object creating module configured to determine the number of server objects by using the number of server nodes, the number of network cards, and the number of processors, and create the server objects according to the number of server objects, so as to form the server objects by using each network card device and each processor.
[0101] Optionally, the task receiving module 501 can include: a demand amount prediction submodule configured to predict the demand amount of the to-be-processed task to the hardware resource according to the first demand amount of the to-be-processed task to the hardware resource at a historical specified time.
[0102] Optionally, the task receiving module 501 can include: a judgment module configured to judge whether the usage amount of the hardware resource of itself reaches a preset threshold; a demand amount random generation submodule configured to, if yes, randomly generate the demand amount for the to-be-processed task; a demand amount prediction submodule configured to, if no, enter a step of predicting the demand amount of the to-be-processed task to the hardware resource according to the first demand amount of the to-be-processed task to the hardware resource at a historical specified time.
[0103] Optionally, at least two data processing units form a data processing unit cluster to jointly schedule the task to be processed to the target server object.
[0104] Optionally, the resource usage determination module 502 may be configured to: determining, for at least two server objects in the server node to which the server object belongs, a first usage and a second usage of the hardware resource by the server objects at a first moment and a second moment, respectively, determining an expected usage using the second usage and the demand, and determining a first deviation and a second deviation between a target usage of the hardware resource and the first usage and the expected usage, respectively; The allocation priority determination module 503 may include: an initial allocation priority determination submodule, configured to determine deviation trend data using the first deviation amount and the second deviation amount, and to determine an initial allocation priority of a server object in the server node to which the server object belongs using the second deviation amount and the deviation trend data; A shared memory writing submodule is used to establish a corresponding relationship between the server object and the initial allocation priority, and write the corresponding relationship into the shared memory; The task allocation module 504 may include: The shared memory reading submodule is used to read the corresponding relationships written by itself and other data processing units from the shared memory, and sort the server objects in each corresponding relationship according to the initial allocation priority in each corresponding relationship; The target server object determination submodule is used to determine the target server object for executing the task to be processed based on the sorting result.
[0105] Optionally, the task assignment module 504 may include: The task sending submodule is used to determine whether the target server object is located in the server node to which it belongs; if so, the pending task is sent to the target server object; if not, the pending task is ignored.
[0106] Optionally, a shared memory area is provided in the memory area of each data processing unit; Shared memory writing submodule, which can include: The local memory writing unit is used to write the corresponding relationship into its own shared memory area; The exchange unit is used to exchange data in the shared memory area with other data processing units through the network, so as to write the corresponding relationship in its own shared memory area into the shared memory area of other data processing units, and write the corresponding relationship in the shared memory area of other data processing units into its own shared memory area.
[0107] For the description of the features in the embodiment corresponding to the task allocation device, reference can be made to the relevant description of the embodiment corresponding to the task allocation method, which will not be repeated here.
[0108] This embodiment further provides a data processing unit, which may include: memory for storing computer programs; The processor is used to implement the above-mentioned task scheduling method when executing a computer program.
[0109] Please refer to Figure 6 , Figure 6 This is a structural block diagram of a data processing unit provided in an embodiment of the present invention. The data processing unit 1 may include: Processor module 11; Programmable logic device module 12 is connected to network port 14 and bus interface 13. Network port 14 can be used to connect to external devices and other data processing units 1, and bus interface 13 can be used to connect to server nodes. A network transmission module (not shown) can be provided within programmable logic device module 12. This network transmission module can support TCP and RDMA (RoCEv2) (remote direct memory access) protocols to implement data transmission.
[0110] The processor module 11 and the programmable logic device module 12 can be used together to execute the above-mentioned task scheduling method. For example, the programmable logic device module 12 can receive pending tasks sent by an external device via the network port 14 and transmit the pending tasks to the processor module 11. The processor module 11 can perform the following steps: determining the hardware resource demand of the pending tasks; determining, for at least two server objects in the server cluster, first and second hardware resource usages of the server objects at a first moment and a second moment, respectively; determining an expected usage using the second usage and the demand; and determining first and second deviations between the target hardware resource usage and the first and expected usages, respectively; wherein the first moment is a historically designated moment and the second moment is a task dispatch moment; determining deviation trend data using the first and second deviations; and determining the allocation priority of the server objects using the second deviation and the deviation trend data; and determining the target server object for executing the pending task based on the allocation priority of each server object. Subsequently, the processor module 11 can transmit the pending task and target server object information to the programmable logic device 12, which then transmits the pending task to the target server object in the server node via the bus interface 13.
[0111] Specifically, the processor module 11 may include a processor 111, an accelerator 112, and a memory 113. The processor 111 and the accelerator 112 may jointly execute the steps required by the processor module 11. The accelerator 112 may also perform some computational tasks on behalf of the processor 111, such as calculating expected usage, a first deviation, a second deviation, and deviation trend data. The memory 113 may include a shared memory area 1131. The programmable logic device 12 may include a programmable logic device 121 (e.g., an FPGA, Field Programmable Gate Array). The programmable logic device 121 may be connected to the processor 111 via a bus structure (e.g., a PCIe bus). The programmable logic device 121 may also be connected to the network port 14 and the bus interface 13.
[0112] It is worth noting that when multiple data processing units 1 form a data processing unit cluster and jointly perform task scheduling, the interaction process between the processor module 11 and the programmable logic device module 12 can be: 1. The programmable logic device module 12 receives a task to be processed sent by an external device through the network and transmits it to the memory 113 of the processor module 11.
[0113] 2. The processor module 11 takes out the pending tasks from the memory 113 and determines the initial allocation priority of the server object in the corresponding server node, establishes a correspondence between the server object and the initial allocation priority, and writes the correspondence into the shared memory area 1131 in the memory 113.
[0114] 3. The processor module 11 controls the programmable logic device module 12 to send the corresponding relationship in the shared memory area 1131 to other data processing units 1 through the network port 14, and controls the programmable logic device module 12 to write the corresponding relationship sent by other data processing units 1 through the network port 14 into the shared memory area 1131.
[0115] 4. The processor module 11 reads the corresponding relationships written by itself and other data processing units 1 from the shared memory area 1131 , and sorts the server objects in each corresponding relationship according to the initial allocation priority in each corresponding relationship.
[0116] 5. The processor module 11 determines the target server object for executing the task to be processed based on the sorting result. If the target server object is the server object to which it belongs, the task to be processed is sent to the programmable logic device module 12, and the programmable logic device module 12 sends the task to be processed to the server node through the bus interface 13.
[0117] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned task scheduling method embodiments when running.
[0118] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0119] An embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any one of the above-mentioned task scheduling method embodiments are implemented.
[0120] An embodiment of the present invention also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of any of the above-mentioned task scheduling method embodiments.
[0121] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0122] The above is a detailed introduction to a task scheduling method, device, data processing unit, program product, and medium provided by the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, the present invention can also be improved and modified in several ways, and these improvements and modifications also fall within the scope of protection of the present invention.
Claims
1. A task scheduling method, characterized in that: Applied to a data processing unit, the method includes: Receive a task to be processed and determine the hardware resource requirements of the task to be processed; For at least two server objects in a server cluster, determining first and second usages of hardware resources by the server objects at a first moment and a second moment, respectively, determining an expected usage using the second usage and the demand, and determining first and second deviations between a target usage of the hardware resource and the first and expected usage, respectively; wherein the first moment is a historically designated moment and the second moment is a task issuance moment; determining deviation trend data using the first deviation amount and the second deviation amount, and determining an allocation priority of the server object using the second deviation amount and the deviation trend data; The target server object for executing the task to be processed is determined according to the allocation priority of each server object, and the task to be processed is sent to the target server object.
2. The task scheduling method according to claim 1, characterized in that: Determining deviation trend data using the first deviation amount and the second deviation amount, and determining an allocation priority of the server object using the second deviation amount and the deviation trend data, includes: Determine the cumulative deviation and the deviation change rate corresponding to the server object at at least two moments using the first deviation and the second deviation; The allocation priority of each server object is determined using the second deviation, the cumulative deviation, and the deviation change rate.
3. The task scheduling method according to claim 2, characterized in that: Determining the cumulative deviation and the deviation change rate corresponding to the server object at at least two moments using the first deviation and the second deviation includes: determining the accumulated deviation using the second deviation, a first deviation at a first moment in a preset time period before the second moment, and a time interval; The deviation change rate is determined using the second deviation, the first deviation at a first moment before the second moment, and the time interval.
4. The task scheduling method according to claim 2, wherein: Determining the allocation priority of each server object by using the second deviation, the cumulative deviation, and the deviation change rate includes: The first parameter, the second parameter, and the third parameter are used to perform weighted summation on the second deviation, the cumulative deviation, and the deviation change rate to obtain the allocation priority of the server object.
5. The task scheduling method according to claim 4, characterized in that: Also includes: A parameter adjustment amount corresponding to the second deviation amount is matched in a preset parameter value adjustment relationship, and the first parameter, the second parameter, and the third parameter are adjusted using the parameter adjustment amount.
6. The task scheduling method according to claim 1, wherein: Determining the allocation priority of the server object by using the second deviation amount and the deviation trend data includes: For at least two hardware resources, determining the sub-priorities corresponding to the hardware resources by using the second deviation amounts and deviation trend data corresponding to the hardware resources; According to the preset weights corresponding to the hardware resources, the sub-priorities are weighted and summed to obtain the allocation priority of the server object.
7. The task scheduling method according to claim 6, characterized in that: Determining a target server object for executing the task to be processed according to the allocation priority of each server object includes: The server object having the highest allocation priority is used as the target server object.
8. The task scheduling method according to claim 1, wherein: Before determining deviation trend data using the first deviation amount and the second deviation amount, the method further includes: Among the server objects, the server objects whose expected usage does not exceed the total amount of hardware resources are selected as candidate server objects; For the candidate server object, the step of determining deviation trend data using the first deviation amount and the second deviation amount, and determining the allocation priority of the server object using the second deviation amount and the deviation trend data is performed.
9. The task scheduling method according to claim 1, wherein: Before determining the hardware resource requirements of the task to be processed, the method further includes: When it is determined that the usage of any hardware resource by any server object reaches a preset threshold, the step of determining the demand for the hardware resource by the task to be processed is entered; When it is determined that the usage of any hardware resource by any server object does not reach a preset threshold, the to-be-processed task is sent to a server object with idle hardware resources.
10. The task scheduling method according to claim 1, wherein: Also includes: Obtain the number of server nodes in the server cluster, and obtain the number of network cards and processors in the server nodes; The number of server objects is determined by using the number of server nodes, the number of network cards, and the number of processors, and the server objects are created according to the number of server objects, so as to form the server objects by using each network card device and each processor.
11. The task scheduling method according to claim 1, wherein: Determining the hardware resource requirements of the task to be processed includes: The demand amount of the task to be processed for the hardware resource is predicted according to the first demand amount of the task to be processed for the hardware resource at the historical specified moment.
12. The task scheduling method according to claim 11, characterized in that: Also includes: Determine whether the usage of its own hardware resources has reached the preset threshold; If yes, randomly generate the demand quantity for the task to be processed; If not, the process proceeds to the step of predicting the demand for hardware resources by the task to be processed according to the first demand for hardware resources by the task to be processed at the specified historical moment.
13. The task scheduling method according to any one of claims 1 to 12, characterized in that: At least two data processing units form a data processing unit cluster, and jointly dispatch the to-be-processed task to the target server object.
14. The task scheduling method according to claim 13, wherein: For at least two server objects in a server cluster, determining a first usage and a second usage of a hardware resource by the server objects at a first moment and a second moment, respectively, determining an expected usage using the second usage and the demand, and determining a first deviation and a second deviation between a target usage of the hardware resource and the first usage and the expected usage, respectively, including: determining, for at least two server objects in the server node to which the server object belongs, a first usage and a second usage of a hardware resource by the server objects at a first moment and a second moment, respectively, determining an expected usage using the second usage and the demand, and determining a first deviation and a second deviation between a target usage of the hardware resource and the first usage and the expected usage, respectively; The step of determining deviation trend data using the first deviation and the second deviation, and determining the allocation priority of the server object using the second deviation and the deviation trend data, includes: Determining deviation trend data using the first deviation amount and the second deviation amount, and determining an initial allocation priority of a server object in the server node to which the server object belongs using the second deviation amount and the deviation trend data; Establishing a correspondence between the server object and the initial assigned priority, and writing the correspondence into a shared memory; Determining a target server object for executing the task to be processed according to the allocation priority of each server object includes: Reading the corresponding relationships written by the shared memory and other data processing units, and sorting the server objects in each corresponding relationship according to the initial allocation priority in each corresponding relationship; The target server object for executing the task to be processed is determined according to the sorting result.
15. The task scheduling method according to claim 14, characterized in that: Sending the pending task to the target server object includes: Determine whether the target server object is located at the server node to which it belongs; If yes, the pending task is sent to the target server object; If not, the pending task is ignored.
16. The task scheduling method according to claim 14, characterized in that: A shared memory area is provided in the memory area of each of the data processing units; Writing the corresponding relationship into the shared memory includes: Writing the corresponding relationship into its own shared memory area; The data in the shared memory area is exchanged with other data processing units through the network to write the corresponding relationship in its own shared memory area into the shared memory area of other data processing units, and the corresponding relationship in the shared memory area of other data processing units is written into its own shared memory area.
17. A task scheduling device, characterized in that: Applied to a data processing unit, the device comprises: A task receiving module, configured to receive tasks to be processed and determine the hardware resource requirements of the tasks to be processed; a resource usage determination module configured to determine, for at least two server objects in a server cluster, first and second usages of hardware resources by the server objects at a first moment and a second moment, respectively, determine an expected usage using the second usage and the demand, and determine a first deviation and a second deviation between a target usage of the hardware resource and the first usage and the expected usage, respectively; wherein the first moment is a historically designated moment and the second moment is a task issuance moment; an allocation priority determination module, configured to determine deviation trend data using the first deviation amount and the second deviation amount, and determine an allocation priority of the server object using the second deviation amount and the deviation trend data; The task allocation module is used to determine the target server object for executing the task to be processed according to the allocation priority of each server object, and send the task to be processed to the target server object.
18. A data processing unit, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the task scheduling method according to any one of claims 1 to 16 when executing the computer program.
19. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the task scheduling method according to any one of claims 1 to 16 is implemented.
20. A non-volatile computer-readable storage medium, characterized in that The non-volatile computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are loaded and executed by the processor, the task scheduling method according to any one of claims 1 to 16 is implemented.
Citation Information
Patent Citations
Method for saving energy of data center under cloud environment
CN104199736A
Stock trading platform load balancing control system
CN116302582A
Resource scheduling method and device, equipment, storage medium and program product
CN119149196A
Intelligent server management method and system in edge computing environment
CN119576594A
Resource scheduling method of server and electronic equipment
CN120256114A