Task scheduling method and device based on reinforcement learning and electronic equipment

By adopting a task scheduling method based on reinforcement learning in the cloud computing environment, decompose tasks, identify dependencies and optimize resource allocation, the problems of low task processing efficiency and neglect of dependencies in the existing technology are solved, and efficient tasks and optimized scheduling are achieved.

CN120029733APending Publication Date: 2025-05-23BEIJING C&W ELECTRONICS GRP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510078480.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing task scheduling methods are difficult to adapt to dynamically changing resource states and task requirements in cloud computing environments, resulting in low task processing efficiency and neglecting the dependencies between tasks, affecting the overall processing efficiency.

Method used

The task scheduling method based on reinforcement learning is adopted. By obtaining the target task and decomposing it into multiple subtasks, the dependency information of the subtasks and the environment information of the surrounding cloud servers are obtained. The reinforcement learning method is used to prioritize the subtasks, and the subtasks are allocated in priority order to achieve efficient task scheduling.

Benefits of technology

By decomposing tasks, identifying dependencies and optimizing resource allocation, task processing efficiency is improved, task waiting and resource competition are avoided, and optimal task scheduling and optimal resource utilization are achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029733A_ABST
    Figure CN120029733A_ABST
Patent Text Reader

Abstract

The invention discloses a task scheduling method and device based on reinforcement learning and electronic equipment, and relates to the field of data processing. The method comprises the steps of obtaining a target task, and decomposing the target task into a plurality of sub-tasks; task dependency information of each subtask and environment information of surrounding cloud servers are obtained, wherein the surrounding cloud servers are other cloud servers in the communication range; based on the task dependency information and the environment information, adopting a reinforcement learning method to carry out priority ranking on each sub-task to obtain a task priority sequence; according to a task priority sequence, distributing a corresponding surrounding cloud server to each sub-task to obtain a task scheduling sequence and a corresponding relationship between the sub-tasks and the surrounding cloud servers; and executing task scheduling according to the task scheduling sequence and the corresponding relationship. By implementing the technical scheme provided by the invention, the task processing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and specifically to a task scheduling method, device and electronic equipment based on reinforcement learning. Background Art

[0002] With the rapid development of big data and cloud computing technologies, the growing demand for data processing has driven the widespread application of cloud computing platforms. Cloud computing technology, with its efficient resource management and powerful computing capabilities, provides flexible services for enterprises of all sizes. These services not only include data storage and processing, but also cover complex task scheduling and resource allocation issues. In a cloud environment, task scheduling has become a key technology to achieve optimal resource utilization and improve service response efficiency.

[0003] At present, although the existing task scheduling methods can handle resource allocation and task management problems to a certain extent, they usually face several major defects. First, many traditional scheduling algorithms are mainly based on static rules, which are difficult to adapt to the dynamically changing resource status and task requirements in cloud computing environments, resulting in low efficiency in task processing. Secondly, these methods often ignore the dependencies between tasks, resulting in scheduling results that may not be optimal, affecting the overall processing efficiency. Therefore, the existing task scheduling methods have the problem of low efficiency.

[0004] Therefore, there is an urgent need for a task scheduling method, device and electronic device based on reinforcement learning. Summary of the invention

[0005] The present application provides a task scheduling method, device and electronic equipment based on reinforcement learning, which improves task processing efficiency.

[0006] In the first aspect of the present application, a task scheduling method based on reinforcement learning is provided, the method comprising: obtaining a target task and decomposing the target task into multiple subtasks; obtaining task dependency information of each of the subtasks and environmental information of surrounding cloud servers, the surrounding cloud servers being other cloud servers within the communication range; based on the task dependency information and the environmental information, using a reinforcement learning method to prioritize each of the subtasks to obtain a task priority order; according to the task priority order, assigning a corresponding surrounding cloud server to each of the subtasks to obtain a task scheduling order and a corresponding relationship between the subtasks and the surrounding cloud servers; and performing task scheduling according to the task scheduling order and the corresponding relationship.

[0007] By adopting the above technical solution, by obtaining the target task and decomposing it into multiple subtasks, complex tasks can be simplified into multiple subtasks, improving the task processing efficiency. At the same time, by obtaining the dependency information of the subtasks and the environmental information of the surrounding cloud servers, the constraint conditions and available resources for task execution can be comprehensively understood, providing a necessary information basis for subsequent task scheduling. On this basis, using the reinforcement learning method to prioritize the subtasks can fully consider the dependency relationship between tasks and the resource requirements of tasks, obtain a reasonable task priority order, and avoid the situation of task waiting or resource competition. Assigning the corresponding surrounding cloud servers to the subtasks according to the task priority order can realize the parallel execution of tasks on multiple surrounding cloud servers, give full play to the advantages of distributed computing, and greatly shorten the overall completion time of tasks. Finally, performing task scheduling according to the obtained task scheduling order and the corresponding relationship between the subtasks and the surrounding cloud servers can ensure that tasks are completed orderly and efficiently according to the optimal plan, achieving the purpose of optimizing task execution performance and efficiency.

[0008] Optionally, based on the task dependency information and the environmental information, using the reinforcement learning method to prioritize each of the subtasks to obtain a task priority order specifically includes: determining a first subtask and a second subtask according to the task dependency information, the first subtask and the second subtask being any two subtasks with dependencies among the multiple subtasks, and the second subtask being the successor task of the first subtask; calculating the average communication time between the first subtask and the second subtask; establishing a pessimistic cost matrix according to the average communication time; calculating the subtask priorities of each of the subtasks based on the pessimistic cost matrix; and training using the Sarsa algorithm based on the subtask priorities to obtain the task priority order.

[0009] By adopting the above technical solution, by determining any two subtasks with a dependency relationship according to the task dependency information, the sequential constraints between tasks can be accurately identified, providing a basis for subsequent task sorting. Calculating the average communication time between these two subtasks can quantify the data transmission overhead between tasks and evaluate the pros and cons of task sorting and scheduling schemes. On this basis, establishing a pessimistic cost matrix can estimate the maximum completion time of each subtask when executed on different servers as a whole, providing a reference for the calculation of task priorities. By comprehensively considering the pessimistic cost matrix and the average execution time of the subtasks, a priority metric that takes into account both task importance and execution efficiency can be obtained, making the priority sorting more comprehensive and accurate. Finally, using the Sarsa algorithm to perform reinforcement learning training on the subtask priorities can adaptively adjust the task priorities through continuous trial and error and optimization, obtain a dynamically optimal task priority order, and improve the overall performance of task scheduling.

[0010] Optionally, the calculation formula for calculating the average communication time of the first subtask and the second subtask is specifically: ; Among them, i is the i-th subtask, j is the j-th subtask, is the average communication time, is the average communication startup time, data i,j is the amount of data transferred between the ith subtask and the jth subtask, is the average transmission rate of each of the surrounding cloud servers.

[0011] By adopting the above technical solution and providing a specific average communication time calculation formula, the communication overhead between any two subtasks can be accurately estimated based on the data transmission volume between subtasks, the average communication startup time and the average transmission rate of the cloud server.

[0012] Optionally, the formula for establishing a pessimistic cost matrix according to the average communication time is specifically: ; Among them, t i represents the i-th subtask, the i-th subtask is the first subtask, t j represents the jth subtask, the jth subtask is the second subtask, t j t i The successor task, succ(t i ) represents the i-th subtask t i The set of direct successor tasks, P is the cloud server set P={p 1 , p 2 , ..., p m}; p k is the kth cloud server, p m For the mth cloud server, PCT(t i , p k ) are the values ​​in the pessimistic cost matrix, PCT(t i , p k ) means that when the i-th subtask t i Select the kth cloud server p k When i The maximum value of the longest path from the successor to the end task; wherein the end task is a subtask without a subsequent task, and w(j, l) is the jth subtask on the cloud server p l The processing time on PCT(t j , p m ) means that when the jth subtask tj Select the mth cloud server p m When j The maximum value of the longest path from the successor to the end task.

[0013] By adopting the above technical solution and providing a specific pessimistic cost matrix calculation formula, the longest execution time required from the start of each subtask to the end of the entire task set can be estimated when each subtask is executed on different servers. This estimation takes into account the difference in execution time of subtasks on different servers, as well as the dependencies and communication overhead between tasks, so it can more comprehensively and accurately evaluate the pros and cons of task sorting and scheduling schemes. In the calculation process, a pessimistic strategy is adopted, that is, data transmission must occur between subtasks, so the estimated result obtained is an upper bound, which provides a certain margin for task scheduling and helps to ensure that the task can be completed within the expected time. At the same time, by recursively calculating the maximum end time of the task, the critical path in the task set can be identified.

[0014] Optionally, the formula for calculating the priority of each subtask based on the pessimistic cost matrix is ​​specifically: ; Among them, rank PCT (t i ) is the i-th subtask t i The priority of the task is w(i, k), which is the priority of the i-th subtask on processor P. k The processing time on , m is the number of cloud servers.

[0015] By adopting the above technical solution and providing a specific subtask priority calculation formula, the pessimistic execution time and average execution time of subtasks on different cloud servers can be comprehensively considered to obtain a comprehensive priority measurement. Among them, the pessimistic execution time reflects the impact of the task on the completion time of the entire task set, and the average execution time reflects the computational amount of the task itself and the demand for resources. The combination of the two can more accurately evaluate the importance and urgency of the task. By calculating the priority of each subtask, a complete task priority sequence can be obtained, which provides a direct basis for the subsequent task sorting. By scheduling tasks in order from high to low priority, the tasks that have the greatest impact on the completion time of the entire task set can be executed first, key resources can be released as soon as possible, and favorable execution conditions can be created for subsequent tasks, thereby shortening the completion time of the entire task set and improving the overall efficiency of task scheduling.

[0016] Optionally, based on the subtask priority, the task priority order is obtained by training with the Sarsa algorithm, specifically including: taking the subtask priority as the immediate reward value of the Sarsa algorithm; determining the successor subtask of each subtask according to the task dependency information, and updating the Q table of the Sarsa algorithm according to the successor subtask to obtain the number of training rounds; judging whether the Q table converges or the number of training rounds is greater than or equal to a preset number of training rounds; if it is determined that the Q table converges or the number of training rounds is greater than or equal to a preset number of training rounds, stopping updating the Q table to obtain a target Q table; according to the target Q table, determining the successor subtask with the largest Q value for each subtask to obtain the task priority order.

[0017] By adopting the above technical solution and using the Sarsa algorithm to perform reinforcement learning training on task priorities, the task priorities can be adaptively adjusted under given task dependencies and environmental information to obtain a dynamically optimal task sorting solution. Taking the subtask priority as an immediate reward value can guide the Sarsa algorithm in the direction of the shortest overall execution time. During the training process, by constantly trying different priority sortings and updating the priority strategy according to the actual task completion time, the Sarsa algorithm can gradually find a near-optimal task sorting solution. By setting the number of training rounds and convergence conditions, the training time and accuracy of the algorithm can be controlled, and a balance between time efficiency and sorting quality can be achieved in practical applications. The final target Q table contains the optimal successor tasks that should be selected under different task states, which can directly guide online task scheduling, quickly give the optimal task priority, and improve the real-time and continuity of task scheduling.

[0018] Optionally, the method further includes: recording the task allocation relationship between each of the subtasks and each of the surrounding cloud servers, and storing the task allocation relationship in a database; and updating the remaining resource information of each of the surrounding cloud servers in the database.

[0019] By adopting the above technical solution, by recording the task allocation relationship and storing it in the database, it is possible to easily track and query the execution of tasks, and provide a basis for abnormal detection during task execution. The remaining resource information of the cloud server is also updated in the database, which can be matched with the task allocation relationship, comprehensively record the system status during task scheduling, provide real-time resource constraint information for task scheduling, and ensure the feasibility and efficiency of scheduling decisions.

[0020] In the second aspect of the present application, a task scheduling device based on reinforcement learning is provided, which includes: an acquisition module and a processing module, wherein: the acquisition module is used to acquire a target task and decompose the target task into multiple subtasks; the acquisition module is also used to acquire task dependency information of each of the subtasks and environmental information of surrounding cloud servers, and the surrounding cloud servers are other cloud servers within the communication range; the processing module is used to prioritize each of the subtasks based on the task dependency information and the environmental information using a reinforcement learning method to obtain a task priority order; the processing module is also used to assign corresponding surrounding cloud servers to each of the subtasks according to the task priority order, and obtain a task scheduling order and a corresponding relationship between the subtasks and the surrounding cloud servers; the processing module is also used to perform task scheduling according to the task scheduling order and the corresponding relationship.

[0021] In the third aspect of the present application, an electronic device is provided, including a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes any one of the methods described above.

[0022] In a fourth aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores instructions, and when the instructions are executed, any of the methods described above is executed.

[0023] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. By obtaining the target task and decomposing it into multiple subtasks, complex tasks can be simplified into multiple subtasks, thereby improving the efficiency of task processing. At the same time, by obtaining the dependency information of the subtasks and the environmental information of the surrounding cloud servers, the constraints and available resources of task execution can be fully understood, providing the necessary information basis for subsequent task scheduling. On this basis, the reinforcement learning method is used to prioritize the subtasks, which can fully consider the dependencies between tasks and the demand for resources of tasks, obtain a reasonable task priority order, and avoid the situation of task waiting or resource competition. According to the task priority order, the corresponding surrounding cloud servers are assigned to the subtasks, which can realize the parallel execution of tasks on multiple surrounding cloud servers, give full play to the advantages of distributed computing, and greatly shorten the overall completion time of the task. Finally, according to the obtained task scheduling order and the corresponding relationship between the subtasks and the surrounding cloud servers, the task scheduling can be performed to ensure that the tasks are completed in an orderly and efficient manner according to the optimal solution, so as to achieve the purpose of optimizing the task execution performance.

[0024] 2. By determining any two dependent subtasks based on task dependency information, the order constraints between tasks can be accurately identified, providing a basis for subsequent task sorting. By calculating the average communication time of the two subtasks, the data transmission overhead between tasks can be quantified, and the pros and cons of task sorting and scheduling schemes can be evaluated. On this basis, a pessimistic cost matrix is ​​established, which can estimate the maximum completion time of each subtask when executed on different servers as a whole, providing a reference for the calculation of task priority. By comprehensively considering the pessimistic cost matrix and the average execution time of subtasks, a priority metric that takes into account both task importance and execution efficiency can be obtained, making priority sorting more comprehensive and accurate. Finally, the Sarsa algorithm is used to perform reinforcement learning training on subtask priorities. Through continuous trial and error and optimization, the task priority can be adaptively adjusted to obtain a dynamic optimal task priority order, thereby improving the overall performance of task scheduling.

[0025] 3. By providing a specific pessimistic cost matrix calculation formula, the longest execution time required from the start of each subtask to the end of the entire task set can be estimated when each subtask is executed on different servers. This estimation takes into account the difference in execution time of subtasks on different servers, as well as the dependencies and communication overhead between tasks, so it can more comprehensively and accurately evaluate the pros and cons of task sorting and scheduling schemes. In the calculation process, a pessimistic strategy is adopted, that is, data transmission must occur between subtasks, so the estimated result is an upper bound, which provides a certain margin for task scheduling and helps ensure that the task can be completed within the expected time. At the same time, by recursively calculating the maximum end time of the task, the critical path in the task set can be identified. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a flowchart of a task scheduling method based on reinforcement learning disclosed in an embodiment of the present application; Figure 2 It is a module schematic diagram of a task scheduling device based on reinforcement learning disclosed in an embodiment of the present application; Figure 3 It is a structural schematic diagram of an electronic device disclosed in an embodiment of the present application.

[0027] Explanation of the reference numerals: 201, acquisition module; 202, processing module; 300, electronic device; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. DETAILED DESCRIPTION

[0028] In order to enable technicians in this field to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.

[0029] In the description of the embodiments of the present application, words such as "for example" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "for example" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "for example" or "for example" is intended to present related concepts in a specific way.

[0030] In the description of the embodiments of the present application, the meaning of the term "multiple" refers to two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals. In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "include", "comprise", "have" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.

[0031] This application provides a task scheduling method based on reinforcement learning. Figure 1 , Figure 1 1 is a flow chart of a task scheduling method based on reinforcement learning provided in an embodiment of the present application. The method is applied to a cloud server, which is a virtual server that executes a task scheduling program in a cloud computing environment. The method includes steps S101 to S104, which are as follows: Step S101: Obtain a target task and decompose the target task into multiple subtasks.

[0032] In step S101, the cloud server first obtains the target task to be executed. The target task is usually a complex task with large computational workload. The cloud server can obtain the target task in a variety of ways, such as direct submission by the user. The user submits the task to be executed to the cloud server in a certain format (such as JSON, XML, etc.) and specifies the basic information of the task, such as the task type, input data, expected output results, etc. The cloud platform manages a global task queue, sorts the tasks submitted by the user or generated by the system according to certain rules (such as priority, submission time, etc.) and stores them in the queue. After obtaining the target task, the cloud server analyzes and splits it. Since the target task is usually more complex, it is difficult and time-consuming to execute directly, so it is divided into multiple smaller, independent subtasks. The subtasks can be parallel or there can be dependencies between them, that is, some subtasks must be executed after other subtasks are completed.

[0033] The specific implementation method of task splitting depends on the type and characteristics of the task. For example, for a large-scale data processing task, it can be split according to the logical division or physical distribution of the data. Assuming that a 10GB text file needs to be counted for word frequency, we can divide the file into 1,000 sub-files by line or size, each sub-file is about 10MB, and then create a sub-task for counting word frequency for each sub-file. The sub-tasks can be executed on multiple computing nodes at the same time, each counting the word frequency of the file it is responsible for, and finally the results are merged to count the word frequency information of the entire file.

[0034] In practice, task splitting is usually handled by a dedicated task scheduling module, which will comprehensively consider the characteristics of the task, the distribution of data, and the situation of computing resources, and use a predefined splitting strategy to generate an execution plan for the subtasks.

[0035] Step S102: Obtain task dependency information of each subtask and environmental information of surrounding cloud servers, where the surrounding cloud servers are other cloud servers within the communication range.

[0036] In step S102, the cloud server obtains dependency information between subtasks. Dependency information describes the execution order and data dependency between subtasks. A task can be represented as a directed acyclic graph (DAG), where nodes represent subtasks and edges represent dependency relationships. If the execution of subtask A depends on the output of subtask B, a directed edge from B to A is added to the graph.

[0037] The cloud server can obtain the dependency information of tasks in various ways. For example, when the user explicitly specifies, the user clearly provides the dependency relationship between subtasks when submitting a task, such as specifying that a certain subtask must be executed after another subtask is completed. At the same time, the cloud server can also automatically analyze and infer. The cloud server automatically infers the dependency relationship between subtasks by statically analyzing or dynamically tracing the execution process of tasks. For example, if the input data of subtask A comes from the output of subtask B, it can be inferred that A depends on B.

[0038] After obtaining the task dependency information, the cloud server obtains the environmental information of the surrounding cloud servers. The environmental information reflects the available computing resources during task execution, mainly including the types and quantities of computing resources, such as the specifications and numbers of resources such as CPU, GPU, memory, and storage; the current load and utilization rate of computing resources, such as CPU utilization rate, memory occupancy rate, and network bandwidth usage.

[0039] The cloud server communicates with the surrounding cloud servers and exchanges the environmental information of the surrounding cloud servers regularly to understand the computing environment of the entire cluster. For example, the server can use the heartbeat mechanism to regularly report its own resource usage to a central coordination node, and the coordination node aggregates this information to form a global resource view for the scheduler to query.

[0040] Take a specific example to illustrate: Suppose there are 3 cloud servers S1, S2, and S3 in the current system, and each server has different CPU and memory configurations. Currently, a target task T is split into 5 subtasks t1, t2, t3, t4, and t5, where t2 depends on the completion of t1 and t3, t4 depends on t2, and t5 depends on t4. The cloud server first obtains and stores these task dependencies, which can be represented as a matrix. Then, the cloud server collects the current resource usage of S1, S2, and S3 through the heartbeat mechanism, such as the CPU utilization rate of S1 is 80% and the memory occupancy is 60%, those of S2 are 50% and 70%, and those of S3 are 30% and 40%.

[0041] Step S103: Based on the task dependency information and environmental information, use the reinforcement learning method to prioritize each subtask to obtain the task priority order.

[0042] In step S103, based on the task dependency information and the environment information, a reinforcement learning method is used to prioritize each subtask to obtain a task priority order, which specifically includes: determining the first subtask and the second subtask according to the task dependency information, the first subtask and the second subtask being any two subtasks with dependencies among multiple subtasks; calculating the average communication time between the first subtask and the second subtask; establishing a pessimistic cost matrix based on the average communication time; calculating the subtask priority of each subtask based on the pessimistic cost matrix; and obtaining the task priority order based on the subtask priority by using the Sarsa algorithm for training.

[0043] In a possible implementation manner, the calculation formula for calculating the average communication time of the first subtask and the second subtask is specifically as follows: ; Among them, i is the i-th subtask, j is the j-th subtask, is the average communication time, is the average communication startup time, is the average transmission rate of each surrounding cloud server.

[0044] Specifically, the cloud server determines each pair of dependent subtasks (called the first subtask i and the second subtask j) according to the task dependency relationship; where the second subtask j is the successor task of the first subtask i. Then, the average communication time between them is calculated using the above method. is the average communication time, is the average communication startup time, data i,j is the amount of data transferred between two subtasks, is the average transmission rate between each cloud server. This step takes into account the data transmission overhead when the task is executed on different servers.

[0045] For example, assuming that subtasks t1 and t2 have a dependency relationship, the output data volume of t1 is 500MB, the average communication startup time is 0.1 seconds, and the average transmission rate between cloud servers is 100MB / s, then the average communication time between t1 and t2 is 0.1+500 / 100=5.1 seconds.

[0046] In a possible implementation, the formula for establishing the pessimistic cost matrix according to the average communication time is specifically: ; Among them, t i represents the i-th subtask, the i-th subtask is the first subtask, t j represents the jth subtask, the jth subtask is the second subtask, succ(t i ) represents the i-th subtask ti The set of direct successor tasks, P is the cloud server set P={p 1 , p 2 , ..., p m}; p k is the kth cloud server, p m For the mth cloud server, PCT(t i , p k ) are the values ​​in the pessimistic cost matrix, PCT(t i , p k ) means that when the i-th subtask t i Select the kth cloud server p k When i The maximum value of the longest path from the successor to the end task; where the end task is a subtask without a subsequent task, and w(j, l) is the jth subtask on the cloud server p l The processing time on PCT(t j , p m ) means that when the jth subtask t j Select the mth cloud server p m When j The maximum value of the longest path from the successor to the end task.

[0047] Specifically, using the average communication time between tasks as the weight, the cloud server constructs an n×m matrix PCM, where n is the number of subtasks and m is the number of cloud servers. The element PCT(t i , p k ) means: if the i-th subtask t i Assigned to the kth cloud server p k , then from t i The maximum estimated execution time from the start to the end of the entire task. The PCT value is calculated using a recursive formula: where t j Yes i The subsequent task, w j,l Yes j In the cloud server l The formula recursively calculates the execution time of i is the maximum estimated completion time allocated to each cloud server in the root subtree. Intuitively, PCT(t i , p k ) is calculated by considering two aspects: task t i It is in the cloud server p k The execution time w(t i, p k ) and task t i Each direct successor task tj At its best processor p l The estimated completion time on the project, PCT(t j , p l )), plus t i With t j The possible communication time between t is the average communication time of the above steps. i The execution time on the current server, the completion time of each successor task on its best server, and the communication time with the successor task are finally taken as the maximum value in various cases. This calculation method can provide a pessimistic but tight upper bound on the completion time.

[0048] In a possible implementation, based on the pessimistic cost matrix, the formula for calculating the priority of each subtask is specifically: ; Among them, rank PCT (t i ) is the i-th subtask t i The priority of the task is w(i, k), which is the priority of the i-th subtask on processor P. k The processing time on , m is the number of cloud servers.

[0049] Specifically, based on the PCM matrix, the cloud server uses the above formula to calculate each subtask t i Priority The priority consists of two items: the first item is t i The average PCT value on all servers reflects the impact of the task on the critical path; the second term is t i The average processing time on all servers reflects the computational effort of the task itself. Intuitively, the higher the priority of a task, the more likely it is to be a critical task that affects the overall execution time and should be executed first.

[0050] In one possible implementation, based on the subtask priority, the Sarsa algorithm is used for training to obtain a task priority order, specifically including: using the subtask priority as the immediate reward value of the Sarsa algorithm; determining the successor subtask of each subtask based on the task dependency information, and updating the Q table of the Sarsa algorithm based on the successor subtask to obtain the number of training rounds; judging whether the Q table converges or the number of training rounds is greater than or equal to the preset number of training rounds; if it is determined that the Q table converges or the number of training rounds is greater than or equal to the preset number of training rounds, stopping updating the Q table to obtain a target Q table; according to the target Q table, determining the successor subtask with the largest Q value for each subtask to obtain the task priority order.

[0051] Specifically, the cloud server determines the set of successor subtasks for each subtask based on the task dependencies. Then, the Q table is initialized. The Q table is a two-dimensional matrix with rows representing states (i.e., current tasks) and columns representing actions (i.e., selected successor tasks). The matrix element Q(s, a) represents the long-term value estimate of taking action a in state s. The initial value of the Q table can be set to 0 or a small random number. At the same time, a maximum number of training rounds N and a Q table convergence threshold ε are set. Next, the cloud server starts the iterative training process. In each round of iteration: First, starting from the initial task, according to the current Q table, the ε-greedy strategy is used to select the next subtask. That is, the successor task with the largest Q value is selected with a probability of 1-ε, and a successor task is randomly selected with a probability of ε. This can achieve a balance between exploring new strategies and using existing strategies.

[0052] Next, execute the selected successor task and observe the immediate reward r (i.e., the priority of the subtask) and the new state s' of the environment (i.e., the current task becomes the task that has just been executed).

[0053] Then, according to the update formula of the Sarsa algorithm, the Q(s, a) value in the Q table is updated: the updated Q(s, a) is Q(s, a)+α[r+γQ(s', a')-Q(s, a)]; Among them, s is the state before executing a, a is the actual action taken in state s, s' and a' are the new state after executing a and the new action selected in state s'. α is the learning rate, which controls the amplitude of each update; γ is the discount factor, which controls the importance of future rewards. If the new state s' is a terminal state (that is, there is no successor task), return to the initial state; otherwise, set the current state s to s', the current action a to a', and repeat the above steps. After each round of iteration, determine whether the Q table converges (for example, the change in Q value is less than ε) or whether the number of training rounds reaches the preset number of training rounds N. If so, the training ends; otherwise, start the next round of iteration.

[0054] For example, suppose the initial task order is t 1 , t 2 , t 3 , t 4 , in a certain round, the Sarsa algorithm chooses to exchange t 1 and t 2 The order of , get the new order t 2 , t 1 , t 3 , t 4 According to the priority calculation formula, the total priority of the new sort is 70, which is higher than the initial sort of 69. Therefore, Q(t 1 , t 2 , t3 , t 4 ), exchange t 1 , t 2 The task ordering will be updated and improved according to the difference in priority between the old and new orderings. After multiple rounds of exploration, the Sarsa algorithm will converge to a task ordering with the highest total priority.

[0055] Step S104: assigning corresponding surrounding cloud servers to each subtask in order of task priority, and obtaining a task scheduling order and a corresponding relationship between the subtasks and the surrounding cloud servers.

[0056] In step S104, the cloud server builds a server queue based on the environmental information. The server queue contains all the surrounding cloud servers that are currently available for executing tasks and are arranged in a certain order. The sorting methods include sorting according to the computing power of the server, such as CPU frequency, memory size and other indicators. The server with stronger computing power is ranked higher. Sort according to the current load of the server. The server with lighter load is ranked higher to give priority to the use of idle resources. Sort according to the network distance between the server and the main server. The closer the distance, the higher the server is to reduce communication overhead. The cloud server can select a combination of one or more sorting methods according to actual needs. At the same time, the cloud server also builds an ordered task queue according to the task priority order. The subtask with the highest priority is ranked at the head of the queue, and so on. Next, the cloud server starts from the head of the task queue and assigns execution servers to the subtasks one by one. For each subtask, the cloud server adopts the following strategy: first check whether the subtask has a predecessor task, and if so, check which server its predecessor task is assigned to. If all predecessor tasks have been assigned, the current subtask is assigned to the server where the predecessor task is located first. This can reduce data transmission between tasks and improve execution efficiency. If the predecessor task has not been assigned, or the predecessor task is scattered across multiple servers, select the server with the highest ranking that meets the task resource requirements from the server queue. If the remaining resources of the selected server are not enough to execute the current subtask, continue to check the next server in the queue until a server that meets the requirements is found. Assign the subtask to the selected server and record this assignment relationship. Update the server's remaining resource information and adjust the order of the server queue if necessary. The cloud server repeats the above process until all subtasks are assigned to the execution server. Finally, a complete task scheduling solution is obtained, including the execution order and execution server of each subtask.

[0057] For example, suppose there are 5 subtasks {T1, T2, T3, T4, T5}, and their priority order is T1>T2>T3>T4>T5. There are also 3 surrounding cloud servers {S1, S2, S3}, and their initial order is S1>S2>S3. First, cloud servers are allocated starting from T1. Assuming that T1 has no predecessor task, it is directly assigned to the highest-ranked S1 server. Next is T2. Assuming that T2 depends on T1, T2 is also assigned to the S1 server so that it can be executed on the same server as T1. For T3, assuming that its predecessor tasks T1 and T2 are both on S1, but the remaining resources of S1 are not enough to execute T3, the suboptimal S2 server is selected to execute T3. The allocation of T4 and T5 is also carried out according to the same principle until all tasks have clear execution servers. The final scheduling scheme may be: T1->S1; T2->S1; T3->S2; T4->S2; T5->S3.

[0058] Step S105: Execute task scheduling according to the task scheduling sequence and the corresponding relationship.

[0059] In step S105, the cloud server distributes the subtasks to the designated execution servers in sequence according to the task scheduling order generated in S104 and the mapping relationship between the subtasks and the surrounding cloud servers, and coordinates the communication and synchronization between the cloud servers to ensure that the tasks are correctly executed according to the predetermined order and dependencies.

[0060] In a possible implementation, the method further includes: recording the task allocation relationship between each subtask and each surrounding cloud server, and storing the task allocation relationship in a database; and updating the remaining resource information of each surrounding cloud server in the database.

[0061] Specifically, the cloud server designs a task allocation relationship table in the database to store which execution server each subtask is assigned to. In addition to recording the task allocation relationship, the cloud server updates the remaining resource information of each cloud server in real time. To this end, a server resource information table is designed in the database to record the hardware configuration and current available resources of each cloud server; the current available resources include the current CPU utilization and the current memory utilization.

[0062] Reference Figure 2The present application also provides a task scheduling device based on reinforcement learning, which is a cloud server, and the cloud server includes an acquisition module 201 and a processing module 202; wherein: the acquisition module 201 is used to acquire the target task and decompose the target task into multiple subtasks; the acquisition module 201 is also used to acquire the task dependency information of each subtask and the environmental information of the surrounding cloud servers, and the surrounding cloud servers are other cloud servers within the communication range; the processing module 202 is used to prioritize each subtask based on the task dependency information and the environmental information, and obtain the task priority order by using the reinforcement learning method; the processing module 202 is also used to assign the corresponding surrounding cloud servers to each subtask according to the task priority order, and obtain the task scheduling order and the corresponding relationship between the subtasks and the surrounding cloud servers; the processing module 202 is also used to perform task scheduling according to the task scheduling order and the corresponding relationship.

[0063] In a possible implementation, the processing module 202 prioritizes each subtask based on the task dependency information and the environment information by using a reinforcement learning method to obtain a task priority sequence, specifically including: the processing module 202 determines the first subtask and the second subtask according to the task dependency information, the first subtask and the second subtask are any two subtasks with dependencies among multiple subtasks, and the second subtask is the successor task of the first subtask; the processing module 202 calculates the average communication time between the first subtask and the second subtask; the processing module 202 establishes a pessimistic cost matrix based on the average communication time; the processing module 202 calculates the subtask priority of each subtask based on the pessimistic cost matrix; the processing module 202 obtains the task priority sequence based on the subtask priority by using the Sarsa algorithm training.

[0064] In a possible implementation manner, the calculation formula for calculating the average communication time of the first subtask and the second subtask by the processing module 202 is specifically as follows: ; Among them, i is the i-th subtask, j is the j-th subtask, is the average communication time, is the average communication startup time, data i,j is the amount of data transferred between the ith subtask and the jth subtask, is the average transmission rate of each surrounding cloud server.

[0065] In a possible implementation manner, the processing module 202 establishes a pessimistic cost matrix according to the average communication time. Specifically, the formula is: ; Among them, t i represents the i-th subtask, the i-th subtask is the first subtask, tj represents the jth subtask, the jth subtask is the second subtask, t j t i The successor task, succ(t i ) represents the i-th subtask t i The set of direct successor tasks, P is the cloud server set P={p 1 , p 2 , ..., p m}; p k is the kth cloud server, p m For the mth cloud server, PCT(t i , p k ) are the values ​​in the pessimistic cost matrix, PCT(t i , p k ) means that when the i-th subtask t i Select the kth cloud server p k When i The maximum value of the longest path from the successor to the end task; where the end task is a subtask without a subsequent task, and w(j, l) is the jth subtask on the cloud server p l The processing time on PCT(t j , p m ) means that when the jth subtask t j Select the mth cloud server p m When j The maximum value of the longest path from the successor to the end task.

[0066] In a possible implementation, the processing module 202 calculates the priority of each subtask based on the pessimistic cost matrix using the following formula: ; Among them, rank PCT (t i ) is the i-th subtask t i The priority of the task is w(i, k), which is the priority of the i-th subtask on processor P. k The processing time on , m is the number of cloud servers.

[0067] In one possible implementation, the processing module 202 uses the Sarsa algorithm for training to obtain a task priority order based on the subtask priority, specifically including: the processing module 202 uses the subtask priority as the immediate reward value of the Sarsa algorithm; the processing module 202 determines the successor subtask of each subtask based on the task dependency information, and updates the Q table of the Sarsa algorithm based on the successor subtask to obtain the number of training rounds; the processing module 202 determines whether the Q table converges or the number of training rounds is greater than or equal to the preset number of training rounds; if the processing module 202 determines that the Q table converges or the number of training rounds is greater than or equal to the preset number of training rounds, the processing module 202 stops updating the Q table to obtain a target Q table; the processing module 202 determines the successor subtask with the largest Q value for each subtask based on the target Q table to obtain the task priority order.

[0068] In a possible implementation, it further includes: the processing module 202 records the task allocation relationship between each subtask and each surrounding cloud server, and stores the task allocation relationship in a database; and updates the remaining resource information of each surrounding cloud server in the database.

[0069] It should be noted that: when the device provided in the above embodiment realizes its function, only the division of the above functional modules is used as an example. In actual application, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.

[0070] The present application also provides an electronic device. Figure 3 , Figure 3 The electronic device 300 may include: at least one processor 301 , at least one network interface 304 , a user interface 303 , a memory 305 , and at least one communication bus 302 .

[0071] The communication bus 302 is used to realize the connection and communication between these components.

[0072] The user interface 303 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 303 may also include a standard wired interface and a wireless interface.

[0073] The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).

[0074] Among them, the processor 301 may include one or more processing cores. The processor 301 uses various interfaces and lines to connect various parts in the entire server, and executes various functions of the server and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 305, and calling data stored in the memory 305. Optionally, the processor 301 can be implemented in at least one hardware form of digital signal processing (Digital Signal Processing, DSP), field programmable gate array (Field-Programmable Gate Array, FPGA), and programmable logic array (Programmable Logic Array, PLA). The processor 301 can integrate one or a combination of a central processing unit (Central Processing Unit, CPU), a graphics processing unit (Graphics Processing Unit, GPU) and a modem. Among them, the CPU mainly processes the operating system, user interface and application programs; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 301, and it can be implemented separately through a chip.

[0075] Among them, the memory 305 may include a random access memory (Random Access Memory, RAM) and may also include a read-only memory (Read-Only Memory). Optionally, the memory 305 includes a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 305 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 305 may optionally also be at least one storage device located away from the aforementioned processor 301. Refer to Figure 3 , the memory 305 as a computer storage medium may include an operating system, a network communication module, a user interface module and an application program of a task scheduling method based on reinforcement learning.

[0076] exist Figure 3In the electronic device 300 shown, the user interface 303 is mainly used to provide an input interface for the user and obtain the data input by the user; and the processor 301 can be used to call the application program storing a task scheduling method based on reinforcement learning in the memory 305, and when executed by one or more processors 301, the electronic device 300 executes one or more of the methods described in the above embodiments. It should be noted that for the aforementioned method embodiments, for the sake of simple description, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required for the present application.

[0077] The present application also provides a computer-readable storage medium, which stores instructions. When executed by one or more processors 301, the electronic device 300 executes one or more of the methods described in the above embodiments.

[0078] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0079] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are only schematic, such as the division of units, which is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.

[0080] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0081] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0082] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a memory and includes several instructions for a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned memory includes: various media that can store program codes, such as USB flash drives, mobile hard drives, magnetic disks or optical disks.

[0083] The above is only an exemplary embodiment of the present disclosure and cannot be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made according to the teachings of the present disclosure are still within the scope of the present disclosure. After considering the disclosure of the specification and the truth of practice, those skilled in the art will easily think of other embodiments of the present disclosure.

[0084] This application is intended to cover any variation, use or adaptation of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary technical means in the art not described in the present disclosure. The description and examples are to be regarded as exemplary only, and the scope and spirit of the present disclosure are defined by the claims.

Claims

1. A task scheduling method based on reinforcement learning, characterized in that: The method comprises: Obtaining a target task, and decomposing the target task into multiple subtasks; Obtaining task dependency information of each of the subtasks and environmental information of surrounding cloud servers, where the surrounding cloud servers are other cloud servers within the communication range; Based on the task dependency information and the environment information, a reinforcement learning method is used to prioritize each of the subtasks to obtain a task priority sequence; According to the task priority order, a corresponding surrounding cloud server is assigned to each of the subtasks to obtain a task scheduling order and a corresponding relationship between the subtasks and the surrounding cloud servers; The task scheduling is performed according to the task scheduling sequence and the corresponding relationship.

2. The method according to claim 1, characterized in that: The step of using a reinforcement learning method to prioritize each of the subtasks based on the task dependency information and the environment information to obtain a task priority sequence specifically includes: Determine a first subtask and a second subtask according to the task dependency information, wherein the first subtask and the second subtask are any two subtasks having dependencies among the plurality of subtasks, and the second subtask is a successor task of the first subtask; Calculating average communication time of the first subtask and the second subtask; Establishing a pessimistic cost matrix based on the average communication time; Based on the pessimistic cost matrix, calculating the subtask priority of each of the subtasks; Based on the subtask priorities, the task priority sequence is obtained by using Sarsa algorithm training.

3. The method according to claim 2, characterized in that The calculation formula for calculating the average communication time of the first subtask and the second subtask is specifically: ; Among them, i is the i-th subtask, j is the j-th subtask, is the average communication time, is the average communication startup time, data i,j is the amount of data transferred between the ith subtask and the jth subtask, is the average transmission rate of each of the surrounding cloud servers.

4. The method according to claim 3, characterized in that The formula for establishing the pessimistic cost matrix according to the average communication time is specifically: ; Among them, t i represents the i-th subtask, the i-th subtask is the first subtask, t j represents the jth subtask, the jth subtask is the second subtask, t j t i The successor task, succ(t i ) represents the i-th subtask t i The set of direct successor tasks, P is the cloud server set P = {p1, p2, ..., p m }; p k is the kth cloud server, p m For the mth cloud server, PCT(t i , p k ) are the values ​​in the pessimistic cost matrix, PCT(t i , p k ) means that when the i-th subtask t i Select the kth cloud server p k When i The maximum value of the longest path from the successor to the end task; wherein the end task is a subtask without a subsequent task, and w(j, l) is the jth subtask on the cloud server p l The processing time on PCT(t j , p m ) means that when the jth subtask t j Select the mth cloud server p m When j The maximum value of the longest path from the successor to the end task.

5. The method according to claim 4, characterized in that The formula for calculating the priority of each subtask based on the pessimistic cost matrix is ​​specifically: ; Among them, rank PCT (t i ) is the i-th subtask t i The priority of the task is w(i, k), which is the priority of the i-th subtask on processor P. k The processing time on , m is the number of cloud servers.

6. The method according to claim 2, characterized in that Based on the subtask priorities, the task priority sequence is obtained by using the Sarsa algorithm training, which specifically includes: Using the subtask priority as the instant reward value of the Sarsa algorithm; Determine the subsequent subtasks of each of the subtasks according to the task dependency information, and update the Q table of the Sarsa algorithm according to the subsequent subtasks to obtain the number of training rounds; Determine whether the Q table converges or the number of training rounds is greater than or equal to a preset number of training rounds; If it is determined that the Q table converges or the number of training rounds is greater than or equal to the preset number of training rounds, then stopping updating the Q table to obtain a target Q table; According to the target Q table, the subsequent subtask with the largest Q value is determined for each of the subtasks to obtain the task priority order.

7. The method according to claim 1, characterized in that The method further comprises: Recording the task allocation relationship between each of the subtasks and each of the surrounding cloud servers, and storing the task allocation relationship in a database; The remaining resource information of each of the surrounding cloud servers is updated into the database.

8. A task scheduling device based on reinforcement learning, characterized in that: The device comprises an acquisition module (201) and a processing module (202), wherein: The acquisition module (201) is used to acquire a target task and decompose the target task into multiple subtasks; The acquisition module (201) is further used to acquire task dependency information of each of the subtasks and environmental information of surrounding cloud servers, wherein the surrounding cloud servers are other cloud servers within the communication range; The processing module (202) is used to prioritize each of the subtasks based on the task dependency information and the environment information using a reinforcement learning method to obtain a task priority sequence; The processing module (202) is further used to allocate a corresponding surrounding cloud server to each of the subtasks according to the task priority order, and obtain a task scheduling order and a corresponding relationship between the subtasks and the surrounding cloud servers; The processing module (202) is further configured to execute task scheduling according to the task scheduling sequence and the corresponding relationship.

9. An electronic device, characterized in that: The electronic device (300) comprises a processor (301), a memory (305), a user interface (303) and a network interface (304), wherein the memory (305) is used to store instructions, the user interface (303) and the network interface (304) are used to communicate with other devices, and the processor (301) is used to execute the instructions stored in the memory (305) so that the electronic device (300) executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1 to 7 is performed.