Task scheduling method and device, equipment, medium and program product

By combining processor load and network connection quality information with the target intelligent agent, the target processor and resource quantity of the task are determined, which solves the problems of processing pressure on terminal devices and data transmission latency, and realizes efficient and reliable task scheduling and system performance improvement.

CN121029331APending Publication Date: 2025-11-28CHINA MOBILE COMM CORP GUANGXI CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410675073.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-28
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Terminal devices face enormous processing pressure when handling massive computing tasks, and data transmission consumes resources and causes latency when offloading tasks to edge servers or cloud servers. How to comprehensively consider the load data of each processor to determine the task offloading target is an urgent technical problem to be solved.

Method used

By comprehensively considering the processor's load information and the network connection quality information between the terminal device and the processor, the target processor and resource quantity of the task are determined by the target agent, and the task scheduling is performed using the pruned and trained agent.

Benefits of technology

It achieves efficient and reliable task scheduling, can find the global optimal solution, improve system efficiency, and reduce resource consumption and latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121029331A_ABST
    Figure CN121029331A_ABST
Patent Text Reader

Abstract

The invention discloses a task scheduling method and device, equipment, a medium and a program product, and belongs to the field of data processing.The task scheduling method comprises the steps that a first task, a task queue where the first task is located, load data of a first candidate processor and network connection quality information between terminal equipment and other candidate processors are obtained, the first candidate processor comprises terminal equipment, at least one edge server and a cloud server; the first task, the task queue, the load data and the network connection quality information are input into a target agent, a first decision result is obtained, and the first decision result comprises a first target processor for processing the first task and a first target resource quantity allocated by the terminal equipment for the first task; allocating a first target resource quantity to the first task; and scheduling the first task to a first target processor for processing by using the first target resource quantity. Efficient and reliable task scheduling can be realized, and the system efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of data processing, and particularly relates to a task scheduling method and device, equipment, medium and program product. BACKGROUND

[0002] With the development of science and technology, terminal equipment generates massive data and computing tasks every day. If the massive computing tasks are all placed in the local of the terminal equipment for processing, it will cause huge processing pressure on the terminal equipment, and the scarce resources of the terminal equipment are also difficult to process complex computing-intensive tasks. Based on this, the terminal equipment often unloads computing tasks to edge servers or cloud servers for processing, but unloading computing tasks to edge servers or cloud servers for processing needs data transmission, which will consume resources and cause time delay. Therefore, how to comprehensively consider the load data of each processor to determine which processor to unload the task to for processing is a technical problem to be solved. SUMMARY

[0003] The embodiments of the present application provide a task scheduling method, device, equipment, medium and program product, which can realize efficient and reliable task scheduling and improve system efficiency.

[0004] In a first aspect, the embodiments of the present application provide a task scheduling method applied to a terminal equipment, and the task scheduling method comprises:

[0005] obtaining load data of a first task, a task queue where the first task is located, a first candidate processor, and network connection quality information between the terminal equipment and other candidate processors, the first candidate processor comprising the terminal equipment, at least one edge server and a cloud server, and the other candidate processors being processors other than the terminal equipment in the first candidate processor;

[0006] inputting the first task, the task queue where the first task is located, the load data of the first candidate processor and the network connection quality information between the terminal equipment and the other candidate processors into a target agent to obtain a first decision result, the first decision result comprising a first target processor processing the first task and a first target resource amount allocated by the terminal equipment for the first task, the target agent being an agent obtained by training an initial agent according to training samples by the cloud server, the training samples comprising a plurality of second tasks, a task queue where the second tasks are located, load data of second candidate processors, network connection quality information between the terminal equipment and other candidate processors, and a second decision result, the other candidate processors being processors other than the terminal equipment in the second candidate processors;

[0007] allocating the first target resource amount for the first task;

[0008] The first task is scheduled to the first target processor for processing by using the first target resource amount.

[0009] In some embodiments of the present application, before the first task, the task queue where the first task is located, the load data of the first candidate processor, and the network connection quality information between the terminal device and other candidate processors are input into the target agent to obtain the first decision result, the method further comprises:

[0010] The target agent sent by the cloud server is received, and the target agent is an agent obtained by the cloud server by performing pruning training on the trained agent according to a set pruning threshold and iteration number.

[0011] In some embodiments of the present application, the first task, the task queue where the first task is located, the load data of the first candidate processor, and the network connection quality information between the terminal device and other candidate processors are input into the target agent to obtain the first decision result, comprising:

[0012] According to the first task, the load data of the first candidate processor, and the network connection quality data between the terminal device and other candidate processors, at least one performance index corresponding to the processing of the first task by the terminal device, at least one edge server and the cloud server is calculated respectively;

[0013] According to at least one performance index and at least one performance index of other tasks in the task queue corresponding to the first task, the first decision result is determined.

[0014] In some embodiments of the present application, the at least one performance index includes the time delay required for processing the first task and the resource quantity required by the terminal device, and according to the first task, the load data of the first candidate processor, and the network connection quality data between the terminal device and other candidate processors, at least one performance index corresponding to the processing of the first task by the terminal device, at least one edge server and the cloud server is calculated respectively, comprising:

[0015] According to the first task and the load data of the terminal device, the first time delay required for processing the first task by the terminal device and the first resource quantity required by the terminal device are calculated;

[0016] According to the first task, the load data of at least one edge server, and the network connection quality data between the terminal device and at least one edge server, the second time delay required for processing the first task by at least one edge server and the second resource quantity required by the terminal device are calculated;

[0017] According to the first task, the load data of the cloud server, and the network connection quality data between the terminal device and the cloud server, a third time delay required when the first task is scheduled to the cloud server for processing and a third resource quantity required by the terminal device are calculated;

[0018] According to at least one performance indicator and at least one performance indicator of other tasks in a task queue corresponding to the first task, a first decision result is determined, including:

[0019] According to the first time delay, the first resource quantity, the second time delay, the second resource quantity, the third time delay, the third resource quantity, and the time delay and resource quantity required by each of the other tasks in the task queue corresponding to the first task, the first decision result is determined.

[0020] In some embodiments of the present application, according to the first time delay, the first resource quantity, the second time delay, the second resource quantity, the third time delay, the third resource quantity, and the time delay and resource quantity required by each of the other tasks in the task queue corresponding to the first task, the first decision result is determined, including:

[0021] According to the first time delay, the first resource quantity, the time delay and resource quantity required by each of the other tasks in the task queue, and the first weight value corresponding to the time delay and the second weight value corresponding to the resource quantity, a first reward value corresponding to the case that the first task is scheduled to the terminal device for processing is calculated;

[0022] According to the second time delay, the second resource quantity, the time delay and resource quantity required by each of the other tasks in the task queue, and the first weight value corresponding to the time delay and the second weight value corresponding to the resource quantity, a second reward value corresponding to the case that the first task is scheduled to at least one edge server for processing is calculated;

[0023] According to the third time delay, the third resource quantity, the time delay and resource quantity required by each of the other tasks in the task queue, and the first weight value corresponding to the time delay and the second weight value corresponding to the resource quantity, a third reward value corresponding to the case that the first task is scheduled to the cloud server for processing is calculated;

[0024] According to the first reward value, the second reward value, and the third reward value, a first decision result corresponding to the first task when a long-term average reward function is maximized is determined.

[0025] In a second aspect, embodiments of the present application provide a task scheduling method applied to a cloud server, the task scheduling method comprising:

[0026] obtain a training sample, the training sample comprising a plurality of second tasks, a task queue in which the second tasks are located, load data of a second candidate processor, network connection quality information between the terminal device and other candidate processors, and a second decision result, the second candidate processor comprising the terminal device, at least one edge server, and a cloud server, the other candidate processors being processors other than the terminal device in the second candidate processor, and the second decision result comprising a second target processor for processing the second tasks and a second target resource amount allocated by the terminal device for the second tasks;

[0027] train the initial agent according to the training sample to obtain a trained agent;

[0028] determine the trained agent as a target agent;

[0029] send the target agent to the terminal device;

[0030] in a case where the terminal device determines, by using the target agent, that a first target processor for processing a first task is the cloud server, receive the first task sent by the terminal device;

[0031] process the first task to obtain a processing result;

[0032] send the processing result to the terminal device.

[0033] In some embodiments of the present application, before the trained agent is determined as the target agent, the method further comprises:

[0034] set a pruning threshold and a number of iterations;

[0035] prune the trained agent according to the pruning threshold and the number of iterations to obtain a pruned agent;

[0036] determine the trained agent as the target agent, comprising:

[0037] determine the pruned agent as the target agent.

[0038] In some embodiments of the present application, pruning the trained agent according to the pruning threshold and the number of iterations to obtain a pruned agent comprises:

[0039] step A: take the trained teacher model as a first student model, and prune the first student model; when a threshold of model iteration pruning does not reach the pruning threshold, cyclically execute the following steps B and C, and the trained teacher model is the trained agent:

[0040] Step B: Continue to train the teacher model and store the state-action pairs, including the task, the queue where the task is located, the load data of the candidate processor, the network connection quality data between the terminal device and the candidate processor, and the decision result, into the experience replay pool;

[0041] Step C: Sample the first student model from the experience replay pool and adjust the policy distillation using the divergence loss function;

[0042] Step D: Measure the sparsity of each layer of the pruned first student model to obtain the model sparsity;

[0043] Step E: Create a second student model based on the model sparsity;

[0044] Step F: Iterate steps G to H N times to train the second student model using the teacher's experience;

[0045] Step G: Continue to train the teacher model and store the state-action pairs into the experience replay pool;

[0046] Step H: Sample the second student model from the experience replay pool to train the second student model;

[0047] Step I: Calculate the size difference between the first student model and the second student model, and if the size difference is less than the pruning threshold, use the second student model as the target agent, and if the size difference is greater than or equal to the pruning threshold, repeat steps A to I.

[0048] In a third aspect, the embodiments of the present application provide a task scheduling device applied to a terminal device, the task scheduling device comprising:

[0049] The acquisition module is configured to acquire a first task, a task queue where the first task is located, a first candidate processor, and network connection quality information between the terminal device and other candidate processors, the first candidate processor including the terminal device, at least one edge server, and a cloud server, and the other candidate processors being processors other than the terminal device in the first candidate processor;

[0050] The decision module is used to input the first task, the task queue in which the first task is located, the first candidate processor, and the network connection quality information between the terminal device and other candidate processors into the target agent to obtain the first decision result. The first decision result includes the first target processor that processes the first task and the first target resource amount allocated by the terminal device for the first task. The target agent is the agent obtained by the cloud server after training the initial agent based on the training samples. The training samples include multiple second tasks, the task queue in which the second tasks are located, the second candidate processors, the network connection quality information between the terminal device and other candidate processors, and the second decision result. The other candidate processors are the processors other than the terminal device among the second candidate processors.

[0051] The allocation module is used to allocate the first target resource amount to the first task.

[0052] The scheduling module is used to schedule the first task to the first target processor for processing using the first target resource quantity.

[0053] Fourthly, embodiments of this application provide a task scheduling device applied to a cloud server. The task scheduling device includes:

[0054] The acquisition module is used to acquire training samples. The training samples include multiple second tasks, the task queue where the second tasks are located, the load data of the second candidate processors, the network connection quality information between the terminal device and other candidate processors, and the second decision result. The second candidate processors include the terminal device, at least one edge server, and a cloud server. The second decision result includes the second target processor for processing the second task and the second target resource amount allocated by the terminal device for the second task. Other candidate processors are processors other than the terminal device among the second candidate processors.

[0055] The first training module is used to train the initial agent based on the training samples to obtain the trained agent.

[0056] The determination module is used to identify the trained agent as the target agent.

[0057] The sending module is used to send the target intelligent agent to the terminal device;

[0058] The receiving module is used to receive the first task sent by the terminal device when the terminal device determines, using the target intelligent agent, that the first target processor for processing the first task is a cloud server.

[0059] The processing module is used to process the first task and obtain the processing result;

[0060] The sending module is also used to send the processing results to the terminal device.

[0061] Fifthly, embodiments of this application provide a task scheduling device, the device including: a processor and a memory storing computer program instructions;

[0062] The processor implements the task scheduling method of any of the above embodiments when executing computer program instructions.

[0063] Sixthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the task scheduling method of any of the above embodiments.

[0064] In a seventh aspect, embodiments of this application provide a computer program product, wherein instructions in the computer program product, when executed by a processor of an electronic device, cause the electronic device to perform the task scheduling method of any of the above embodiments.

[0065] The task scheduling method, apparatus, device, medium, and program product according to embodiments of this application can rely on a target intelligent agent to make intelligent decisions, comprehensively consider the processor's load information and the network connection quality information between the terminal device and the processor to determine the first target processor for processing the first task and the first target resource amount allocated by the terminal device for the first task, and then perform task scheduling based on the decision results. This enables efficient and reliable task scheduling and improves system performance. Attached Figure Description

[0066] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0067] Figure 1 A flowchart illustrating a task scheduling method provided in an embodiment of this application;

[0068] Figure 2 A flowchart illustrating another task scheduling method provided in an embodiment of this application;

[0069] Figure 3 This is a schematic diagram of the structure of a task scheduling device provided in an embodiment of this application;

[0070] Figure 4 This is a schematic diagram of another task scheduling device provided in an embodiment of this application;

[0071] Figure 5 This is a schematic diagram of the structure of the task scheduling device provided in the embodiments of this application. Detailed Implementation

[0072] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.

[0073] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.

[0074] With the development of science and technology, terminal devices generate massive amounts of data and computational tasks every day. If these massive computational tasks were all processed locally on the terminal device, it would place an enormous processing burden on it, and the scarce resources of the terminal device would struggle to handle complex, computationally intensive tasks. Therefore, terminal devices often offload computational tasks to edge servers or cloud servers. However, offloading computational tasks to edge servers or cloud servers requires data transmission, which consumes resources and causes latency. Therefore, determining which processor to offload a task to by comprehensively considering the load data of each processor is a pressing technical challenge that needs to be solved.

[0075] Related technologies provide task scheduling strategies based on heuristic algorithms. Traditional heuristic-based task scheduling strategies work by optimizing the task scheduling problem to find the optimal or near-optimal solution, thus achieving efficient task scheduling. This algorithm, based on heuristic rules or heuristic search strategies, evaluates candidate solutions and iteratively searches for the optimal one. While this algorithm can find relatively good solutions, it may not guarantee finding the globally optimal solution. The algorithm needs to be selected and adjusted according to the specific problem and application scenario to achieve the best scheduling performance.

[0076] This application provides a task scheduling method, apparatus, device, medium, and program product. It relies on a target intelligent agent to make intelligent decisions, comprehensively considering processor load information and network connection quality information between the terminal device and the processor to determine a first target processor for processing a first task and a first target resource amount allocated by the terminal device for the first task. Task scheduling can then be performed based on the decision results. This achieves efficient and reliable task scheduling, finds the globally optimal solution, and improves system performance.

[0077] Figure 1 A flowchart illustrating a task scheduling method provided in an embodiment of this application;

[0078] refer to Figure 1 This application introduces a task scheduling method provided by an embodiment, applied to a terminal device. The task scheduling method includes:

[0079] S110, obtain the first task, the task queue where the first task is located, the load data of the first candidate processor, and the network connection quality information between the terminal device and other candidate processors. The first candidate processor includes the terminal device, at least one edge server, and a cloud server. The other candidate processors are processors other than the terminal device among the first candidate processors.

[0080] S120, the first task, the task queue in which the first task is located, the load data of the first candidate processor, and the network connection quality information between the terminal device and other candidate processors are input into the target intelligent agent to obtain the first decision result. The first decision result includes the first target processor that processes the first task and the first target resource amount allocated by the terminal device for the first task. The target intelligent agent is the intelligent agent obtained by the cloud server after training the initial intelligent agent based on the training samples. The training samples include multiple second tasks, the task queue in which the second task is located, the load data of the second candidate processors, the network connection quality information between the terminal device and other candidate processors, and the second decision result. The other candidate processors are the processors other than the terminal device among the second candidate processors.

[0081] S130, allocate the first target resource amount for the first task;

[0082] S140, the first task is scheduled to the first target processor for processing using the first target resource quantity.

[0083] According to the task scheduling method provided in this application, intelligent decision-making can be performed by a target intelligent agent. This method comprehensively considers the processor's load information and the network connection quality information between the terminal device and the processor to determine the first target processor for processing the first task and the first target resource amount allocated by the terminal device for the first task. Task scheduling can then be performed based on the decision results. This method achieves efficient and reliable task scheduling, finds the globally optimal solution, and improves system performance.

[0084] Regarding S110 above, at the beginning of each time slice, the terminal device stores the newly generated tasks into its maintained task queue. At the beginning of each time slot, the user equipment generates new tasks following a Bernoulli probability distribution, with each time slice having a length of δ. In each time slot, one task is retrieved from the task queue and processed sequentially according to the FIFO principle. The user equipment needs to maintain its own task queue Q. n ={Task1,Task2,…,Task N}

[0085] This application expresses the calculation of the unloading decision as a Markov decision equation: where the environment state is a quadruple containing information about the task to be processed, the task queue, the load status of candidate servers (including terminal devices, at least one edge server, and cloud server), and the network connection quality information between the terminal devices and other candidate servers.

[0086] state = (Task k Q n ,L e ,L I )

[0087] The first task, the task queue in which the first task is located, the load data of the first candidate processor, and the network connection quality information between the terminal device and other candidate processors constitute the environment state quadruple. The environment state quadruple can be directly obtained after the terminal device generates a task.

[0088] Regarding S120 above, the target agent can decide whether to process the first task locally on the terminal device or offload it to an available edge server or the cloud, and determine the number of energy units, i.e., the amount of resources, allocated to the user device. The decision-making action space is:

[0089] action = (c k ,e k )

[0090] c k ∈{0,1,2,…,|E|,|E|+1}

[0091] e k ∈{0,1,2,…,enmax}

[0092] Where |E| represents the number of edge servers that can provide services; en max This represents the maximum energy unit that a user device can allocate. action = 0 indicates that the task is executed locally on the terminal device, action ∈ {1,2,…,|E|} indicates that the task is offloaded to an edge server for execution, and action = |E|+1 indicates that the task is offloaded to the cloud for processing.

[0093] In some embodiments of this application, the target intelligent agent can calculate at least one performance index corresponding to scheduling the first task to the terminal device, at least one edge server, and the cloud server for processing, based on the first task, the load data of the first candidate processor, and the network connection quality data between the terminal device and other candidate processors.

[0094] The first decision result is determined based on at least one performance metric and at least one performance metric of other tasks in the task queue corresponding to the first task.

[0095] In some embodiments of this application, at least one performance metric includes the latency required to process the first task and the amount of resources required by the terminal device.

[0096] The step of calculating at least one performance metric corresponding to scheduling the first task to the terminal device, at least one edge server, and a cloud server for processing, based on the load data of the first task, the first candidate processor, and the network connection quality data between the terminal device and other candidate processors, includes:

[0097] Based on the first task and the load data of the terminal device, calculate the first latency required to schedule the first task to the terminal device for processing and the first amount of resources required by the terminal device.

[0098] Based on the first task, the load data of at least one edge server, and the network connection quality data between the terminal device and at least one edge server, calculate the second latency required to schedule the first task to at least one edge server for processing and the second amount of resources required by the terminal device.

[0099] Based on the first task, the load data of the cloud server, and the network connection quality data between the terminal device and the cloud server, calculate the third latency required to schedule the first task to the cloud server for processing and the amount of third resources required by the terminal device.

[0100] The first decision result is determined based on at least one performance metric and at least one performance metric of other tasks in the task queue corresponding to the first task, including:

[0101] The first decision result is determined based on the first delay, the first resource quantity, the second delay, the second resource quantity, the third delay, the third resource quantity, and the delay and resource quantity required by other tasks in the task queue corresponding to the first task.

[0102] The calculations for latency and resource usage are as follows:

[0103] Tasks in the task queue are retrieved using the letter 'k', and are denoted as Task. k = (B k D k ). Among them B k D represents the amount of data (i.e., the number of bytes in the program) that needs to be transferred to unload the computing task. k Represents the task processing task k Number of calculation cycles required.

[0104] When offloading tasks from user device n to edge server s, the uplink transmission rate of the user device can be calculated using the following formula, and the same formula applies to calculating the uplink transmission rate from the edge server to the cloud:

[0105]

[0106] W is the uplink bandwidth, H s,n P represents the channel gain caused by path loss and shadow attenuation when user equipment n offloads tasks to edge service node s, ω0 is the background noise during transmission, and i and j are used to index the edge service node and user equipment, respectively. s,n This refers to the power used by terminal device n to transmit data when it offloads a task to edge server s.

[0107]

[0108] Where e k It is the energy unit allocated for data transmission. Maximum data transmission power is a constant determined by the hardware conditions of the device.

[0109] The task scheduling algorithm needs to perform task unloading actions c in time slice k. k and the number of energy units allocated e k The decision is made based on two factors. The energy cost of the user equipment is equivalent to the number of energy units allocated.

[0110] energy k =e k

[0111] When offloading a task to be processed locally on the terminal device:

[0112] The required latency is: using Indicates Task k The queuing delay from task generation to processing, when the task is executed locally on the terminal device, is:

[0113]

[0114] Among them, f local Calculate the CPU frequency for user devices.

[0115] The amount of resources required, i.e., the amount of resources needed to process the task locally. k The energy expenditure in the process is expressed as:

[0116]

[0117] Where ε represents the computational cycle that can be operated by consuming a unit of energy.

[0118] When offloading tasks to edge servers for processing:

[0119] The required latency is: the total latency for processing the task. It mainly consists of the queuing latency of the task locally, the transmission latency from user device n to edge server s, the task computation latency, and the transmission latency from the edge server to return the computation results. composition:

[0120]

[0121] Among them, f edge This represents the CPU frequency used by the edge server for computing tasks. This represents the transmission latency spent by user device n transmitting task data to edge server s:

[0122]

[0123] The required resources, i.e. the energy cost of data transmission from user device n to edge server s, are expressed as:

[0124]

[0125] Task k The total energy cost of offloading processing to edge servers is expressed as follows:

[0126]

[0127] When offloading tasks to a cloud server for processing:

[0128] The required latency is: Total latency It can be calculated using the following formula:

[0129]

[0130] in, The transmission latency incurred by terminal device n in submitting task data to cloud c:

[0131]

[0132] Where r(c,n) represents the rate at which the edge user device transmits task data to the cloud, and its calculation method is similar to that of the transmission rate between the user device and the edge server. cloud The CPU frequency allocated to the cloud for executing tasks.

[0133] The energy overhead of data transmission from user device n to edge server c is expressed as:

[0134]

[0135] Task k The total energy cost of offloading processing to edge servers is expressed as follows:

[0136]

[0137] In some embodiments of this application, a first decision result is determined based on a first delay, a first resource quantity, a second delay, a second resource quantity, a third delay, a third resource quantity, and the delay and resource quantity required by other tasks in the task queue corresponding to the first task, including:

[0138] Based on the first latency, the first resource quantity, the latency and resource quantity required by other tasks in the task queue, the first weight value corresponding to the latency, and the second weight value corresponding to the resource quantity, calculate the first reward value corresponding to scheduling the first task to the terminal device for processing.

[0139] Based on the second latency, the second resource quantity, the latency and resource quantity required by other tasks in the task queue, the first weight value corresponding to the latency, and the second weight value corresponding to the resource quantity, calculate the second reward value corresponding to scheduling the first task to at least one edge server for processing.

[0140] Based on the third latency, the third resource quantity, the latency and resource quantity required by other tasks in the task queue, the first weight value corresponding to the latency, and the second weight value corresponding to the resource quantity, calculate the third reward value corresponding to scheduling the first task to the cloud server for processing.

[0141] Based on the first reward value, the second reward value, and the third reward value, determine the first decision result corresponding to the first task that maximizes the long-term average reward function.

[0142] Specifically, continuing with the formula above, the first decision result can be explained by the following formula.

[0143] The reward function of the target agent is calculated as follows:

[0144]

[0145] In the above formula and Weights are assigned to latency overhead and energy overhead, respectively, and follow the principle of...

[0146] The agent's goal is to find a policy π: state → action that maximizes the long-term average reward function, where T is the total number of time slices in the studied time period.

[0147]

[0148] Regarding the above S130 and S140, after determining the first target processor and the first target resource quantity corresponding to the first task, task scheduling can be performed accordingly. When the first target processor is a terminal device, it means that the task is executed locally by the terminal device. When the first target processor is a server among at least one edge server, the first task is unloaded to the corresponding server for processing, and the processing result fed back by the edge server is received. When the first target processor is a cloud server, the first task is unloaded to the cloud server for processing, and the processing result fed back by the cloud server is received.

[0149] In some embodiments of this application, before inputting the first task, the task queue in which the first task is located, the first candidate processor, and the network connection quality information between the terminal device and other candidate processors into the target agent to obtain the first decision result, the method further includes:

[0150] Receive the target agent sent by the cloud server. The target agent is the agent obtained by the cloud server after training the agent according to the set pruning threshold and number of iterations.

[0151] On the one hand, redundant parameters of the agent's neural network can be removed while maintaining inference accuracy, achieving size compression. This makes it easier to deploy on user devices with limited storage resources. On the other hand, the pruned agent requires fewer computing resources during decision-making and inference, while speeding up inference, making it more suitable for user devices with limited computing resources. The pruning and training process of the agent will be described in the following embodiments.

[0152] The aforementioned target agent is solved and optimized based on the Rainbow algorithm. The Rainbow algorithm combines multiple techniques, including experience replay, dual Q-learning, priority experience replay, duel network structure, and N-step learning, to improve the learning efficiency and performance of the reinforcement learning agent. Through the combination of these key elements, the Rainbow algorithm can better address the problems existing in traditional reinforcement learning algorithms, helping the agent make more accurate decisions in various complex environments.

[0153] The specific steps are as follows:

[0154] 1. Input: Minibatch sample size k for each experience replay, step size η, replay period K, total time T

[0155] 2: Initialize the experience replay pool D, with priority P = {p1, p2, ...}, constant β, and TD bias δ. j Q and Q target The network parameters are θ and θ target Sample weight w, change in sample weight Δ

[0156] 3: Observe the initial state (state(0)) and select action (action(0)).

[0157] 4: Time t changes from 1 to T, entering a loop.

[0158] 5: Observe state(t) and R(t)

[0159] 6: Store the data (state(t-1), action(t-1), state(t), R(t)) into the experience replay pool D.

[0160] 7: Perform an experience replay every K steps.

[0161] 8: Collect k samples sequentially, loop through one minibatch, and use j as an indicator.

[0162] 9: Sampling a sample point j according to the probability distribution ~ P(j) = p j / ∑ i p i

[0163] 10: Calculate the importance weight w of this sample. j =(N·P(j)) -β / max w

[0164] 11: Update the TD bias for this sample point.

[0165] δ j =R j +Q target (s j argmaxa Q(s j ,a))-Q(s j-1 ,a j-1 )

[0166] 12: Update the priority p of this sample point based on the TD deviation. j =|δ j |

[0167] 13: Cumulative weight change

[0168] 14: End the current sample weighting and sample the next sample.

[0169] 15: Update network parameters θ = θ + η·Δ, and reset Δ = 0.

[0170] 16: Update the weights θ of the target policy network by step size. target =θ

[0171] 17: End of update

[0172] 18: Choose the next action based on the new strategy.

[0173] 19: Apply the new action to the environment, obtain new data, and enter a new cycle.

[0174] Figure 2 A flowchart illustrating another task scheduling method provided in an embodiment of this application;

[0175] refer to Figure 2 This application provides a task scheduling method based on embodiments, applied to a cloud server. The task scheduling method includes:

[0176] S210, Obtain training samples. The training samples include multiple second tasks, the task queue where the second tasks are located, the load data of the second candidate processors, the network connection quality information between the terminal device and other candidate processors, and the second decision result. The second candidate processors include the terminal device, at least one edge server, and a cloud server. The other candidate processors are processors other than the terminal device among the second candidate processors. The second decision result includes the second target processor for processing the second task and the second target resource amount allocated by the terminal device for the second task.

[0177] S220, The initial agent is trained based on the training samples to obtain the trained agent;

[0178] S230, the trained agent is identified as the target agent;

[0179] S240, the target intelligent agent is sent to the terminal device;

[0180] S250, when the terminal device determines, using the target intelligent agent, that the first target processor for processing the first task is a cloud server, it receives the first task sent by the terminal device.

[0181] S260, process the first task and obtain the processing result;

[0182] S270 sends the processing result to the terminal device.

[0183] In some embodiments of this application, the method further includes, before using the trained agent as the target agent:

[0184] Set the trimming threshold and number of iterations;

[0185] Based on the pruning threshold and the number of iterations, the trained agent is pruned and trained to obtain the pruned and trained agent.

[0186] The trained agent is used as the target agent, including:

[0187] Use the trained agent as the target agent.

[0188] In some embodiments of this application, the trained agent is pruned according to a pruning threshold and the number of iterations to obtain a pruned trained agent, including:

[0189] Step A: Use the trained teacher model as the first student model and prune it. When the pruning threshold for model iteration does not reach the pruning threshold, repeat steps B and C below. The trained teacher model becomes the trained agent.

[0190] Step B: Continue training the teacher model and store the state-action pairs in the experience replay pool. The state-action pairs include the task, the queue in which the task is located, the load data of the candidate processor, the network connection quality data between the terminal device and the candidate processor, and the decision result.

[0191] Step C: Sample from the experience replay pool and adjust the first student model using strategy distillation with the divergence loss function;

[0192] Step D: Measure the sparsity of each layer of the clipped first student model to obtain the model sparsity;

[0193] Step E: Create a second student model based on model sparsity;

[0194] Step F: Iterate steps G through H N times to train the second student model using the teacher's experience;

[0195] Step G: Continue training the teacher model and store the state-action pairs in the experience replay pool;

[0196] Step H: Sample and train the second student model from the experience replay pool;

[0197] Step I: Calculate the size difference between the first student model and the second student model. If the size difference is less than the pruning threshold, use the second student model as the target agent. If the size difference is greater than or equal to the pruning threshold, repeat Step A to Step I.

[0198] Specifically, the lightweight pruning of the agent consists of four main stages: creating a teacher model, experience replay, iterative model pruning, and model compression. First, creating the Rainbow algorithm requires training a large-scale sparse network model Q... teacher The accumulated experience samples are then iteratively trained and stored in the experience replay pool D. Let... These are state-action experience replay pairs generated by the teacher model. These are the actions generated by the student model (i.e., the lightweight model that needs to be trained). The network parameters of the student model are adjusted using the KL divergence metric to match the desired behavior. and The output distribution between them.

[0199] The model pruning and shrinking steps are trained based on samples in the experience replay pool and do not directly interact with the environment. In each iteration, we base our work on Q... teacher The model is pruned based on accumulated experience, and then fine-tuned using policy distillation. In each time slot, the sparsity of each layer after pruning satisfies the following:

[0200]

[0201] t is the cutting round, g t It is the sparsity of the model in this round, g initial and g final These represent the initial sparsity of the model and the final sparsity after pruning, respectively, where n is the sparsity of g. t Reaching g final The number of cutting rounds, where Δ is the cutting frequency.

[0202] At regular intervals, the model pruning module evaluates the model based on its performance. If the model's optimization result for the objective function falls below a predetermined minimum, the model is fine-tuned using policy distillation with the KL-divergence loss function and trained until it recovers and maximizes the long-term average reward function.

[0203] The model compression stage involves creating a small, dense neural network based on the output of the model iterative pruning module and training it. After training, the model size is compared with the model built in the previous iteration: if the difference is less than a predefined threshold T, the pruning algorithm has converged; otherwise, the above process is repeated. The specific process of model pruning is as follows:

[0204] Input: The trained teacher model Q teacher ;

[0205] Output: The final compressed model Q ans ;

[0206] initialization:

[0207] 1: Initialize the student model Q0 = Q according to the teacher model. teacher ;

[0208] 2: Set the pruning threshold T and the number of iterations N;

[0209] 3: Initialize the experience replay pool D, converged = False;

[0210] Iterative process:

[0211] 4: Repeat steps 4 to 13 until convergence, where the current training round is i;

[0212] 5: If the threshold for model iteration pruning is not reached, repeat steps 6 and 7:

[0213] 6: Training the teacher model Q teacher Store the state-action pair in the experience replay pool D;

[0214] 7: Sample model Q from D i The model was fine-tuned using policy distillation with a KL divergence loss function;

[0215] 8: For the cropped Q i The sparsity of each layer of the model is measured and denoted as M;

[0216] 9: Create a small-scale dense neural network model Q according to M. i+1 ;

[0217] 10: Iterate steps 11 to 12 N times, using the teacher's experience to train the newly generated model Q. i+1 :

[0218] 11: Training Q teacher Store the state-action pair in the experience replay pool D;

[0219] 12: Sample model Q from D i+1Perform training;

[0220] 13: Q i = Q i+1 ;

[0221] 14: When the size of the pruned model meets the set threshold, i.e., size(Q i ) - size(Q i+1 ) < T, set converged = True, Q ans = Q i+1 , and the iteration ends.

[0222] According to the task scheduling method of the embodiments of the present application, intelligent decision-making can be relied on the target agent, and the first target processor for processing the first task and the first target resource amount allocated by the terminal device for the first task can be determined by comprehensively considering the load information of the processor and the network connection quality information between the terminal device and the processor. Furthermore, task scheduling can be performed according to the decision result. It can achieve efficient and reliable task scheduling and improve the system efficiency.

[0223] Figure 3 It is a schematic structural diagram of a task scheduling device provided by an embodiment of the present application;

[0224] Refer to Figure 3 , and introduce a task scheduling device provided by an embodiment of the present application, which is applied to a terminal device. The task scheduling device includes:

[0225] An acquisition module 301, configured to acquire the first task, the task queue where the first task is located, the first candidate processor, and the network connection quality information between the terminal device and other candidate processors. The first candidate processor includes the terminal device, at least one edge server, and a cloud server, and the other candidate processors are the processors in the first candidate processor except the terminal device;

[0226] A decision-making module 302, configured to input the first task, the task queue where the first task is located, the first candidate processor, and the network connection quality information between the terminal device and other candidate processors into the target agent to obtain a first decision result. The first decision result includes the first target processor for processing the first task and the first target resource amount allocated by the terminal device for the first task. The target agent is an agent obtained by training an initial agent by a cloud server according to training samples. The training samples include multiple second tasks, the task queues where the second tasks are located, the second candidate processors, the network connection quality information between the terminal device and other candidate processors, and second decision results. The other candidate processors are the processors in the second candidate processor except the terminal device;

[0227] An allocation module 303, configured to allocate the first target resource amount for the first task;

[0228] The scheduling module 304 is used to schedule the first task to the first target processor for processing using the first target resource quantity.

[0229] In some embodiments of this application, the task scheduling device further includes:

[0230] The receiving module is used to receive the target agent sent by the cloud server. The target agent is the agent obtained by the cloud server through training the trained agent according to the set pruning threshold and number of iterations.

[0231] In some embodiments of this application, the decision module 302 includes:

[0232] The computing module is used to calculate at least one performance index corresponding to scheduling the first task to the terminal device, at least one edge server, and the cloud server for processing, based on the load data of the first task, the first candidate processor, and the network connection quality data between the terminal device and other candidate processors.

[0233] The determination module is used to determine a first decision result based on at least one performance indicator and at least one performance indicator of other tasks in the task queue corresponding to the first task.

[0234] In some embodiments of this application, the at least one performance metric includes the latency required to process the first task and the amount of resources required by the terminal device, and the calculation module is specifically used for:

[0235] Based on the first task and the load data of the terminal device, calculate the first latency required to schedule the first task to the terminal device for processing and the first amount of resources required by the terminal device.

[0236] Based on the first task, the load data of at least one edge server, and the network connection quality data between the terminal device and at least one edge server, calculate the second latency required to schedule the first task to at least one edge server for processing and the second amount of resources required by the terminal device.

[0237] Based on the first task, the load data of the cloud server, and the network connection quality data between the terminal device and the cloud server, calculate the third latency required to schedule the first task to the cloud server for processing and the amount of third resources required by the terminal device.

[0238] Determine the unit, specifically for:

[0239] The first decision result is determined based on the first delay, the first resource quantity, the second delay, the second resource quantity, the third delay, the third resource quantity, and the delay and resource quantity required by other tasks in the task queue corresponding to the first task.

[0240] In some embodiments of this application, the determining unit is specifically used for:

[0241] Based on the first latency, the first resource quantity, the latency and resource quantity required by other tasks in the task queue, the first weight value corresponding to the latency, and the second weight value corresponding to the resource quantity, calculate the first reward value corresponding to scheduling the first task to the terminal device for processing.

[0242] Based on the second latency, the second resource quantity, the latency and resource quantity required by other tasks in the task queue, the first weight value corresponding to the latency, and the second weight value corresponding to the resource quantity, calculate the second reward value corresponding to scheduling the first task to at least one edge server for processing.

[0243] Based on the third latency, the third resource quantity, the latency and resource quantity required by other tasks in the task queue, the first weight value corresponding to the latency, and the second weight value corresponding to the resource quantity, calculate the third reward value corresponding to scheduling the first task to the cloud server for processing.

[0244] Based on the first reward value, the second reward value, and the third reward value, determine the first decision result corresponding to the first task that maximizes the long-term average reward function.

[0245] The task scheduling apparatus according to the embodiments of this application can rely on the target intelligent agent to make intelligent decisions, comprehensively consider the processor's load information and the network connection quality information between the terminal device and the processor to determine the first target processor for processing the first task and the first target resource amount allocated by the terminal device for the first task, and then perform task scheduling based on the decision results. It can achieve efficient and reliable task scheduling and improve system performance.

[0246] Figure 4 This is a schematic diagram of another task scheduling device provided in an embodiment of this application;

[0247] refer to Figure 4 This application provides a task scheduling device for use on a cloud server. The device includes:

[0248] The acquisition module 401 is used to acquire training samples. The training samples include multiple second tasks, the task queue where the second tasks are located, the load data of the second candidate processors, the network connection quality information between the terminal device and other candidate processors, and the second decision result. The second candidate processors include the terminal device, at least one edge server, and a cloud server. The second decision result includes the second target processor for processing the second task and the second target resource amount allocated by the terminal device for the second task. The other candidate processors are processors other than the terminal device among the second candidate processors.

[0249] The first training module 402 is used to train the initial agent based on the training samples to obtain the trained agent;

[0250] The determination module 403 is used to determine the trained agent as the target agent;

[0251] The sending module 404 is used to send the target intelligent agent to the terminal device;

[0252] The receiving module 405 is used to receive the first task sent by the terminal device when the terminal device determines, using the target intelligent agent, that the first target processor for processing the first task is a cloud server.

[0253] Processing module 406 is used to process the first task and obtain the processing result;

[0254] The sending module 404 is also used to send the processing result to the terminal device.

[0255] In some embodiments of this application, the apparatus further includes:

[0256] The settings module is used to set the clipping threshold and the number of iterations;

[0257] The second training module is used to perform pruning training on the trained agent based on the pruning threshold and the number of iterations, so as to obtain the pruned trained agent.

[0258] Module 403 is specifically used for:

[0259] Use the trained agent as the target agent.

[0260] In some embodiments of this application, the second training module is specifically used for:

[0261] Step A: Use the trained teacher model as the first student model and prune it. If the pruning threshold for the model iteration is not reached, repeat steps B and C below. The trained teacher model becomes the trained agent.

[0262] Step B: Continue training the teacher model and store the state-action pairs in the experience replay pool. The state-action pairs include the task, the queue in which the task is located, the load data of the candidate processor, the network connection quality data between the terminal device and the candidate processor, and the decision result.

[0263] Step C: Sample from the experience replay pool and adjust the first student model using strategy distillation with the divergence loss function;

[0264] Step D: Measure the sparsity of each layer of the clipped first student model to obtain the model sparsity;

[0265] Step E: Create a second student model based on model sparsity;

[0266] Step F: Iterate steps G through H N times to train the second student model using the teacher's experience;

[0267] Step G: Continue training the teacher model and store the state-action pairs in the experience replay pool;

[0268] Step H: Sample and train the second student model from the experience replay pool;

[0269] Step I: Calculate the size difference between the first student model and the second student model. If the size difference is less than the pruning threshold, use the second student model as the target agent. If the size difference is greater than or equal to the pruning threshold, repeat Step A to Step I.

[0270] The task scheduling apparatus according to the embodiments of this application can rely on the target intelligent agent to make intelligent decisions, comprehensively consider the processor's load information and the network connection quality information between the terminal device and the processor to determine the first target processor for processing the first task and the first target resource amount allocated by the terminal device for the first task, and then perform task scheduling based on the decision results. It can achieve efficient and reliable task scheduling and improve system performance.

[0271] Figure 5 This is a schematic diagram of the structure of the task scheduling device provided in the embodiments of this application;

[0272] The task scheduling device may include a processor 501 and a memory 502 storing computer program instructions.

[0273] Specifically, the processor 501 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.

[0274] Memory 502 may include mass storage for data or instructions. For example, and not limitingly, memory 502 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 502 may include removable or non-removable (or fixed) media. Where appropriate, memory 502 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 502 is non-volatile solid-state memory.

[0275] Memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Therefore, typically, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the methods according to one aspect of this disclosure.

[0276] The processor 501 implements the task scheduling method described in the above embodiments by reading and executing computer program instructions stored in the memory 502.

[0277] In one example, the task scheduling device may further include a communication interface 503 and a bus 510. Wherein, as... Figure 5 As shown, the processor 501, memory 502, and communication interface 503 are connected through bus 510 and complete communication with each other.

[0278] The communication interface 503 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.

[0279] Bus 510 includes hardware, software, or both, that couples components of a device for determining base station configuration parameters together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 510 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, this application contemplates any suitable bus or interconnect.

[0280] The task scheduling device executes the task scheduling method in the embodiments of this application, thereby achieving... Figure 1 , Figure 2 The task scheduling method.

[0281] Furthermore, in conjunction with the task scheduling methods in the above embodiments, this application embodiment can provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the task scheduling methods in the above embodiments.

[0282] This application also provides a computer program product, wherein when the instructions in the computer program product are executed by the processor of an electronic device, the electronic device performs the task scheduling method of any of the above embodiments.

[0283] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.

[0284] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.

[0285] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0286] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.

[0287] The above description is merely a specific implementation of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.

Claims

1. A task scheduling method, characterized in that, Applied to a terminal device, the method includes: The system acquires a first task, the task queue in which the first task is located, the load data of a first candidate processor, and the network connection quality information between the terminal device and other candidate processors. The first candidate processor includes the terminal device, at least one edge server, and a cloud server. The other candidate processors are processors other than the terminal device among the first candidate processors. The first task, the task queue in which the first task is located, the load data of the first candidate processor, and the network connection quality information between the terminal device and other candidate processors are input into the target agent to obtain a first decision result. The first decision result includes the first target processor that processes the first task and the first target resource amount allocated by the terminal device for the first task. The target agent is the agent obtained by the cloud server after training the initial agent based on training samples. The training samples include multiple second tasks, the task queue in which the second tasks are located, the load data of the second candidate processors, the network connection quality information between the terminal device and other candidate processors, and the second decision result. The other candidate processors are the processors other than the terminal device among the second candidate processors. Allocate the first target resource amount to the first task; The first task is scheduled to the first target processor for processing using the first target resource quantity.

2. The method according to claim 1, characterized in that, Before inputting the first task, the task queue in which the first task is located, the first candidate processor, and the network connection quality information between the terminal device and other candidate processors into the target agent to obtain the first decision result, the method further includes: The system receives the target agent sent by the cloud server. The target agent is an agent obtained by the cloud server through pruning and training the trained agent according to the set pruning threshold and number of iterations.

3. The method according to claim 1, characterized in that, The step of inputting the first task, the task queue in which the first task is located, the load data of the first candidate processor, and the network connection quality information between the terminal device and other candidate processors into the target agent to obtain the first decision result includes: Based on the first task, the load data of the first candidate processor, and the network connection quality data between the terminal device and other candidate processors, calculate at least one performance index corresponding to scheduling the first task to the terminal device, the at least one edge server, and the cloud server for processing. The first decision result is determined based on the at least one performance metric and at least one performance metric of other tasks in the task queue corresponding to the first task.

4. The method according to claim 3, characterized in that, The at least one performance metric includes the latency required to process the first task and the amount of resources required by the terminal device. The step of calculating at least one performance metric corresponding to scheduling the first task to the terminal device, the at least one edge server, and the cloud server for processing, based on the load data of the first task, the first candidate processor, and the network connection quality data between the terminal device and other candidate processors, includes: Based on the first task and the load data of the terminal device, calculate the first latency required to schedule the first task to the terminal device for processing and the first amount of resources required by the terminal device. Based on the first task, the load data of the at least one edge server, and the network connection quality data between the terminal device and the at least one edge server, calculate the second latency required to schedule the first task to the at least one edge server for processing and the second amount of resources required by the terminal device. Based on the first task, the load data of the cloud server, and the network connection quality data between the terminal device and the cloud server, calculate the third latency required to schedule the first task to the cloud server for processing and the third amount of resources required by the terminal device. Determining the first decision result based on the at least one performance indicator and at least one performance indicator of other tasks in the task queue corresponding to the first task includes: The first decision result is determined based on the first delay, the first resource quantity, the second delay, the second resource quantity, the third delay, the third resource quantity, and the delay and resource quantity required by other tasks in the task queue corresponding to the first task.

5. The method according to claim 4, characterized in that, The step of determining the first decision result based on the first delay, the first resource quantity, the second delay, the second resource quantity, the third delay, the third resource quantity, and the delay and resource quantity required by other tasks in the task queue corresponding to the first task includes: Based on the first latency, the first resource quantity, the latency and resource quantity required by each of the other tasks in the task queue, the first weight value corresponding to the latency, and the second weight value corresponding to the resource quantity, calculate the first reward value corresponding to scheduling the first task to the terminal device for processing. Based on the second latency, the second resource quantity, the latency and resource quantity required by each of the other tasks in the task queue, the first weight value corresponding to the latency, and the second weight value corresponding to the resource quantity, calculate the second reward value corresponding to scheduling the first task to the at least one edge server for processing. Based on the third latency, the third resource quantity, the latency and resource quantity required by other tasks in the task queue, the first weight value corresponding to the latency, and the second weight value corresponding to the resource quantity, calculate the third reward value corresponding to scheduling the first task to the cloud server for processing. Based on the first reward value, the second reward value, and the third reward value, determine the first decision result corresponding to the first task that maximizes the long-term average reward function.

6. A task scheduling method, characterized in that, Applied to a cloud server, the method includes: Acquire training samples, which include multiple second tasks, the task queue where the second tasks are located, the load data of the second candidate processors, the network connection quality information between the terminal device and other candidate processors, and the second decision result. The second candidate processors include the terminal device, at least one edge server, and the cloud server. The other candidate processors are processors other than the terminal device among the second candidate processors. The second decision result includes the second target processor for processing the second task and the second target resource amount allocated by the terminal device for the second task. The trained agent is obtained by training the initial agent based on the training samples. The trained agent is identified as the target agent; Send the target intelligent agent to the terminal device; When the terminal device determines, using the target intelligent agent, that the first target processor for processing the first task is the cloud server, it receives the first task sent by the terminal device. Process the first task and obtain the processing result; The processing result is sent to the terminal device.

7. The method according to claim 6, characterized in that, Before using the trained agent as the target agent, the method further includes: Set the trimming threshold and number of iterations; Based on the pruning threshold and the number of iterations, the trained agent is subjected to pruning training to obtain the pruned trained agent; The step of using the trained agent as the target agent includes: The trained agent is used as the target agent.

8. The method according to claim 7, characterized in that, The step of performing pruning training on the trained agent based on the pruning threshold and the number of iterations to obtain the pruned agent includes: Step A: Use the trained teacher model as the first student model and prune it. When the pruning threshold for model iteration does not reach the pruning threshold, repeat steps B and C below. The trained teacher model is the trained agent. Step B: Continue training the teacher model and store the state-action pairs in the experience replay pool. The state-action pairs include the task, the queue in which the task is located, the load data of the candidate processor, the network connection quality data between the terminal device and the candidate processor, and the decision result. Step C: Sample from the experience replay pool and adjust the first student model using strategy distillation with divergence loss function; Step D: Measure the sparsity of each layer of the cropped first student model to obtain the model sparsity; Step E: Create a second student model based on the sparsity of the aforementioned model; Step F: Iterate steps G to H N times, using the teacher's experience to train the second student model; Step G: Continue training the teacher model and store the state-action pairs in the experience replay pool; Step H: Sample and train the second student model from the experience replay pool; Step I: Calculate the size difference between the first student model and the second student model. If the size difference is less than the pruning threshold, use the second student model as the target agent. If the size difference is greater than or equal to the pruning threshold, repeat steps A to I.

9. A task scheduling device, characterized in that, Applied to a terminal device, the device includes: The acquisition module is used to acquire a first task, the task queue in which the first task is located, a first candidate processor, and network connection quality information between the terminal device and other candidate processors. The first candidate processor includes the terminal device, at least one edge server, and a cloud server. The other candidate processors are processors other than the terminal device among the first candidate processors. The decision module is used to input the first task, the task queue in which the first task is located, the first candidate processor, and the network connection quality information between the terminal device and other candidate processors into the target agent to obtain a first decision result. The first decision result includes the first target processor that processes the first task and the first target resource amount allocated by the terminal device for the first task. The target agent is an agent obtained by the cloud server after training an initial agent based on training samples. The training samples include multiple second tasks, the task queue in which the second tasks are located, second candidate processors, and the network connection quality information between the terminal device and other candidate processors, as well as the second decision result. The other candidate processors are processors other than the terminal device among the second candidate processors. The allocation module is used to allocate the first target resource amount to the first task; The scheduling module is used to schedule the first task to the first target processor for processing using the first target resource quantity.

10. A task scheduling device, characterized in that, The device, applied to a cloud server, includes: The acquisition module is used to acquire training samples, which include multiple second tasks, the task queue where the second tasks are located, the load data of the second candidate processors, the network connection quality information between the terminal device and other candidate processors, and the second decision result. The second candidate processors include the terminal device, at least one edge server, and the cloud server. The second decision result includes the second target processor for processing the second task and the second target resource amount allocated by the terminal device for the second task. The other candidate processors are processors other than the terminal device among the second candidate processors. The first training module is used to train the initial agent based on the training samples to obtain the trained agent. A determination module is used to determine the trained agent as the target agent; A sending module is used to send the target intelligent agent to the terminal device; A receiving module is configured to receive the first task sent by the terminal device when the terminal device determines, using the target intelligent agent, that the first target processor for processing the first task is the cloud server. The processing module is used to process the first task and obtain the processing result; The sending module is also used to send the processing result to the terminal device.

11. A task scheduling device, characterized in that, The device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the task scheduling method as described in any one of claims 1 to 5 or the task scheduling method as described in any one of claims 6 to 8.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the task scheduling method as described in any one of claims 1 to 5 or the task scheduling method as described in any one of claims 6 to 8.

13. A computer program product, characterized in that, When the instructions in the computer program product are executed by the processor of the electronic device, the electronic device performs the task scheduling method as described in any one of claims 1 to 5 or implements the task scheduling method as described in any one of claims 6 to 8.