Task scheduling method and electronic device
Patent Information
- Application Number
- CN202610961810.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-22
AI Technical Summary
[0004]本申请实施例的目的是提供一种任务调度方法和电子设备,能够避免在固定云边分层处理的方式下由于节点的网络质量波动而导致的无法及时响应的问题,大大提高任务处理的可用性的问题
[0010]基于上述技术方案,本申请提供的任务调度方法可以获取端节点、边缘节点以及云节点之间的多条链路的网络质量信息,并确定待执行任务的执行复杂度。然后基于多条链路的网络质量信息以及待执行任务的执行复杂度,可以确定端节点、边缘节点以及云节点中待执行任务对应的第一执行节点。也就是说,第一执行节点是结合了节点之间链路的网络状态以及任务的执行复杂度来动态确定的。如此向第一执行节点发送待执行任务以及第一指示信息,可以指示在当前网络质量下,具有处理待执行任务能力的第一执行节点来处理该待执行任务,可以避免在固定云边分层处理的方式下由于节点的网络质量波动而导致的无法及时响应的问题,大大提高任务处理的可用性。
Smart Images

Figure CN122802506A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer technology, specifically relating to a task scheduling method and an electronic device. Background Technology
[0002] Currently, service robots face a contradiction when performing reasoning tasks such as facial recognition, speech understanding, and scene perception: insufficient computing power on the edge and excessive latency on the cloud. How to dynamically allocate computing resources is an urgent problem to be solved.
[0003] Current technologies employ a fixed cloud-edge hierarchical processing approach to allocate computing resources, meaning simple tasks are assigned to edge processing and complex tasks to cloud processing. However, this hierarchical strategy is static and fixed, and in scenarios with fluctuating network quality, it cannot respond in a timely manner, easily leading to task backlog or response timeouts. Summary of the Invention
[0004] The purpose of this application is to provide a task scheduling method and electronic device that can avoid the problem of untimely response caused by network quality fluctuations of nodes in a fixed cloud-edge layered processing method, and greatly improve the availability of task processing.
[0005] To solve the above-mentioned technical problems, this application is implemented as follows: In a first aspect, embodiments of this application provide a task scheduling method, the method comprising: acquiring network quality information of multiple links and determining the execution complexity of a task to be executed, wherein the links include links between end nodes and edge nodes, and links between end nodes and cloud nodes; determining a first execution node corresponding to the task to be executed based on the network quality information of the multiple links and the execution complexity of the task to be executed, wherein the first execution node is one of the end node, the edge node, and the cloud node; and sending the task to be executed and first indication information to the first execution node, wherein the first indication information is used to instruct the first execution node to process the task to be executed.
[0006] Secondly, embodiments of this application provide a task scheduling device, which includes a communication unit and a processing unit. The communication unit is used to acquire network quality information of multiple links, wherein the links include links between end nodes and edge nodes, and links between end nodes and cloud nodes. The processing unit is used to determine the execution complexity of a task to be executed. The processing unit is also used to determine a first execution node corresponding to the task to be executed based on the network quality information of the multiple links and the execution complexity of the task to be executed, wherein the first execution node is one of the end node, the edge node, and the cloud node. The communication unit is also used to send the task to be executed and first indication information to the first execution node, wherein the first indication information is used to instruct the first execution node to process the task to be executed.
[0007] Thirdly, embodiments of this application provide an electronic device including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the method described in the first aspect.
[0008] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0009] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.
[0010] Based on the above technical solution, the task scheduling method provided in this application can obtain network quality information of multiple links between end nodes, edge nodes, and cloud nodes, and determine the execution complexity of the task to be executed. Then, based on the network quality information of multiple links and the execution complexity of the task to be executed, the first execution node corresponding to the task to be executed in the end node, edge node, and cloud node can be determined. That is to say, the first execution node is dynamically determined by combining the network status of the links between nodes and the execution complexity of the task. By sending the task to be executed and the first indication information to the first execution node, it can instruct the first execution node with the ability to process the task to be executed under the current network quality to process the task. This can avoid the problem of untimely response caused by the fluctuation of the network quality of nodes in the fixed cloud-edge layered processing method, and greatly improve the availability of task processing. Attached Figure Description
[0011] Figure 1This application provides a schematic diagram of the architecture of a task scheduling system. Figure 2 A flowchart illustrating a task scheduling method provided in an embodiment of this application; Figure 3 A flowchart illustrating another task scheduling method provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of a task scheduling device provided in an embodiment of this application; Figure 5 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0013] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0014] It should be noted that in the embodiments of this application, the words "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the words "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0015] In the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0016] Currently, service robots face a contradiction when performing inference tasks such as facial recognition, speech understanding, and scene perception: insufficient edge computing power and excessively high cloud latency. For example, when all tasks are assigned to the cloud for processing, if the cloud is in a congested or weak network environment, the inference latency will be extremely high (e.g., exceeding 500 milliseconds), severely impacting the real-time performance of interactions. Furthermore, when all tasks are assigned to the edge for processing, the limited computing power of the chip prevents inference for complex tasks. Therefore, how to dynamically allocate computing resources to improve the inference capabilities of service robots is a pressing issue that needs to be addressed.
[0017] Current technologies employ a fixed cloud-edge hierarchical processing approach to allocate computing resources, meaning simple tasks are assigned to edge processing and complex tasks to cloud processing. However, this hierarchical strategy is static and fixed, making it unable to respond promptly in scenarios with fluctuating network quality, easily leading to task backlog or response timeouts. For example, if the network signal of a cloud node drops from -60 dBm to -80 dBm, the cloud node will lack the ability to process tasks, and the task cannot be switched to an edge node in time, resulting in a sudden increase in response latency and a sharp decline in user experience.
[0018] Furthermore, the existing technical solutions rely solely on static task classification and fixed computing power allocation, lacking the ability to dynamically perceive the real-time complexity of tasks. This prevents the identification of complexity differences for the same type of task in different application scenarios, such as the difference in complexity between face recognition in unobstructed, simple backgrounds and face recognition in complex, occluded scenarios. Consequently, this results in a mismatch between computing resource allocation and actual task computing power requirements, insufficient accuracy in computing power scheduling, and overall low resource utilization efficiency.
[0019] Most related collaborative scheduling schemes are limited to supporting two-level collaboration between the end and the cloud or between the edge and the cloud, failing to effectively utilize the mid-level computing resources of edge nodes; under adverse network conditions such as network congestion and signal fluctuations, there are no degradation and alternative scheduling mechanisms, and the task adaptation and fault tolerance capabilities are insufficient.
[0020] Furthermore, the relevant scheduling decisions are usually made on cloud nodes, and scheduling decisions require network round-trip time. Thus, under adverse network conditions such as network congestion and signal fluctuations, the delay of scheduling decisions may exceed the time of task execution, forming a scheduling paradox.
[0021] Therefore, the task scheduling method provided in this application can obtain network quality information of multiple links between end nodes, edge nodes, and cloud nodes, and determine the execution complexity of the task to be executed. Then, based on the network quality information of multiple links and the execution complexity of the task to be executed, the first execution node corresponding to the task to be executed in the end node, edge node, and cloud node can be determined. That is, the first execution node is dynamically determined by combining the network status of the links between nodes and the execution complexity of the task. By sending the task to be executed and the first indication information to the first execution node, it can instruct the first execution node with the ability to process the task to be executed under the current network quality to process the task. This can avoid the problem of untimely response caused by the fluctuation of the network quality of nodes in the fixed cloud-edge layered processing method, and greatly improve the availability of task processing.
[0022] The task scheduling method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.
[0023] Figure 1 This is a schematic diagram of the architecture of a task scheduling system provided in an embodiment of this application. Figure 1 As shown, the task scheduling system includes end node 101, edge node 102 and cloud node 103.
[0024] In this configuration, end node 101 and edge node 102 are connected via a communication link, and edge node 102 and cloud node 103 are connected via a communication link. This communication link can be a wired communication link or a wireless communication link; this application does not limit the type of link.
[0025] It should be noted that, Figure 1 The task scheduling system is only described by example. In actual application scenarios, the number of end nodes 101, edge nodes 102 and cloud nodes 103 may be more or less, and this application does not impose any restrictions on this.
[0026] In some embodiments, the end node 101 is configured with a scheduling system 1011, which is used to obtain network quality information of multiple links and determine the execution complexity of the task to be executed; based on the network quality information of multiple links and the execution complexity of the task to be executed, it determines the first execution node corresponding to the task to be executed, and sends the task to be executed and first indication information to the first execution node to instruct the first execution node to process the task to be executed.
[0027] It is understood that the scheduling system provided in this application embodiment is set at the end node. That is to say, in this application embodiment, the acquisition of network quality information, task execution complexity assessment and decision on the execution node of the task to be executed are all performed by the local scheduling system of the end node. In other words, in this application embodiment, the end node dominates the scheduling, which can ensure that the task scheduling decision can be completed without relying on network connectivity.
[0028] Optionally, in this embodiment, the end node 101 may be deployed with a lightweight model (e.g., MobileNet / TinyBERT) to handle low-complexity tasks with a latency less than a first latency threshold, such as 5 milliseconds.
[0029] In some embodiments, end node 101 can be a service robot, such as a computing motherboard, sensor, mobile device, IoT device, smart home device, etc. End node 101 is the source of data generation and action execution in edge computing, and also the foundation for real-time processing and response.
[0030] Understandably, end node 101 is real-time, capable of collecting and processing data in real time and responding quickly. For example, in smart homes, smart speakers, smart light bulbs, and other devices can be considered end nodes 101; they can receive user commands and react accordingly.
[0031] Optionally, the end node 101 in this embodiment can also be a user-side entity used to receive signals, or transmit signals, or both receive and transmit signals. End node 101 is used to provide users with one or more of voice services and data connectivity services. End node 101 can also be referred to as user equipment (UE), terminal equipment, terminal, access terminal, user unit, user station, mobile station, remote station, remote terminal, mobile device, user terminal, wireless communication equipment, user agent, or user device. End node 101 can be a vehicle-to-everything (V2X) device, such as a smart car, intelligent car, digital car, unmanned car, driverless car, pilotless car, or automobile, self-driving car, or autonomous car, pure electric vehicle (EV), hybrid electric vehicle (HEV), range-extended electric vehicle (REEV), plug-in hybrid electric vehicle (PHEV), or new energy vehicle. End node 101 can also be a device-to-device (D2D) device, such as an electricity meter or water meter.
[0032] End node 101 can also be a mobile station (MS), subscriber unit, drone, Internet of Things (IoT) device, station (ST) in a WLAN, cellular phone, smartphone, cordless phone, wireless data card, tablet computer, session initiation protocol (SIP) phone, wireless local loop (WLL) station, personal digital assistant (PDA) device, laptop computer, machine type communication (MTC) terminal, handheld device with wireless communication capabilities, computing device or other processing device connected to a wireless modem, vehicle-mounted device, or wearable device (also known as wearable smart device). End node 101 can also be a terminal device in next-generation communication systems, such as a terminal device in a 5G system or a terminal device in a future evolved PLMN, or a terminal device in an NR system.
[0033] In some embodiments, edge node 102 can be a multi-access edge computing server or edge gateway deployed within a local area network to provide high-bandwidth, low-latency near-field computing power. Edge node 102 can be referred to as an edge device, which is a device located at the edge of the network capable of performing computation and data processing, and also capable of data interaction with end node 101 and cloud node 103. Edge node 102 can be understood as the network node where the edge device is located, a neuron in the entire network.
[0034] Optionally, in this embodiment, the edge node 102 may be deployed with a medium-sized model (e.g., ResNet50 / BERT-base) to process medium-complexity tasks with a latency less than a second latency threshold, such as 50 milliseconds.
[0035] Optionally, the edge node 102 can be a device located on the access network side of the communication system and having wireless transceiver capabilities, or a chip or chip system that can be installed on the device. Edge node 102 includes, but is not limited to: access points (APs) in WiFi systems, such as wireless access network elements, home gateways, routers, servers, switches, bridges, etc.; evolved NodeBs (eNBs), radio network controllers (RNCs), NodeBs (NBs), base station controllers (BSCs), base transceiver stations (BTSs), home base stations (e.g., home evolved NodeBs, or home NodeBs (HNBs)); base band units (BBUs); wireless relay nodes; wireless backhaul nodes; transmission and reception points (TRPs) or transmission points (TPs); 5G base stations, such as gNBs in new radio (NR) systems, or transmission points (TRPs or TPs); one or a group of antenna panels (including multiple antenna panels) of a base station in a 5G system; or network nodes constituting gNBs or transmission points, such as base band units (BBUs) or distributed units (TRPs). Edge node 102 also includes base stations in different networking modes, such as master-evolved NodeB (MeNB) and secondary base stations (secondary eNB, SeNB, or secondary gNB, SgNB). Edge node 102 also includes different types, such as terrestrial base stations, airborne base stations, and satellite base stations.
[0036] In some embodiments, cloud node 103 can be a computer cluster, a remote data center, a server, or a public cloud server, used to provide computing and storage resources, but it will be affected by wide area network latency.
[0037] Optionally, in this embodiment, the edge node 102 may be deployed with a large model (e.g., ViT / GPT) for processing highly complex tasks, with a latency less than a third latency threshold, such as 200 milliseconds.
[0038] Optionally, cloud node 103 can be an independent physical server, a server cluster composed of multiple physical servers, a distributed file system, or at least one of the following cloud servers that provide basic cloud computing services: cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Of course, the above is merely an exemplary description of cloud node 103. Cloud node 103 can also be a network element with data processing capabilities (including but not limited to core network elements, wireless access network elements, or transmission network elements), or a data processing device, processor, or computing device configured in a network element. This application embodiment does not limit this to any particular type.
[0039] Optionally, the number of cloud nodes 103 can be more or less, and this embodiment does not limit this. Of course, cloud nodes 103 can also include other functional servers to provide more comprehensive and diversified services.
[0040] It should be noted that, in this embodiment of the application, the task to be executed can be migrated between the end node 101, the edge node 102, and the cloud node 103. For example, when the end node 101 is processing the task to be executed, if a sudden increase in the load on the end node 101 is detected, the system can trigger a task migration mechanism, such as migrating the task to be executed to the edge node 102 to continue the subsequent inference, thereby achieving seamless handover and avoiding the delay caused by the recalculation of the task to be processed.
[0041] It should be noted that the various embodiments of this application can be referenced or learned from each other. For example, the same or similar steps, method embodiments, system embodiments and device embodiments can be referenced from each other without limitation.
[0042] Figure 2 This is a flowchart illustrating a task scheduling method provided in an embodiment of this application. Figure 2 As shown, the method includes the following steps S201-S203.
[0043] S201. The scheduling system obtains network quality information of multiple links and determines the execution complexity of the task to be executed.
[0044] The links include those between end nodes and edge nodes, as well as those between end nodes and cloud nodes.
[0045] Optionally, network quality information includes, but is not limited to, at least one of the following: bandwidth (in megabits per second (Mbps)), round trip time (RTT) (in milliseconds (ms)), number of packet losses and packet loss rate during the probe period, and network latency jitter.
[0046] In one possible implementation, the scheduling system in the end node can send a first probe packet (e.g., lightweight probe packets) to the edge node. This first probe packet is used to probe the network quality information of the link between the end node and the edge node. Correspondingly, the edge node receives the first probe packet and immediately returns a first response packet, carrying its own reception timestamp. The end node, upon receiving the first response packet and combining it with the timestamps of its locally sent and received first probe packets, can determine the network quality information of the link.
[0047] Similarly, the scheduling system in the end node can send a second probe message to the cloud node. This second probe message is used to probe the network quality information of the link between the end node and the cloud node. Correspondingly, the edge node receives the second probe message and immediately returns a second response message, carrying its own reception timestamp. The end node, by receiving the second response message and combining it with the timestamps of its locally sent and received second probe messages, can determine the network quality information of the link.
[0048] Optionally, end nodes can send probe messages to edge nodes or cloud nodes at preset intervals to update network quality information and provide real-time basis for task scheduling decisions. For example, an end node can send the first probe message to an edge node every 100 milliseconds, thus updating the link's network quality information every 100 milliseconds.
[0049] Optionally, in one possible implementation, the network quality information of the link can be smoothed by exponentially weighted moving average to filter out instantaneous network jitter and improve the reliability of the output network quality information.
[0050] In some embodiments, determining the execution complexity of a task to be executed includes: the scheduling system can determine the data size of the task to be executed, the number of floating point operations per second (FLOPS), and the real-time load rate of the end nodes; and weight the data size of the task to be executed, the number of floating point operations per second (FLOPS), and the real-time load rate of the end nodes to obtain the execution complexity of the task to be executed.
[0051] Among them, execution complexity can be used to reflect the computational scale, processing difficulty, and computing resource consumption level of the task to be executed.
[0052] The data scale of the task to be executed can be image resolution, audio / video duration, etc. Of course, the above is only an exemplary description of the data scale. The task to be executed can be a task in an audio / video scene, or a task in a point cloud registration (high concurrency, computationally intensive) scenario in simultaneous localization and mapping (SLAM), a task in a multi-sensor fusion pose estimation scenario in complex terrain (high real-time requirements), or a task in a large language model (LLM) path reasoning scenario in semantic navigation (memory intensive). This application does not impose any restrictions on the application scenario corresponding to the task to be executed.
[0053] The number of floating-point operations per second (Floating-point operations per second) during model inference can be estimated based on the task type to be executed. It's important to note that Floating-point operations per second can be understood as a unit for evaluating computational speed, and can be used as a metric to describe hardware performance, such as evaluating the computing power of a GPU, i.e., the computational speed it can generate for the model.
[0054] In one possible implementation, the execution complexity of the task to be executed can be specifically quantified using the following formula 1: C = α×A + β×B + γ×D (Formula 1) Where C represents the execution complexity of the task to be executed, A represents the data size of the task to be executed, B represents the number of floating-point operations per second, D represents the real-time load rate of the end node (e.g., the real-time load rate of the central processing unit (CPU) / graphics processing unit (GPU) on the end node), and α, β, and γ are dynamic adaptive weights obtained by online regression fitting from historical scheduling data. The execution complexity C of the task to be executed can be a complexity level, such as "low complexity", "medium complexity", and "high complexity".
[0055] For example, when the real-time load rate of the GPU on the end node is greater than 80%, the weight of γ increases sharply. At this time, the execution complexity of the task to be executed can be rated as "high complexity". In this case, the computing power resources on the end node are not enough to handle the task to be executed. Therefore, the task to be executed is scheduled to the edge node or cloud node for processing to protect the computing power resources of the end node.
[0056] S202. The scheduling system determines the first execution node corresponding to the task to be executed based on the network quality information of multiple links and the execution complexity of the task to be executed.
[0057] The first execution node can be one of the end node, edge node, or cloud node.
[0058] In some embodiments, the scheduling system can determine the network quality score of each link in the multiple links based on the network quality information of the multiple links; input the network quality score of each link and the execution complexity of the task to be executed into the scheduling strategy model to determine the first execution node corresponding to the task to be executed.
[0059] The scheduling strategy model can be a lightweight decision model trained with a deep Q-network (DQN) of reinforcement learning. The number of parameters of this model is one-tenth that of the standard DQN model. In other words, this model occupies very little computing power in the end node, so as to ensure that the scheduling system in the end node can quickly determine the first node corresponding to the task to be executed (i.e., output scheduling action) while ensuring the computing power of the main business of the end node.
[0060] To elaborate further, DQN is a cutting-edge algorithm that combines deep learning (extracting features from complex states) with reinforcement learning decision-making (Q-Learning), and can solve the optimal scheduling decision problem in a high-dimensional continuous state space.
[0061] In one possible implementation, the three network metrics—bandwidth, round-trip latency, and packet loss rate—are each subjected to threshold normalization, mapping each metric to the 0-1 range. Based on preset weighting coefficients corresponding to bandwidth, latency, and packet loss rate, a quantitative network quality score in the 0-100 range is calculated using a linear weighted fusion method, thereby achieving a comprehensive quantitative assessment of the multi-dimensional network status.
[0062] Optionally, the network quality score of each link and the execution complexity of the task to be executed are input into the scheduling strategy model to determine the first execution node corresponding to the task to be executed. This may include: determining the expected benefit of each node in executing the task to be executed based on the network quality score of each link, the execution complexity of the task to be executed, and the load rate of each node; and determining the node corresponding to the largest expected benefit as the first execution node corresponding to the task to be executed.
[0063] It should be noted that a higher expected return indicates lower latency and better performance for that node. Therefore, selecting the node with the highest expected return as the first execution node maximizes the computing resources available for the tasks to be executed.
[0064] In one example, consider a strong network scenario with a network quality score of 75 (good), an execution complexity of 80 (high complexity) for the task to be executed, and node loads of 30% for end nodes, 40% for edge nodes, and 10% for cloud nodes. The state vector S (network quality score for each link, execution complexity of the task to be executed, and load rate of each node), i.e., S=[75, 80, 0.3, 0.4, 0.1], is input into the DQN scheduling strategy model. After inference, the model outputs the expected revenue Q values for end nodes, edge nodes, and cloud nodes. For example, the expected revenue Q1 for end nodes is 28, Q2 for edge nodes is 65, and Q3 for end nodes is 89. The DQN scheduling strategy model then compares the Q values of the three nodes and selects the cloud node with the highest Q value as the first execution node.
[0065] In another example, consider a weak network scenario with a network quality score of 25 (poor), a task execution complexity of 40 (medium complexity), and node loads of 20% for end nodes, 35% for edge nodes, and 80% for cloud nodes. The state vector S (network quality score for each link, execution complexity of the task, and load rate of each node), i.e., S=[25, 40, 0.2, 0.35, 0.8], is input into the DQN scheduling strategy model. After inference, the model outputs the expected revenue Q values for end nodes, edge nodes, and cloud nodes. For example, the expected revenue Q1 for end nodes is 72, Q2 for edge nodes is 58, and Q3 for end nodes is 21. The DQN scheduling strategy model then compares the Q values of the three nodes and selects the end node with the highest Q value as the first execution node.
[0066] Optionally, in some scenarios, the DQN scheduling decision model can schedule tasks with low complexity or those experiencing network interruptions to end nodes (i.e., the robot's local node) for execution. In other words, for tasks with low complexity or those experiencing network interruptions, the DQN scheduling strategy model can select end nodes as the first execution node. Tasks with medium complexity or those with moderate network quality can be scheduled to edge nodes for execution. Similarly, tasks with high complexity and good network quality can be scheduled to cloud nodes for execution. Of course, the above schemes are merely exemplary methods for determining the first execution node in the DQN model. The DQN scheduling decision model in this application can also employ other dynamic optimization schemes for selecting the first execution node for a task, and this application does not impose any limitations on these schemes.
[0067] It should be noted that since the DQN scheduling strategy model selects the node with the highest expected benefit as the first execution node, that is, selects the node with the highest expected benefit to process the task to be executed, even if the execution complexity of the task to be executed is low, the DQN scheduling strategy model can select the edge node as the first execution node when the end node load is extremely high and the network quality score of the edge node is excellent.
[0068] S203. The scheduling system sends the task to be executed and the first instruction information to the first execution node.
[0069] The first instruction information is used to instruct the first execution node to process the task to be executed.
[0070] In one example, taking an edge node as the first execution node and a complex scene facial recognition artificial intelligence (AI) inference task as an example, the scheduling system can encapsulate the data packet of the task to be executed, including raw facial image data, inference model parameters, task timeout constraints, and other business data. The scheduling system generates a first instruction message, which instructs the edge node to enable its local AI inference model, call the specified facial recognition model, and execute the current task according to a preset priority. The scheduling system sends the task to be executed and the first instruction message to the edge node through a communication link. After receiving the message, the edge node parses the first instruction message, loads the model according to the message, and schedules its own computing resources to complete the facial recognition task.
[0071] In another example, taking the robot's local end node as the first execution node and the task to be executed as a speech and semantic understanding task, the scheduling system can encapsulate the data packet of the task to be executed, including raw speech data, inference model parameters, task timeout constraints, and other business data. The scheduling system generates first instruction information, which instructs the end node to call the local speech parsing model, start the real-time inference thread, and prohibit the task from being processed in the cloud. The scheduling system sends the task to be executed and the first instruction information to the edge node through the communication link. After receiving the first instruction information, the edge node parses it, loads the model according to the first instruction information, and schedules its own computing resources to complete the speech and semantic understanding task.
[0072] Based on the above technical solution, the task scheduling method provided in this application can obtain network quality information of multiple links between end nodes, edge nodes, and cloud nodes, and determine the execution complexity of the task to be executed. Then, based on the network quality information of multiple links and the execution complexity of the task to be executed, the first execution node corresponding to the task to be executed in the end node, edge node, and cloud node can be determined. That is to say, the first execution node is dynamically determined by combining the network status of the links between nodes and the execution complexity of the task. By sending the task to be executed and the first indication information to the first execution node, it can instruct the first execution node with the ability to process the task to be executed under the current network quality to process the task. This can avoid the problem of untimely response caused by the fluctuation of the network quality of nodes in the fixed cloud-edge layered processing method, and greatly improve the availability of task processing.
[0073] As one possible embodiment provided in this application, combined with Figure 2 ,like Figure 3 As shown, the embodiments of this application also include an implementation scheme for migrating execution nodes during the execution of a task to be processed, specifically including the following S301-S302.
[0074] S301. The scheduling system monitors the execution status of tasks to be executed.
[0075] The execution status reflects the progress of the task to be executed on the first execution node and the network status.
[0076] In some embodiments, when the first execution node is executing the task to be executed, the scheduling system can continuously monitor the execution progress of the task to be executed and the network status of each node (including the first execution node).
[0077] For example, taking the first execution node as an edge node and the task to be executed as a complex face recognition AI inference task, the execution status includes the execution progress of the complex face recognition AI inference task and the real-time network status between the current end node and the edge node.
[0078] To monitor the execution progress of complex face recognition AI inference tasks, edge nodes can report real-time information to the scheduling system, including the initialization completion rate, the loading progress of the AI inference model, the percentage of inference computation completed (e.g., 30%, 70%, or 100%), whether there are any pauses, timeouts, or inference failures during execution. For example, the task execution progress could be 60%, indicating normal operation with no timeouts.
[0079] To monitor the real-time network status between end nodes and edge nodes, the scheduling system can continuously detect the real-time round-trip latency, packet loss rate, and available bandwidth of the links between end nodes and edge nodes. For example, the network status could be bandwidth RTT 45ms, packet loss rate 0.3%, and 15Mbps.
[0080] It should be noted that the scheduling system's continuous monitoring of the execution status of pending tasks can be used for both short-term decision-making and long-term feedback. For short-term decision-making, the scheduling system can determine whether a pending task is at a transferable breakpoint based on its execution status. For example, the inference process of a pending task can be implemented through steps 1-3, and the first execution node has just completed step 1 (e.g., the first execution node has just extracted the intermediate layer features of the pending task). If the execution progress of the pending task is greater than or equal to a first threshold (e.g., 90%), then even in the event of a network mutation, the task migration mechanism will not be triggered, meaning the pending task will not be migrated from the first execution node to the second execution node. For long-term feedback, the scheduling system can record the total latency of the pending tasks as a reward value for the online reinforcement learning of the DQN scheduling strategy model, continuously optimizing the DQN scheduling strategy model.
[0081] S302. The scheduling system determines whether to move the task to be executed from the first execution node to the second execution node based on the execution status of the task to be executed.
[0082] The second execution node is one of the end nodes, edge nodes, and cloud nodes other than the first execution node. For example, taking a scenario that includes end node 1, edge node 1, edge node 2, edge node 3, cloud node 1, and cloud node 2, if the first execution node is edge node 1, then the second execution node can be one of the end node 1, edge node 2, edge node 3, cloud node 1, and cloud node 2.
[0083] In one possible implementation, if the execution progress of the task to be executed is less than a first threshold and the network quality of the first execution node suddenly changes, the task to be executed is migrated from the first execution node to the second execution node.
[0084] For example, with a first threshold of 90%, the first execution node being an edge node, the task to be executed being a complex face recognition AI inference task, and the second execution node being a robot local node, the scheduling system assigns the complex face recognition AI inference task to the edge node for execution. When the execution progress is 45% (i.e., the core inference task of the complex face recognition AI inference task has not been completed), and the network was originally normal, a sudden fluctuation in the network signal causes the network quality score to plummet from 82 points to 28 points (a sudden change in network quality occurs, latency increases sharply, and packet loss rate increases). In this situation, the scheduling system determines that if the complex face recognition AI inference task continues to be executed on the edge node, timeouts and stuttering will occur. Therefore, the scheduling system immediately migrates the unfinished complex face recognition AI inference task from the edge node to the robot local node (i.e., the second execution node), where the robot local node continues to complete the AI inference.
[0085] Alternatively, in another possible implementation, if the execution progress of the task to be executed is less than the first threshold, and the real-time expected benefit of the second execution node executing the task to be executed is greater than the real-time expected benefit of the first execution node executing the task to be executed, the task to be executed is migrated from the first execution node to the second execution node.
[0086] For example, with a first threshold of 90%, the first execution node being an edge node, the task to be executed being a complex face recognition AI inference task, and the second execution node being a robot local node, the scheduling system allocates the complex face recognition AI inference task to the edge node for execution. When the execution progress is 45% (i.e., the core inference task of the complex face recognition AI inference task has not been completed), and the real-time expected return for the edge node (first execution node) is 35, while the real-time expected return for the robot local node (second execution node) is 78, the scheduling system determines that continuing to process the complex face recognition AI inference task by the edge node would result in low efficiency and high latency. Therefore, the scheduling system immediately migrates the unfinished complex face recognition AI inference task from the edge node to the robot local node (i.e., the second execution node), allowing the robot local node to continue completing the AI inference. This improves the processing efficiency of the complex face recognition AI inference task and also frees up the computing resources of the edge node.
[0087] Alternatively, in another possible implementation, if the execution progress of the task to be executed is greater than or equal to a first threshold, the task to be executed will not be migrated from the first execution node to the second execution node. In other words, if the execution progress of the task to be executed is greater than or equal to the first threshold, in order to avoid the migration overhead exceeding the remaining computational overhead, the task migration mechanism will not be triggered even if a network mutation occurs.
[0088] In some embodiments, the above-mentioned migration of the task to be executed from the first execution node to the second execution node may include: the scheduling system determining the execution information when the first execution node processes the task to be executed, the execution information including the intermediate layer feature map of the task to be executed and the context state information of the task to be executed; controlling the first execution node to send the execution information and the second instruction information to the second execution node, the second instruction information being used to instruct the second execution node to continue processing the task to be executed based on the execution information.
[0089] For example, taking a complex face recognition AI inference task as the task to be executed, with the first execution node being an edge node and the second execution node being a robot local end node, the intermediate layer feature map can be the intermediate feature vector extracted by the backbone network of the face recognition model, and the context state information can be the input original image parameters, the model version information, the model inference parameters, the current inference step, the task timeout constraint, and the coordinates of the face detection box, etc.
[0090] Optionally, the complex face recognition AI inference task runs on the edge node (first execution node). The shallow inference of the model has been completed, and the intermediate feature vector has been extracted by the face recognition model backbone network. Simultaneously, contextual state information (including input image parameters, model version information, model inference parameters, current inference step, task timeout constraints, and face detection box coordinates) is recorded. The scheduling system determines this intermediate feature vector and contextual state information, and controls the edge node (first execution node) to send execution information (including the aforementioned intermediate feature vector, contextual state information, and second instruction information) to the robot's local end node (second execution node). This instruction instructs the second execution node not to reload the image from scratch and run the shallow feature extraction again, but to directly utilize the received intermediate feature vector and contextual state information to continue the subsequent face recognition backend classification inference.
[0091] Understandably, the above solution supports breakpoint resume and seamless migration, eliminating the need for task re-initialization and recalculation, thus reducing bandwidth overhead and computing power consumption.
[0092] Based on the above technical solution, the scheduling system monitors the execution status of tasks to be executed and determines whether to migrate the tasks from the first execution node to the second execution node based on the execution status. In this way, the real-time computing power and network conditions of each node in the end, edge, and cloud can be dynamically matched to realize on-demand allocation of computing resources and significantly improve the overall resource utilization efficiency. It also enhances the adaptive fault tolerance capability for complex working conditions such as weak networks, network mutations, and high node loads, and improves the robustness and environmental adaptability of end-edge-cloud collaborative scheduling.
[0093] In one example, taking a scenario where the network signal fluctuates between -60dBm and -85dBm, the traditional cloud solution has a response latency of 480ms when the signal is less than a preset threshold (e.g., -85dBm). The task scheduling method provided in this application can switch the task to be processed to the edge node within a preset time (e.g., 100ms) after the network signal drops, with a response latency of less than 45ms, which is 91% lower than the latency of the fixed cloud solution.
[0094] In another example, five local robot nodes simultaneously execute face recognition and speech understanding tasks, resulting in significant concurrent load on the cloud nodes. The fixed cloud solution experiences a response latency of 380ms at peak concurrency, with some tasks timed out. The task scheduling method provided in this application can automatically allocate low-complexity tasks (e.g., simple background face recognition) to local robot nodes and high-complexity tasks (e.g., complex scene face recognition + speech understanding) to edge nodes based on the network quality of each node and the complexity of the tasks. This reduces the cloud node load by 65%, achieves a response latency of <80ms for all tasks, and eliminates timeouts, significantly improving the reliability of task processing.
[0095] In some embodiments, this application also provides an online reinforcement learning mechanism. For example, the scheduling system receives the execution results of tasks to be executed and updates the scheduling strategy model based on the execution results of the tasks to be executed.
[0096] For example, consider a face recognition AI inference task, with the execution chain migrating from cloud nodes to edge nodes for inference completion. The execution result of the face recognition AI inference task can include the final recognition result, actual execution latency, task success / failure indicators, network quality during task execution, node load, and other information. After completing the face recognition AI inference task, the edge node can send the execution result to the end node. Correspondingly, the end node receives the execution result from the edge node and trains and optimizes the scheduling strategy model based on this result. Furthermore, the scheduling system can also statistically analyze the task execution latency of each node and generate a scheduling effect report based on a preset duration (e.g., by day or by time period). Then, based on the scheduling effect report, the scheduling strategy is dynamically adjusted to prioritize the allocation of similar tasks to nodes with lower latency and better stability, achieving closed-loop continuous optimization.
[0097] Optionally, the scheduling decision model outputs a scheduling action (such as migrating a task from the first execution node to the second execution node), executes the scheduling action, and records the actual latency. Then, the reward is calculated, and the experience vector (e.g., S=[state, action, reward]) is stored in the experience replay pool. The scheduling system can then periodically extract experience vectors from the experience replay pool to update the parameters of the DQN scheduling strategy model, forming a closed-loop continuous optimization.
[0098] Optionally, the above rewards can be calculated using the following formula 2.
[0099] Reward = -(w1 × total delay + w2 × transmission power consumption) (Formula 2) Among them, w1 and w2 are preset weights.
[0100] In some embodiments, this application also provides a network interruption degradation mechanism. For example, in the event of a network interruption, the scheduling system switches the pending tasks to the end node for execution. This ensures that the pending tasks are not interrupted, guaranteeing the reliability of task processing.
[0101] It should be noted that the task scheduling method provided in this application embodiment can be executed by a task scheduling device, or a control module within the task scheduling device for executing the loading task scheduling method. This application embodiment uses the execution of the loading task scheduling method by a task scheduling device as an example to illustrate the task scheduling method provided in this application embodiment.
[0102] This application embodiment can divide the task scheduling device into functional modules or functional units according to the above method examples. For example, each function can be divided into a separate functional module or functional unit, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or in software functional modules or functional units. The module or unit division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.
[0103] like Figure 4 As shown, Figure 4 This is a schematic diagram of the structure of a task scheduling device 40 provided in an embodiment of this application. The device includes a communication unit 401 and a processing unit 402.
[0104] The communication unit 401 is used to acquire network quality information of multiple links, wherein the links include links between end nodes and edge nodes, and links between end nodes and cloud nodes; the processing unit 402 is used to determine the execution complexity of the task to be executed; the processing unit 402 is also used to determine the first execution node corresponding to the task to be executed based on the network quality information of multiple links and the execution complexity of the task to be executed, wherein the first execution node is one of the end node, edge node, and cloud node; the communication unit 401 is also used to send the task to be executed and first indication information to the first execution node, wherein the first indication information is used to instruct the first execution node to process the task to be executed.
[0105] In one possible implementation, the processing unit 402 is specifically used to: determine the data size of the task to be executed, the number of floating-point operations per second, and the real-time load rate of the end node; and perform weighted processing on the data size of the task to be executed, the number of floating-point operations per second, and the real-time load rate of the end node to obtain the execution complexity of the task to be executed, which reflects the computational scale, processing difficulty, and computing resource consumption level of the task to be executed.
[0106] In one possible implementation, the processing unit 402 is specifically used to: determine the network quality score of each link in the multiple links based on the network quality information of the multiple links; input the network quality score of each link and the execution complexity of the task to be executed into the scheduling strategy model to determine the first execution node corresponding to the task to be executed.
[0107] In one possible implementation, the processing unit 402 is specifically used to: determine the expected benefit of each node in executing the task based on the network quality score of each link, the execution complexity of the task to be executed, and the load rate of each node; and determine the node corresponding to the largest expected benefit as the first execution node corresponding to the task to be executed.
[0108] In one possible implementation, the processing unit 402 is further configured to: monitor the execution status of the task to be executed, wherein the execution status is used to reflect the execution progress of the task to be executed at the first execution node and the network status; and based on the execution status of the task to be executed, determine whether to migrate the task to be executed from the first execution node to the second execution node, wherein the second execution node is one of the end node, edge node, and cloud node other than the first execution node.
[0109] In one possible implementation, the processing unit 402 is specifically used to: migrate the task to be executed from the first execution node to the second execution node when the execution progress of the task to be executed is less than a first threshold and the network quality of the first execution node changes abruptly; or, migrate the task to be executed from the first execution node to the second execution node when the execution progress of the task to be executed is less than the first threshold and the real-time expected benefit of the second execution node executing the task to be executed is greater than the real-time expected benefit of the first execution node executing the task to be executed.
[0110] In one possible implementation, the processing unit 402 is specifically used to: determine the execution information when the first execution node processes the task to be executed, the execution information including the intermediate layer feature map of the task to be executed and the context state information of the task to be executed; control the first execution node to send the execution information and the second instruction information to the second execution node, the second instruction information being used to instruct the second execution node to continue processing the task to be executed based on the execution information.
[0111] In one possible implementation, the task scheduling device 40 may further include a storage unit 403. Figure 4 (shown in dashed box in the image), the storage unit 403 stores a program or instruction. When the processing unit 402 executes the program or instruction, the task scheduling device 40 can execute the task scheduling method described in the above method embodiment.
[0112] Through the above description of the embodiments, those skilled in the art will clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0113] The task scheduling device in this application embodiment can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can be servers, network-attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. This application embodiment does not impose specific limitations.
[0114] The task scheduling device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.
[0115] The task scheduling device provided in this application embodiment can achieve... Figures 2 to 3 The various processes implemented by the task scheduling device in the method embodiment will not be described again here to avoid repetition.
[0116] Since the task scheduling device provided in this embodiment can execute the above method, the technical effects it can achieve can be referred to the above method embodiment, and will not be repeated here.
[0117] Figure 5 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Figure 5 As shown, the electronic device includes at least one processor 501, a communication line 502, and at least one communication interface 504, and may also include a memory 503. The processor 501, memory 503, and communication interface 504 are connected via the communication line 502.
[0118] The processor 501 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application, such as one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs).
[0119] Communication line 502 may include a path for transmitting information between the aforementioned components.
[0120] The communication interface 504 is used to communicate with other devices or communication networks. It can use any transceiver-like device, such as Ethernet, radio access network (RAN), wireless local area network (WLAN), etc.
[0121] The memory 503 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of including or storing desired program code having the form of instructions or data structures and accessible by a computer, but not limited thereto.
[0122] In one possible design, the memory 503 can exist independently of the processor 501, meaning the memory 503 can be an external memory of the processor 501. In this case, the memory 503 can be connected to the processor 501 via a communication line 502 to store execution instructions or application code, and its execution is controlled by the processor 501 to implement the task scheduling method provided in this application embodiment. In another possible design, the memory 503 can also be integrated with the processor 501, meaning the memory 503 can be an internal memory of the processor 501. For example, the memory 503 can be a cache, used to temporarily store some data and instruction information.
[0123] As one possible implementation, processor 501 may include one or more CPUs, for example Figure 5 CPU0 and CPU1 in the example. As another possible implementation, the electronic device may include multiple processors, such as... Figure 5 The processors 501 and 507 are included. As another possible implementation, the electronic device may also include an output device 505 and an input device 506. The output device 505 communicates with the processor 501 and can display information in various ways. For example, the output device 505 may be a liquid crystal display (LCD), a light-emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector, etc. The input device 506 communicates with the processor 501 and can receive user input in various ways. For example, the input device 506 may be a mouse, keyboard, touchscreen device, or sensing device, etc.
[0124] Optionally, this application embodiment also provides an electronic device, including a processor 501, a memory 503, and a program or instructions stored in the memory 503 and executable on the processor 501. When the program or instructions are executed by the processor 501, they implement the various processes of the above task scheduling method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0125] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0126] It should be noted that the various embodiments of this application can be referenced or learned from each other. For example, the same or similar steps, method embodiments, system embodiments and device embodiments can be referenced from each other without limitation.
[0127] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described task scheduling method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.
[0128] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, optical disk register, hard disk, optical fiber, compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof, or any other form of computer-readable storage medium well known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and storage medium can reside in an application-specific integrated circuit (ASIC). In the embodiments of this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0129] Since the electronic devices, readable storage media, and computer program products in the embodiments of this application can be applied to the above methods, the technical effects they can achieve can also be referred to the above method embodiments. The embodiments of this application will not be repeated here.
[0130] This application provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above-described task scheduling method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0131] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0132] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0133] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0134] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A task scheduling method, characterized in that, The method includes: The network quality information of multiple links is obtained and the execution complexity of the task to be executed is determined. The links include links between end nodes and edge nodes, and links between end nodes and cloud nodes. Based on the network quality information of the multiple links and the execution complexity of the task to be executed, a first execution node corresponding to the task to be executed is determined. The first execution node is one of the end node, the edge node, and the cloud node. The task to be executed and a first instruction message are sent to the first execution node, wherein the first instruction message is used to instruct the first execution node to process the task to be executed.
2. The task scheduling method according to claim 1, characterized in that, Determining the execution complexity of the task to be executed includes: Determine the data size of the task to be executed, the number of floating-point operations per second, and the real-time load rate of the end node; The execution complexity of the task to be executed is obtained by weighting the data size, the number of floating-point operations per second, and the real-time load rate of the end node. The execution complexity is used to reflect the computational scale, processing difficulty, and computing resource consumption level of the task to be executed.
3. The task scheduling method according to claim 1, characterized in that, The step of determining the first execution node corresponding to the task to be executed based on the network quality information of the multiple links and the execution complexity of the task to be executed includes: Based on the network quality information of the multiple links, determine the network quality score of each of the multiple links; The network quality score of each link and the execution complexity of the task to be executed are input into the scheduling strategy model to determine the first execution node corresponding to the task to be executed.
4. The task scheduling method according to claim 3, characterized in that, The step of inputting the network quality score of each link and the execution complexity of the task to be executed into the scheduling strategy model to determine the first execution node corresponding to the task to be executed includes: Based on the network quality score of each link, the execution complexity of the task to be executed, and the load rate of each node, the expected benefit of each node in executing the task to be executed is determined. The node corresponding to the largest expected return is determined as the first execution node for the task to be executed.
5. The task scheduling method according to any one of claims 1-4, characterized in that, The method further includes: Monitor the execution status of the task to be executed, wherein the execution status is used to reflect the execution progress of the task to be executed at the first execution node and the network status; Based on the execution status of the task to be executed, determine whether to migrate the task to be executed from the first execution node to the second execution node, wherein the second execution node is one of the end node, the edge node, and the cloud node other than the first execution node.
6. The task scheduling method according to claim 5, characterized in that, Based on the execution status of the task to be executed, determining whether to migrate the task to be executed from the first execution node to the second execution node includes: If the execution progress of the task to be executed is less than a first threshold and the network quality of the first execution node suddenly changes, the task to be executed will be migrated from the first execution node to the second execution node. Alternatively, if the execution progress of the task to be executed is less than the first threshold, and the real-time expected benefit of the second execution node executing the task to be executed is greater than the real-time expected benefit of the first execution node executing the task to be executed, the task to be executed is migrated from the first execution node to the second execution node.
7. The task scheduling method according to claim 6, characterized in that, The step of migrating the task to be executed from the first execution node to the second execution node includes: Determine the execution information when the first execution node processes the task to be executed, the execution information including the intermediate layer feature map of the task to be executed and the context state information of the task to be executed; The first execution node is controlled to send the execution information and the second instruction information to the second execution node. The second instruction information is used to instruct the second execution node to continue processing the task to be executed based on the execution information.
8. A task scheduling device, characterized in that, The task scheduling device includes a communication unit and a processing unit; The communication unit is used to acquire network quality information of multiple links, wherein the links include links between end nodes and edge nodes, and links between end nodes and cloud nodes; The processing unit is used to determine the execution complexity of the task to be executed; The processing unit is further configured to determine a first execution node corresponding to the task to be executed based on the network quality information of the multiple links and the execution complexity of the task to be executed, wherein the first execution node is one of the end node, the edge node, and the cloud node; The communication unit is further configured to send the task to be executed and first instruction information to the first execution node, wherein the first instruction information is configured to instruct the first execution node to process the task to be executed.
9. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the task scheduling method as described in claims 1-7.
10. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the task scheduling method as described in claims 1-7.