Task scheduling system and method

By combining the load detection of the task scheduling system with the hardware acceleration device, network traffic is dynamically predicted, which solves the shortcomings of network cards in energy saving and emission reduction, achieves a balance between deep energy saving and performance, and adapts to dynamic traffic changes.

CN120750682AActive Publication Date: 2025-10-03INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511232760.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-10-03
Estimated Expiration
2045-08-29

AI Technical Summary

Technical Problem

Existing network card designs have defects in energy saving. Low power consumption leads to data transmission delays, dynamic frequency and voltage regulation response capabilities are insufficient, and data packet merging is difficult to adapt to dynamic traffic, making it impossible to achieve effective energy saving and emission reduction.

Method used

A task scheduling system is used to dynamically detect load information through the task scheduler, predict future traffic data, and schedule tasks to the hardware acceleration device for execution, avoiding frequent wake-up of the processor core and achieving deep energy saving.

Benefits of technology

It achieves maximum energy saving without the risk of delay, flexibly responds to changes in network traffic, ensures the best balance between energy efficiency and performance, and reduces the possibility of additional power consumption and inadequate response to burst traffic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120750682A_ABST
    Figure CN120750682A_ABST
Patent Text Reader

Abstract

The invention discloses a task scheduling system and method, and relates to the technical field of task scheduling processing, and the method comprises the steps: a task scheduler executes dynamic load detection, and obtains load information corresponding to a to-be-processed task; according to the basic window flow data, the current window flow data and a predetermined historical flow data attenuation coefficient, predicting future window flow data corresponding to each moment in a preset time period from the current moment to a future preset moment; when the future window flow data corresponding to each moment in the preset time period is lower than a preset flow data threshold value, tasks from the current moment to the future preset moment are scheduled to a hardware acceleration device; and the hardware acceleration device executes the task from the current moment to the future preset moment. According to the mode, the future low-load time period is accurately predicted, and the network task is unloaded from the ARM core to the special hardware acceleration device for execution in the window period, so that deep energy conservation is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of task scheduling processing technology, and in particular to a task scheduling system and method. Background Art

[0002] With the rapid growth of internet data centers, a large number of high-data-intensive, high-power network interface cards (NICs) have entered the market. These high-power NICs consume a significant amount of energy. Saving energy and reducing emissions is a major challenge. With countries around the world focusing on energy conservation, energy conservation has become a major focus in the electronics industry. Traditional NIC designs currently rely primarily on low-power idle states, dynamic frequency and voltage regulation, and packet merging.

[0003] However, low-power consumption, dynamic frequency and voltage regulation, and packet merging all have certain drawbacks. For example, the low-power consumption method requires the processor core to be frequently woken up, resulting in data transmission delays and additional power consumption; dynamic frequency and voltage regulation has the problem of insufficient responsiveness to burst traffic; the packet merging method requires setting a fixed threshold for the number of packets to be merged, but the fixed threshold is difficult to adapt to dynamic traffic. Summary of the Invention

[0004] The present application provides a task scheduling system and method to at least solve the problem in related technologies that network cards consume a lot of energy and cannot achieve energy conservation and emission reduction.

[0005] The present application provides a task scheduling system, which includes: a task scheduler and a hardware acceleration device; The task scheduler is configured to perform dynamic load detection and obtain load information corresponding to the task to be processed, wherein the load information includes basic window traffic data and current window traffic data; predict the future window traffic data corresponding to each moment in a preset time period from the current moment to a preset future moment based on the basic window traffic data, the current window traffic data, and a predetermined historical traffic data attenuation coefficient; and schedule the task from the current moment to the preset future moment to the hardware acceleration device when the future window traffic data corresponding to each moment in the preset time period is lower than a preset traffic data threshold. A hardware acceleration device is used to execute tasks from the current time to a preset time in the future.

[0006] The present application also provides a task scheduling method, which is applied to a task scheduling system. The task scheduling system includes a task scheduler and a hardware acceleration device. The method is executed by the task scheduler and includes: Perform dynamic load detection to obtain load information corresponding to the task to be processed, where the load information includes basic window traffic data and current window traffic data; Based on the basic window traffic data, the current window traffic data, and the predetermined historical traffic data attenuation coefficient, the future window traffic data corresponding to each moment in the preset time period from the current moment to the future preset moment is predicted; When the future window traffic data corresponding to each moment in the preset time period is lower than the preset traffic data threshold, the tasks from the current moment to the future preset moment will be scheduled to the hardware acceleration device, so that the hardware acceleration device will execute the tasks from the current moment to the future preset moment.

[0007] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned task scheduling method are implemented.

[0008] The present application also provides a computer program product, including a computer program, which implements the steps of the above-mentioned task scheduling method when executed by a processor.

[0009] Through this application, the task scheduler dynamically performs load detection and predicts window traffic data for each moment in the future based on the base window traffic and current window traffic data in the load information, as well as the decay coefficient of historical traffic data. If the window traffic data at all moments in the future preset time period does not exceed the preset traffic data threshold, all tasks corresponding to the time period from the current moment to the preset future moment can be directly scheduled for execution on the hardware acceleration device. This approach accurately predicts future low-load periods and intelligently offloads network tasks from the ARM core to a dedicated hardware acceleration device during this window, achieving deep energy savings. This solution predicts a continuous, stable low-traffic time window and completely offloads tasks to the hardware acceleration device during this period. During this time, the ARM core can enter a deeper and longer low-power state (even hibernation) without frequent wake-up. This significantly reduces the additional power consumption caused by state switching, achieving true deep energy savings. Furthermore, the decision-making of this application is predictive rather than reactive. Before making a scheduling decision, it predicts that traffic will be "consistently below the threshold" for the entire future period. This means that the possibility of sudden high traffic is essentially eliminated during this preset period. Therefore, there is no need to worry about not being able to respond in time when a burst of traffic hits. This application transforms the unpredictable risk of "burst traffic" into a manageable "deterministic low traffic" window through prediction. Moreover, because this application is a dynamic prediction, it does not rely on any fixed threshold as a criterion for executing actions (although there is a preset traffic threshold, it is used to compare with the dynamic prediction value, rather than as a fixed threshold for the accumulated amount of data packets). The system continuously performs dynamic calculations based on the "basic window traffic", "current window traffic" and "historical attenuation coefficient" to predict traffic at every moment in the future. This method is adaptive and intelligent, and can flexibly respond to dynamic changes in network traffic, thereby maximizing the energy-saving window and achieving the best balance between energy efficiency and performance while ensuring no delay risk. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0011] Figure 1 A schematic diagram of the structure of a task scheduling system provided in an embodiment of the present application; Figure 2 A schematic diagram of another task scheduling system structure provided in an embodiment of the present application; Figure 3A schematic diagram of the overall structure of the task scheduling system provided in an embodiment of the present application; Figure 4 A simplified schematic diagram of the topological decomposition logic of the specific implementation principle of the task scheduling system provided in the embodiment of the present application; Figure 5 A flowchart of a task scheduling method provided in an embodiment of the present application. DETAILED DESCRIPTION

[0012] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0013] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0014] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0015] With the rapid growth of internet data centers, a large number of high-data-intensive, high-power network interface cards (NICs) have entered the market. These high-power NICs consume a significant amount of energy. Saving energy and reducing emissions is a major challenge. With countries around the world focusing on energy conservation, energy conservation has become a major focus in the electronics industry. Traditional NIC designs currently rely primarily on low-power idle states, dynamic frequency and voltage regulation, and packet merging.

[0016] 1. Low power idle state The low-power state of a Peripheral Component Interconnect Express (PCIe) device requires the system driver to explicitly place the device in a low-power state, which in turn allows the PCIe link to transition to a low-power link state. The PCIe specification allows the PCIe link to enter a low-power state without the system driver. This feature is known as active state power management. Generally speaking, both the system driver hardware and the device hardware can detect idle time on the PCIe link and initiate power state transitions.

[0017] The low-power idle state has the following drawbacks: the processor core goes into sleep when there is no data transmission and needs to be woken up when data transmission occurs. This frequent wake-up will cause data transmission delays and additional power consumption.

[0018] 2. Dynamic frequency and voltage modulation Dynamic voltage regulation is a flexible and efficient power management technology that balances performance and power consumption by dynamically adjusting voltage and frequency. It provides optimal operating parameters based on different application scenarios and chip status, thereby minimizing power consumption while ensuring performance.

[0019] Although dynamic frequency and voltage regulation can adjust the frequency and voltage of the network card according to the load, its response capability to burst traffic is insufficient. 3. Data Packet Merging Packet coalescing is a complex networking technology designed to optimize network data transmission efficiency by combining multiple smaller packets into a single larger packet before transmission or reception. This approach significantly improves network performance, reduces network interface processing overhead, and lowers energy consumption, contributing to a more environmentally friendly and cost-effective network environment. The core function of packet coalescing is to aggregate smaller packets into larger ones, thereby reducing the total number of packets that need to be transmitted or processed. This process is crucial for optimizing data flow within a network, especially when large amounts of data are being transmitted. By reducing the number of packets, packet coalescing significantly reduces the computational workload and energy consumption of the network interface, thereby improving network efficiency.

[0020] Packet merging can be done by buffering a small number of packets and then processing them uniformly, but packet merging requires setting a fixed threshold, which will be difficult to adapt to dynamic traffic.

[0021] To solve the above problems, the present invention provides a task scheduling system. Figure 1 As shown, the system includes a task scheduler 101 and a hardware acceleration device 102 .

[0022] The task scheduler 101 is configured to perform dynamic load detection and obtain load information corresponding to the task to be processed, wherein the load information includes basic window traffic data and current window traffic data; predict future window traffic data corresponding to each moment in a preset time period from the current moment to a preset future moment based on the basic window traffic data, the current window traffic data, and a predetermined historical traffic data attenuation coefficient; and schedule the task from the current moment to the preset future moment to the hardware acceleration device 102 when the future window traffic data corresponding to each moment in the preset time period is lower than a preset traffic data threshold. The hardware acceleration device 102 is used to execute tasks from the current time to a preset time in the future.

[0023] Specifically, the task scheduler 101 may perform dynamic load detection periodically or in real time, thereby obtaining load information corresponding to the task to be processed. In a specific example, the load information may include at least basic window traffic data and current window traffic data.

[0024] The task scheduler 101 can predict the future window traffic data corresponding to each moment in a preset time period between the current moment and a preset future moment based on the basic window traffic data, the current window traffic data, and a predetermined historical traffic data attenuation coefficient. The historical traffic data attenuation coefficient is obtained by statistically analyzing training data or a large amount of experimental data based on historical traffic data, and will not be further described here. However, it should be noted that the historical traffic data attenuation coefficient changes according to changes in historical data.

[0025] When it is predicted that the future window traffic data corresponding to each future preset time is lower than the preset traffic data threshold, the tasks from the current time to the future preset time can be directly scheduled to the hardware acceleration device 102. In an optional example, the hardware acceleration device 102 is, for example, a hardware acceleration module engine (Hardware Offload Engine, abbreviated as HOE), without passing through the processor core (ARM core) in the network card.

[0026] The HOE executes all tasks from the current moment to a preset moment in the future. During this period, the network tasks handled by the ARM core are consolidated into dedicated hardware. The ARM core is placed in a deep sleep state during this period, reducing resource consumption.

[0027] In an optional example, the preset time period may actually include only one moment. For example, the next moment is predicted from the current moment. If the traffic at the next moment is lower than a preset traffic threshold, the tasks from the current moment to the next moment may be directly scheduled to the hardware acceleration device 102. Of course, if the current moment is the first moment, it is also necessary to consider whether the tasks at the current moment can be directly scheduled to the hardware acceleration device 102.

[0028] Of course, the preset time period can also be a time period consisting of several moments in the future, such as 10 minutes, 20 minutes, one hour, etc. It can be set according to actual conditions and is not limited here.

[0029] An embodiment of the present application provides a task scheduling system in which a task scheduler dynamically performs load detection and predicts window traffic data for each moment in the future based on the base window traffic and current window traffic data in the load information, as well as the decay coefficient of historical traffic data. If the window traffic data at all moments in a preset future time period does not exceed a preset traffic data threshold, all tasks corresponding to the time period from the current moment to the preset future moment can be directly scheduled for execution on a hardware acceleration device. This approach accurately predicts future low-load periods and intelligently offloads network tasks from the ARM core to a dedicated hardware acceleration device during this window, thereby achieving deep energy savings. This solution predicts a continuous, stable low-traffic time window and completely offloads tasks to the hardware acceleration device during this period. During this period, the ARM core can enter a deeper, longer-term low-power state (even hibernation) without having to be frequently awakened. This significantly reduces the additional power consumption caused by state switching, achieving true deep energy savings. Furthermore, the decision-making of this application is predictive rather than reactive. Before making a scheduling decision, it predicts that traffic will be "consistently below the threshold" for the entire future period. This means that the possibility of sudden high traffic is essentially eliminated during this preset time period. Therefore, there is no need to worry about not being able to respond in time when a burst of traffic hits. This application transforms the unpredictable risk of "burst traffic" into a manageable "deterministic low traffic" window through prediction. Moreover, because this application is a dynamic prediction, it does not rely on any fixed threshold as a criterion for executing actions (although there is a preset traffic threshold, it is used to compare with the dynamic prediction value, rather than as a fixed threshold for the accumulated amount of data packets). The system continuously performs dynamic calculations based on the "basic window traffic", "current window traffic" and "historical attenuation coefficient" to predict traffic at every moment in the future. This method is adaptive and intelligent, and can flexibly respond to dynamic changes in network traffic, thereby maximizing the energy-saving window and achieving the best balance between energy efficiency and performance while ensuring no delay risk.

[0030] In an optional embodiment, the task scheduler 101 is specifically configured to predict future window traffic data corresponding to each moment in a preset time period from the current moment to a future preset moment using the following expression:

[0031] in, is the future window traffic data corresponding to the i-th moment in the preset time period, As the basic window traffic data, is the attenuation coefficient of historical traffic data, L is the ratio between the current window traffic data and the basic window traffic data, The time difference between the current moment and the i-th moment is a constant. The i-th moment is any moment within the preset time period. The basic window traffic data is based on the Quality of Service (QOS) setting.

[0032] In a specific example, for example, the base window traffic data Wbase = 100 Mbps and the current window traffic data is 150 Mbps, then L = 1.5 and β = 0.2.

[0033] In an optional example, the time units of the current moment and the future moment can be minutes, hours or other units. However, in the calculation formula of this application document, because Only the constant of the time difference is taken, so No unit.

[0034] In a specific example, suppose ,So: .

[0035] In another specific example, suppose ,So: .

[0036] In another specific example, suppose ,So: .

[0037] In another specific example, suppose ,So: .

[0038] The above example is only used to illustrate that the formula can be used to predict the future window traffic data corresponding to each moment in a preset time period. The specific budget process can be set according to actual conditions and is not subject to excessive restrictions here.

[0039] In an optional example, considering that in the actual application process, there may be a window flow at a certain moment that is greater than the preset flow data threshold. Therefore, in this application document, the task scheduling system also includes a processor core 103, see Figure 2 As shown, Figure 2 Another task scheduling system provided in an embodiment of the present application, in which the task scheduler 101 is further configured to: When there is at least one future window flow data corresponding to a moment greater than or equal to a preset flow data threshold within a preset time period, determining whether the future window flow data corresponding to at least one moment is greater than the basic window flow data; When it is determined that the future window traffic data corresponding to at least one moment are all less than or equal to the basic window traffic data, the tasks from the current moment to the future preset moment are directly assigned to the processor core 103 for processing.

[0040] Specifically, if there is at least one moment within the preset time period where the future window traffic data is greater than or equal to the preset traffic data threshold, then directly having the HOE process all tasks may result in task processing delays. Therefore, it is necessary to wake up the processor core 103 for processing. However, to a certain extent, reducing resource consumption also considers the balance between service processing delays, etc. It is also necessary to determine whether the future window traffic data corresponding to at least one moment is greater than the base window traffic data.

[0041] In a specific example, if it is determined that the future window traffic data corresponding to at least one moment are all less than or equal to the basic window traffic data, the tasks from the current moment to the future preset moment are directly assigned to the processor core 103 for processing.

[0042] That is, the processor core 103 can directly process tasks from the current moment to a preset future moment, thereby achieving a refined balance between energy saving and performance.

[0043] Specifically, when the load is expected to increase but is still controllable, processor core 103 is selected to operate normally. Without excessively sacrificing potential performance risks, the energy overhead of enabling HOE itself is also saved, thereby achieving more refined energy efficiency management and enhancing the robustness and reliability of the system. This method greatly reduces the performance risks caused by slight deviations in the prediction model. It provides a safety buffer for the system. As long as the rising traffic does not exceed the known and processable baseline, it is considered a safe condition. This makes the entire scheduling system more robust and reliable in the face of network traffic uncertainty. By introducing more detailed judgment conditions, it ensures that the system can make the most reasonable and safest decisions in any traffic prediction scenario, thereby better achieving the goal of "energy conservation and emission reduction" as a whole.

[0044] Further, in actual application, in addition to the above-mentioned situations, other situations may also be included. For example, there is a moment when the future window traffic data is greater than the basic window traffic data at at least one moment. When such a situation occurs, the task scheduler 101 is further configured to: Determine the target moment closest to the current moment from the moment when the future window traffic data is greater than the basic window traffic data; The tasks from the current time to the target time are assigned to the processor core 103 for execution, and the tasks from the target time to a preset future time are assigned to the processor core 103 and the hardware acceleration device 102 for joint execution.

[0045] Specifically, it is to take into account that the preset time period itself is a time period. Within this time period, it is possible that the future window flow data is less than the basic window flow data from the current moment to a certain intermediate moment, and the data flow at the subsequent moment may be greater than the basic window flow data. Therefore, the target moment closest to the current moment can be determined from the moment when the future window flow data is greater than the basic window flow data, and then the tasks from the current moment to the target moment are assigned to the processor core 103 for execution, and the tasks between the target moment and the future preset moment are assigned to the processor core 103 and the hardware acceleration device 102 for joint execution.

[0046] That is, if the processor core 103 can handle the task by itself, the task will be directly handled by the processor core 103. If there may be a task processing delay in the processing by the processor core 103, part of the task can be assigned to the hardware acceleration device 102 for execution.

[0047] Specifically, the portion executed by the hardware acceleration device 102 may be a portion of the task corresponding to the traffic data exceeding the basic window.

[0048] This embodiment achieves lossless processing and performance assurance for burst traffic, ensuring system performance does not degrade even when facing predicted high loads. By enabling the hardware accelerator to proactively intervene and share the load, system throughput is maintained and data processing latency is strictly controlled, completely resolving the paradox of traditional energy-saving technologies where energy conservation reduces performance. Furthermore, this solution optimizes resource utilization, eliminating hardware acceleration resources before the target time; after the target time, the processor core and hardware accelerator work together to maximize their respective capabilities. This ensures seamless and full utilization of both computing resources, including the processor core and hardware accelerator, based on load fluctuations, preventing either from being idle or overloaded, and maximizing overall efficiency. Furthermore, by using a precise target time as the switching point, the system transitions smoothly and promptly from "processor core-only" to "processor core-and-hardware accelerator collaboration." This avoids performance jitter or inconsistencies that can occur during increasing traffic, providing users with a more stable service experience. This ensures worry-free system performance while ensuring extreme energy conservation.

[0049] In addition to the above situation, the embodiment of the present application also includes another situation, that is, when it is determined that the future window traffic data corresponding to each moment in the preset time period are greater than or equal to the preset traffic data threshold, then the task scheduler 101 is also used to assign the tasks corresponding to the basic window traffic data at each moment in the preset time period to the processor core 103 for execution, and assign the tasks that exceed the basic window traffic data at each moment in the preset time period to the hardware acceleration device 102 for execution.

[0050] The specific task allocation principle is similar to that described in the previous article, so I will not go into details here.

[0051] In an optional embodiment, the load information further includes task attribute information corresponding to the task to be processed; the hardware acceleration device 102 includes a protocol processing module, a packet filtering module, and a direct memory access (DMA) optimization module; and the task scheduler 101 is further configured to: Determine, based on the task attribute information, to classify the tasks performed by the hardware acceleration device 102; The classified tasks are sequentially assigned to one or more modules among the protocol processing module, the data packet filtering module, and the direct memory access optimization module for execution.

[0052] Specifically, because the hardware acceleration device 102 includes multiple modules, such as a protocol processing module, a packet filtering module, and a direct memory access module, when the task scheduler 101 assigns some or all tasks within a preset time period to the hardware acceleration device 102 for execution, it is also necessary to classify the tasks based on the task attribute information to determine which module to assign them to.

[0053] In a specific example, the task attribute information includes: the protocol type followed by the data packet to be transmitted and the size of the data packet to be transmitted. Then, the task scheduler 101 is specifically used to: Identify the protocol type followed by the data packet to be transmitted of the target task, where the target task is any task between the current time and a preset time in the future; When the protocol type is the first target protocol type, determining to assign the target task to the direct memory access optimization module for execution; Alternatively, when the protocol type is the second target protocol type, it is determined to allocate the target task to the protocol processing module for execution.

[0054] Alternatively, when the identified protocol type is any protocol type other than the first target protocol type and the second target protocol type, the task scheduler 101 is further configured to: Determine, based on the size of the data packet to be transmitted, whether the number of bytes occupied by the target data packet to be transmitted corresponding to the target task is greater than or equal to the preset target number of bytes; When it is determined that the number of bytes occupied by the target data packet to be transmitted is greater than or equal to the preset target number of bytes, determining to assign the target task to the data packet filtering module for processing; Alternatively, when it is determined that the number of bytes occupied by the target data packet to be transmitted is greater than or equal to a preset target number of bytes, it is determined to allocate the target task to the protocol processing module for execution.

[0055] Specifically, the first target protocol type is, for example, User Datagram Protocol (UDP), the second target protocol type is, for example, Transmission Control Protocol (TCP), and the preset target byte is, for example, 64 bits.

[0056] By precisely routing tasks to the most specialized hardware modules, we avoid the performance waste of "using a general-purpose processor to handle all tasks" or "using a single acceleration module to handle tasks it is not good at." In the embodiments of this application, each module can maximize its performance, significantly reducing task processing latency and improving overall throughput.

[0057] Furthermore, different hardware modules consume varying amounts of power. This solution allows simple tasks to be handled by smaller, more energy-efficient modules, while complex tasks are handled by high-performance modules. For example, this avoids the energy waste of having a large protocol processing module perform a simple DMA task. This allows for a precise match between energy consumption and computing requirements. Furthermore, hardware resources such as protocol processing, filtering, and DMA can be utilized in parallel and evenly, improving overall system resource utilization.

[0058] Furthermore, the solution can also achieve performance optimization based on data characteristics, adopting an efficient strategy of "large packets take the fast channel, small packets take the general channel". For large packets, they are quickly forwarded through the optimized data packet filtering module to maximize bandwidth utilization. For small packets, they are processed by a more powerful protocol processing module to cope with the complex processing logic they may need and optimize the packet processing rate. Secondary scheduling based on data characteristics ensures that the system can maintain high performance when processing unknown protocols. When encountering other types of protocols, the system will not crash or simply discard the data packet, but will start a set of backup, still efficient scheduling strategies. This ensures the robustness and engineering practicality of the system. In an optional embodiment, in addition to classifying tasks and assigning them to different processing modules in the HOE based on the aforementioned rules, the task scheduler 101 can also classify and process tasks based on the performance of each module. Specifically, the protocol processing module can parse / encapsulate the TCP / IP / UDP protocol stack and support Checksum hardware calculation. Checksum refers to a data summary value automatically and efficiently calculated by dedicated hardware circuits. Its core purpose is to detect whether errors occur during data transmission or storage.

[0059] The packet filtering module can implement Virtual Local Area Network (VLAN) / QoS classification and traffic shaping based on programmable matching logic.

[0060] The purpose of VLAN / QoS classification is to provide a basis for subsequent network processing (such as queuing, forwarding, and discarding). After assigning a priority tag, the device knows how to handle the packet. Traffic shaping is a mechanism that controls the output of data flows, aiming to align traffic with the data requirements of the upstream link or contract, thereby avoiding congestion and packet loss.

[0061] The packet filtering module classifies packets and assigns them to different priority queues (high, medium, and low). Based on priority, packets are selectively removed from these queues and sent out. For example, for high-priority queues, the module sends data at the fastest permitted rate. For low-priority queues, the module may restrict the amount of data they can send or temporarily hold them in case of network congestion, prioritizing high-priority packets.

[0062] Direct Memory Access (DMA) allows computer hardware devices (such as network cards, hard drive controllers, sound cards, and graphics processors) to read and write data directly to main memory without the need for the CPU to be actively involved. In embodiments of the present application, DMA can be used to transfer data directly from a device (such as a network card) to the memory space of the application that ultimately needs it, or vice versa, completely avoiding unnecessary copying between the kernel buffer and the user buffer. This significantly reduces memory bandwidth usage, lowers transmission latency, and further improves data transfer efficiency. It also reduces the triggering of ARM core interrupts.

[0063] In a specific example, the task attribute information may also include the data packet type. In addition to classifying tasks based on the protocol type followed by the data packet to be transmitted and the size of the data packet to be transmitted, the task scheduler 101 may also consider the data packet type and the capability attributes of each module, and based on these parameters, select a module from the protocol processing module, the data packet filtering module, and the direct memory access optimization module to perform the corresponding task, or select multiple modules to perform the task collaboratively.

[0064] For example, when a data packet belongs to a large data packet of the first preset type. The first preset type is a transmission type data packet, for example, greater than or equal to 1000 bytes, even if the data packet belongs to a TCP data packet. During the classification process of the task scheduler 101, the data packet can also be assigned to the protocol processing module and the DMA for collaborative processing.

[0065] Specific examples are as follows: Case 1: Large TCP packets (e.g. file downloads, video streaming); Packet size: large (e.g. 1500 bytes MTU); Protocol type: TCP; Traffic data: continuous and stable; Protocol type identification: TCP is identified. TCP requires complex state tracking (such as sequence numbers, acknowledgments, and retransmissions), but its checksum calculation and large-scale data handling are very regular.

[0066] Packet size analysis: This is a large data packet whose primary purpose is to transmit data payloads, not control signaling. Processing it focuses on efficient data movement, not complex analysis. Therefore, the protocol processing unit only needs to process TCP header information (connection establishment, acknowledgment, etc.), while the large data volume is efficiently handled entirely by the DMA hardware. This is because the DMA optimization module excels at "zero-copy" large-scale, continuous data movement. It can move data directly from the network card to the application's memory space, minimizing processor core 103 intervention and memory copy overhead.

[0067] In another specific example, Case 2: small UDP packets (such as DNS queries, VoIP voice packets).

[0068] Packet size: small (e.g., tens to hundreds of bytes); Protocol type: UDP; Traffic data: sudden and sporadic; Scheduler decision process: Protocol analysis: Identified as UDP. UDP is connectionless, and the processing logic is relatively simple without a complex state machine.

[0069] Packet size analysis: Packets are small, and the focus of processing is on rapid identification, classification, and forwarding, rather than data handling efficiency.

[0070] Decision: Such packets are preferentially assigned to the packet filtering module for processing.

[0071] The specific principle is that the packet filtering module's "programmable matching logic" can perform hardware-accelerated deep inspection of packets. It can quickly extract the target IP and port (for example, determine that this is a DNS query sent to port 53).

[0072] Based on pre-set rules, the packet filter module can instantly decide whether to drop a packet, forward it to a specific queue, or apply a QoS tag. For these small packets that require quick decisions, packet filtering is much faster and has lower latency than using more general protocol processing modules.

[0073] Case 3: Protocol interaction control packets (e.g. TCP Synchronize (SYN) / Finish packets, Internet Control Message Protocol (ICMP) packets); Packet size: very small (usually only a header, no data payload); Protocol type: TCP (control flags) or ICMP; Traffic data: irregular; Scheduler decision process: Protocol analysis: Identify packets as TCP SYN packets (requesting to establish a connection) or ICMP echo requests (ping requests). These packets are control signals for network protocols and require complex responses based on the state of the protocol stack.

[0074] Packet size analysis: The packet is very small and contains almost no data payload. The processing focus is on logical judgment rather than data movement.

[0075] Decision: Assign such packets to the protocol processing module for processing.

[0076] Specific principle: Establishing a TCP connection requires maintaining a connection table, allocating resources, generating sequence numbers, etc. These operations involve complex and variable state management and are best handled by the general computing power of the protocol processing unit.

[0077] By considering packet type, the scheduler can make more granular decisions. For example, even with TCP protocols, TCP control packets (such as SYN and ACK) can be assigned to the fast channel for processing to reduce latency, while TCP data packets can be allocated to high-bandwidth channels. For video streams, different processing strategies can be used for I-frames and P / B-frames. The task scheduler also assigns tasks based on the capabilities of each module, ensuring that each task is handled by the module that best excels at it, thereby fully utilizing the potential of each hardware module and avoiding resource misallocation. The coordinated execution of multiple modules achieves internal pipeline parallelism within the hardware, enabling the handling of complex tasks that would be beyond the capabilities of a single module, significantly improving processing power and throughput for complex workloads. Precisely allocating tasks to the most efficient modules avoids the excessive power consumption caused by overkill or underutilization. Tasks are completed with optimal energy consumption, further advancing the goal of energy conservation and emission reduction.

[0078] Figure 3 The overall structure diagram of the task scheduling system in this application document is shown in FIG. The system is described using a network card as an example. Figure 3As shown in the figure, in addition to the aforementioned task scheduler, hardware accelerator, and ARM core, the network card also includes peripheral processing logic, such as a system control processor (SCP), an inertial measurement unit (IMU), a complex programmable logic device (CPLD), a high-speed peripheral component interconnect (PCIe), a fifth-generation double data rate synchronous dynamic random-access memory (DDR5), and a serializer / deserializer (SerDes).

[0079] The functions performed by other peripheral components are not within the scope of this application document and will not be described in detail here.

[0080] In this structural diagram, because some of the processor core work has been assigned to the hardware accelerator by the task scheduler, compared to traditional technologies where the entire processor core performs the work, for example, a high-speed network card handling over 400G of traffic requires full ARM core operation, with all processing logic placed on the ARM core. Both the main frequency and power consumption must be maximized to meet current Standard Performance Evaluation Corporation (SPEC) evaluations, inevitably resulting in high power consumption. This is particularly true in scenarios where some servers typically have eight network cards and eight GPUs, where network card-side power consumption is more significant. In this application, not all tasks are executed by ARM cores, so power consumption in the ARM cores is significantly reduced. This effect is particularly pronounced when the number of network cards is particularly large.

[0081] Figure 4 A simple schematic diagram of the topological decomposition logic of the specific implementation principle of the above system of this application document is shown in FIG. Figure 4 Shown, including: The hardware accelerator includes a protocol processing module, a packet filtering module, and a direct memory access optimization module. The task scheduler (e.g., an Intelligent Task Dispatcher (ITD)) implements the intelligent traffic prediction algorithm described in this application document, which predicts future window traffic data corresponding to each moment in a preset time period between the current moment and a preset future moment. This algorithm, also known as the formula mentioned above, is used.

[0082] Through this intelligent prediction algorithm, it is determined whether the arm core can enter a deep sleep state or wake up operation based on the predicted traffic.

[0083] The above process has been introduced in detail in the previous article, so I will not go into details here.

[0084] The embodiment of the present application provides a task scheduling method, see Figure 5 As shown, the method is applied to a task scheduling system, which includes a task scheduler and a hardware acceleration device. The method is executed by the task scheduler and specifically includes the following method steps: Step S501: Perform dynamic load detection to obtain load information corresponding to the task to be processed.

[0085] The load information includes basic window traffic data and current window traffic data.

[0086] Step S502: predict the future window traffic data corresponding to each moment in a preset time period from the current moment to a future preset moment based on the basic window traffic data, the current window traffic data, and the predetermined historical traffic data attenuation coefficient.

[0087] Step S503: When the future window traffic data corresponding to each moment in the preset time period is lower than the preset traffic data threshold, the task from the current moment to the future preset moment is scheduled to the hardware acceleration device, so that the hardware acceleration device executes the task from the current moment to the future preset moment.

[0088] In an optional embodiment, the task scheduling system further includes a processor core; when there is at least one moment in the preset time period where future window traffic data is greater than or equal to a preset traffic data threshold, the method further includes: Determine whether future window traffic data corresponding to at least one moment is greater than basic window traffic data; When it is determined that the future window traffic data corresponding to at least one moment are all less than or equal to the basic window traffic data, the tasks from the current moment to the future preset moment are directly assigned to the processor core for processing.

[0089] In an optional embodiment, based on the basic window traffic data, the current window traffic data, and the predetermined historical traffic data attenuation coefficient, the future window traffic data corresponding to each moment in a preset time period from the current moment to the future preset moment is predicted, which is specifically expressed by the following expression:

[0090] in, is the future window traffic data corresponding to the i-th moment in the preset time period, As the basic window traffic data, is the attenuation coefficient of historical traffic data, L is the ratio between the current window traffic data and the basic window traffic data, is the time difference between the current moment and the i-th moment, where the i-th moment is any moment within the preset time period.

[0091] In an optional embodiment, when it is determined that there is a moment in at least one moment when the future window flow data is greater than the basic window flow data, the method further includes: Determine the target moment closest to the current moment from the moment when the future window traffic data is greater than the basic window traffic data; The tasks from the current moment to the target moment are assigned to the processor core for execution, and the tasks from the target moment to a preset future moment are assigned to the processor core and the hardware acceleration device for joint execution.

[0092] In an optional embodiment, when it is determined that the future window traffic data corresponding to each moment in the preset time period is greater than or equal to the preset traffic data threshold, the method further includes: The tasks corresponding to the basic window traffic data at each moment in the preset time period are assigned to the processor core for execution, and the tasks exceeding the basic window traffic data at each moment in the preset time period are assigned to the hardware acceleration device for execution.

[0093] In an optional embodiment, the load information further includes task attribute information corresponding to the task to be processed; the hardware acceleration device includes a protocol processing module, a packet filtering module, and a direct memory access optimization module; when the future window traffic data corresponding to each moment in a preset time period is lower than a preset traffic data threshold, the task from the current moment to the future preset moment is dispatched to the hardware acceleration device, specifically including: Determine, based on the task attribute information, to classify the tasks executed by the hardware acceleration device; The classified tasks are sequentially assigned to one or more modules among the protocol processing module, the data packet filtering module, and the direct memory access optimization module for execution.

[0094] In an optional embodiment, the task attribute information includes: the protocol type followed by the data packet to be transmitted and the size of the data packet to be transmitted; and determining the classification of the task to be performed by the hardware acceleration device based on the task attribute information specifically includes: Identify the protocol type followed by the data packet to be transmitted of the target task, where the target task is any task between the current time and a preset time in the future; When the protocol type is the first target protocol type, determining to assign the target task to the direct memory access optimization module for execution; Alternatively, when the protocol type is the second target protocol type, it is determined to allocate the target task to the protocol processing module for execution.

[0095] In an optional embodiment, when the identified protocol type is any protocol type other than the first target protocol type and the second target protocol type, the method further includes: Determine, based on the size of the data packet to be transmitted, whether the number of bytes occupied by the target data packet to be transmitted corresponding to the target task is greater than or equal to the preset target number of bytes; When it is determined that the number of bytes occupied by the target data packet to be transmitted is greater than or equal to the preset target number of bytes, determining to assign the target task to the data packet filtering module for processing; Alternatively, when it is determined that the number of bytes occupied by the target data packet to be transmitted is greater than or equal to a preset target number of bytes, it is determined to allocate the target task to the protocol processing module for execution.

[0096] The specific implementation process of a task scheduling method provided in the embodiments of the present application has been described in detail in the aforementioned embodiments, so it will not be repeated here.

[0097] An embodiment of the present application provides a task scheduling method in which a task scheduler dynamically performs load detection and predicts window traffic data for each moment in the future based on the base window traffic and current window traffic data in the load information, as well as the decay coefficient of historical traffic data. If the window traffic data at all moments in a preset future time period does not exceed a preset traffic data threshold, all tasks corresponding to the time period from the current moment to the preset future moment can be directly scheduled for execution on a hardware acceleration device. This method accurately predicts future low-load periods and intelligently offloads network tasks from the ARM core to a dedicated hardware acceleration device during this window, thereby achieving deep energy savings. This solution predicts a continuous, stable low-traffic time window and completely offloads tasks to the hardware acceleration device during this period. During this period, the ARM core can enter a deeper, longer-term low-power state (even hibernation) without having to be frequently awakened. This significantly reduces the additional power consumption caused by state switching, achieving true deep energy savings. Furthermore, the decision-making of this application is predictive rather than reactive. Before making a scheduling decision, it predicts that traffic will be "consistently below the threshold" for the entire future period. This means that the possibility of sudden high traffic is essentially eliminated during this preset time period. Therefore, there is no need to worry about not being able to respond in time when a burst of traffic hits. This application transforms the unpredictable risk of "burst traffic" into a manageable "deterministic low traffic" window through prediction. Moreover, because this application is a dynamic prediction, it does not rely on any fixed threshold as a criterion for executing actions (although there is a preset traffic threshold, it is used to compare with the dynamic prediction value, rather than as a fixed threshold for the accumulated amount of data packets). The system continuously performs dynamic calculations based on the "basic window traffic", "current window traffic" and "historical attenuation coefficient" to predict traffic at every moment in the future. This method is adaptive and intelligent, and can flexibly respond to dynamic changes in network traffic, thereby maximizing the energy-saving window and achieving the best balance between energy efficiency and performance while ensuring no delay risk.

[0098] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned task scheduling method embodiments when running.

[0099] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0100] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned task scheduling method embodiments are implemented.

[0101] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0102] The above is a detailed introduction to a task scheduling system and method provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only intended to help understand the method and core ideas of the present application. It should be pointed out that, for those skilled in the art, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A task scheduling system, characterized in that: The task scheduling system includes: a task scheduler and a hardware acceleration device; The task scheduler is configured to perform dynamic load detection and obtain load information corresponding to a task to be processed, wherein the load information includes basic window traffic data and current window traffic data; predict future window traffic data corresponding to each moment in a preset time period from the current moment to a preset future moment based on the basic window traffic data, the current window traffic data, and a predetermined historical traffic data attenuation coefficient; and when the future window traffic data corresponding to each moment in the preset time period are all lower than a preset traffic data threshold, dispatch the task from the current moment to the preset future moment to the hardware acceleration device; The hardware acceleration device is used to execute the task from the current moment to the future preset moment.

2. The system according to claim 1, wherein: The task scheduler is specifically used to predict the future window traffic data corresponding to each moment in a preset time period from the current moment to a future preset moment using the following expression: in, is the future window traffic data corresponding to the i-th moment within the preset time period, is the basic window traffic data, is the attenuation coefficient of the historical traffic data, L is the ratio between the current window traffic data and the basic window traffic data, is the time difference between the current moment and the i-th moment, where the i-th moment is any moment within the preset time period.

3. The system according to claim 1, wherein: The task scheduling system further includes a processor core, and the task scheduler is further configured to: When there is at least one future window flow data corresponding to a moment within the preset time period that is greater than or equal to the preset flow data threshold, determining whether the future window flow data corresponding to at least one moment is greater than the basic window flow data; When it is determined that the future window traffic data corresponding to at least one moment are all less than or equal to the basic window traffic data, the tasks from the current moment to the future preset moment are directly allocated to the processor core for processing.

4. The system according to claim 3, characterized in that The task scheduler is also used to: When it is determined that there is a moment in which the future window flow data is greater than the basic window flow data in at least one moment, determining a target moment closest to the current moment from the moments in which the future window flow data is greater than the basic window flow data; The tasks from the current moment to the target moment are assigned to the processor core for execution, and the tasks from the target moment to the future preset moment are assigned to the processor core and the hardware acceleration device for joint execution.

5. The system according to claim 3 or 4, characterized in that The task scheduler is also used to: When it is determined that the future window traffic data corresponding to each moment in the preset time period is greater than or equal to the preset traffic data threshold, the tasks corresponding to the basic window traffic data at each moment in the preset time period are allocated to the processor core for execution, and the tasks exceeding the basic window traffic data at each moment in the preset time period are allocated to the hardware acceleration device for execution.

6. The system according to any one of claims 1 to 4, characterized in that: The load information also includes task attribute information corresponding to the task to be processed; the hardware acceleration device includes a protocol processing module, a data packet filtering module, and a direct memory access optimization module; the task scheduler is further used to: determining, according to the task attribute information, to classify the tasks executed by the hardware acceleration device; The classified tasks are sequentially assigned to one or more modules among the protocol processing module, the data packet filtering module, and the direct memory access optimization module for execution.

7. The system according to claim 6, characterized in that The task attribute information includes: the protocol type followed by the data packet to be transmitted and the size of the data packet to be transmitted; the task scheduler is specifically used to: Identifying a protocol type followed by the to-be-transmitted data packet of a target task, wherein the target task is any task between the current moment and the future preset moment; When the protocol type is the first target protocol type, determining to assign the target task to the direct memory access optimization module for execution; Alternatively, when the protocol type is the second target protocol type, it is determined to allocate the target task to the protocol processing module for execution.

8. The system according to claim 7, characterized in that The task scheduler is further configured to: When the protocol type is identified as any protocol type other than the first target protocol type and the second target protocol type, determining, based on the size of the data packet to be transmitted, whether the number of bytes occupied by the target data packet to be transmitted corresponding to the target task is greater than or equal to a preset target number of bytes; When it is determined that the number of bytes occupied by the target data packet to be transmitted is greater than or equal to the preset target number of bytes, determining to assign the target task to the data packet filtering module for processing; Alternatively, when it is determined that the number of bytes occupied by the target data packet to be transmitted is greater than or equal to the preset target number of bytes, it is determined to allocate the target task to the protocol processing module for execution.

9. A task scheduling method, characterized in that: The method is applied to a task scheduling system, which includes a task scheduler and a hardware acceleration device. The method is executed by the task scheduler and includes: Perform dynamic load detection to obtain load information corresponding to the task to be processed, wherein the load information includes basic window traffic data and current window traffic data; Predicting future window traffic data corresponding to each moment in a preset time period from the current moment to a future preset moment based on the basic window traffic data, the current window traffic data, and a predetermined historical traffic data attenuation coefficient; When the future window traffic data corresponding to each moment in the preset time period is lower than the preset traffic data threshold, the task from the current moment to the future preset moment is scheduled to the hardware acceleration device, so that the hardware acceleration device executes the task from the current moment to the future preset moment.

10. The method according to claim 9, characterized in that The task scheduling system further includes a processor core; when there is future window traffic data corresponding to at least one moment within the preset time period that is greater than or equal to the preset traffic data threshold, the method further includes: Determine whether future window traffic data corresponding to at least one moment is greater than the basic window traffic data; When it is determined that the future window traffic data corresponding to at least one moment are all less than or equal to the basic window traffic data, the tasks from the current moment to the future preset moment are directly allocated to the processor core for processing.

Citation Information

Patent Citations

  • Technologies for offloading and on-loading data for processor / coprocessor arrangements

    CN107408062A

  • Lightweight network forwarding unloading method based on dynamic filtering

    CN117240792A

  • Edge computing resource management method and system, storage medium and electronic equipment

    CN117573352A

  • Edge computing task scheduling method and device, equipment, storage medium and product

    CN118069317A

  • Task scheduling balancing method and device and readable storage medium

    CN119739482A