Task Scheduling Method and System, Electronic Device, Storage Medium
By identifying and marking the computing node type, the host node matches the appropriate computing nodes according to the task type for scheduling, solving the computing bottlenecks and delay problems in the traditional computing architecture and achieving efficient task execution.
Patent Information
- Application Number
- CN202510380707.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-28
AI Technical Summary
Traditional computing architectures have obvious computing bottlenecks and delay problems when facing different computing needs, resulting in poor task execution efficiency.
By identifying the type of computing node and assigning corresponding tags, the host nodes match appropriate computing nodes according to the task type to schedule, and use the computing efficiency differences between multiple node types to achieve efficient task scheduling.
It improves task execution efficiency, is compatible with multiple computing nodes, optimizes the utilization of computing resources, and reduces task queuing time.
Smart Images

Figure CN119883578B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a task scheduling method, a task scheduling system, an electronic device, and a computer-readable storage medium. Background Art
[0002] With the rapid development of big data, artificial intelligence, and high-performance computing, the computing bottleneck and latency problems of traditional computing architectures have become increasingly obvious. To solve the foregoing problems, various computing architectures adapted to application scenarios and different computing requirements have emerged. Just as different computing architectures are mainly designed to adapt to different computing requirements, there are still obvious computing bottlenecks and latency problems when they execute other types of computing requirements, resulting in poor task execution efficiency. Summary of the Invention
[0003] This application provides a task scheduling method, a task scheduling system, an electronic device, and a computer-readable storage medium to at least solve the problem of poor task execution efficiency in related technologies.
[0004] This application provides a task scheduling method, which includes: the host node of the server obtains the node information of the computing node, identifies the node type represented by the node information, and assigns a type label matching the node type to the computing node; wherein, the node type includes at least a first type and a second type, the computing efficiency of the first type is higher than that of the second type, and the data processing efficiency of the second type is higher than that of the first type; the host node obtains the task to be executed and evaluates the task type of the task to be executed; the host node uses the node type matching the task type as the target type, selects the computing node carrying the target type label as the target node, and schedules the task to be executed to the target node so that the target node executes the task to be executed.
[0005] This application also provides a task scheduling system for implementing any of the above task scheduling methods; the task scheduling system includes: a host node, a computing node, and a switch module; the computing node includes computing nodes of multiple node types; the node type includes at least a first type and a second type, the computing efficiency of the first type is higher than that of the second type, and the data processing efficiency of the second type is higher than that of the first type; the switch module is respectively connected to the host node and the computing node, and the host node and the computing node communicate through the switch module.
[0006] This application also provides an electronic device, which includes: a memory for storing a computer program; a processor for implementing the steps of any of the above task scheduling methods when executing the computer program.
[0007] The present application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the steps of any one of the above task scheduling methods are implemented.
[0008] Through the present application, since the task scheduling system can be compatible with computing nodes of multiple node types. That is, the node types at least include a first type and a second type, and the computing efficiency and data processing efficiency of the first type and the second type are different. The host node can identify the node type of the computing node and assign a label. When obtaining a task to be executed, the host node can evaluate the node type suitable for the task to be executed, so as to select a computing node suitable for the task to be executed as the target node, and schedule the task to be executed to the target node for execution. Therefore, the technical problem of poor task execution efficiency can be solved, and the technical effect of being compatible with multiple computing nodes to reasonably schedule the computing nodes for executing the task to be executed, thereby improving the task execution efficiency can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0010] Figure 1 Structural schematic diagram of an embodiment of the task scheduling application scenario of the present application;
[0011] Figure 2 Structural schematic diagram of an embodiment of the task scheduling system of the present application;
[0012] Figure 3 Structural schematic diagram of another embodiment of the task scheduling system of the present application;
[0013] Figure 4 Flow schematic diagram of an embodiment of the task scheduling method of the present application;
[0014] Figure 5 Flow schematic diagram of an embodiment of identifying the node type of the computing node of the present application;
[0015] Figure 6 Flow schematic diagram of another embodiment of the task scheduling method of the present application;
[0016] Figure 7 Structural schematic diagram of an embodiment of the electronic device of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0018] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0019] In order to enable those skilled in the art of this technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0020] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the task scheduling method depends, the specific application environment architecture or specific hardware architecture is described herein.
[0021] Please refer to Figure 1 , Figure 1 , which is a schematic structural diagram of an embodiment of the task scheduling application scenario of the present application.
[0022] In one embodiment, personnel with identities such as users and administrators can send a task execution request to the server 12 through the external device 11. Among them, the task execution request may carry the task to be executed, or the server 12 may obtain the task to be executed in response to obtaining the task execution request, or the server 12 may generate the task to be executed in response to obtaining the task execution request, which is not limited herein. To achieve the acquisition of the task to be executed as described above.
[0023] The external device 11 may include intelligent wearable devices such as smart watches, mobile phones, computers, other servers, etc., which is not limited herein.
[0024] The server 12 is a specific IT device that provides computing power and runs software applications in a network environment, and it can provide computing or application services for other client machines (such as terminal devices such as personal computers and smart phones) in the network. Generally speaking, the server 12 can have the ability to undertake response service requests, undertake services, and guarantee services.
[0025] It should be noted that the server described in this embodiment can represent a single server or a server cluster, etc., which is not limited herein. As the name implies, a server cluster includes multiple servers, and the multiple servers are communicatively connected and can cooperate to implement the task to be executed.
[0026] Further, the server may include a task scheduling system to schedule and execute the task to be executed by the task scheduling system. That is, the embodiment of the present application provides a task scheduling system.
[0027] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of an embodiment of the task scheduling system of the present application.
[0028] In one embodiment, the task scheduling system may include a host node, a switch module, and a computing node.
[0029] A host generally represents the main collective part of a computer except for input and output devices, and is also a control box for placing a motherboard and other main components. A host usually includes a CPU (Central Processing Unit, central processing unit), etc. The host node in this embodiment can be considered as a server host, that is, a special computer system, which can be used to provide services and manage network resources, etc., and can be responsible for storing, processing, and distributing data, etc., and can also run various application programs to provide the required services for the client, that is, at least can implement obtaining a task execution request and executing the task to be executed described above.
[0030] As the name implies, a computing node is a node that can be responsible for undertaking computing tasks and performing data processing and storage within a server. In this embodiment, the computing node may include computing nodes of multiple node types.
[0031] Among them, the node type may at least include a first type and a second type. The computing efficiency of the first type is higher than that of the second type, and the data processing efficiency of the second type is higher than that of the first type. Thus, in this embodiment, it can at least adapt to multiple tasks to be executed. Generally speaking, when the task to be executed focuses on computing efficiency, the computing nodes of the first type can be preferentially scheduled for execution; when the task to be executed focuses on data processing, the computing nodes of the second type can be preferentially scheduled for execution. That is to say, when there are no idle or available computing nodes of the node type that is more suitable for the task to be executed, the computing nodes of other node types can also be scheduled to execute the task to be executed, so as to reduce the queuing time of the task to be executed and also improve the task scheduling and task execution efficiency.
[0032] A switch means "a switch", that is, it can be considered that a switch is a network device for forwarding electrical (optical) signals. The switch can provide an exclusive electrical signal path for any two network nodes of the access device. In this embodiment, the switch module can be respectively connected to the host node and the computing node, and the host node and the computing node communicate with each other through the switch module.
[0033] Furthermore, in this embodiment, there can also be computing nodes of three, four, etc. node types, which is beneficial to adapting to more application scenarios and computing requirements. And in this embodiment, by connecting the host node and the computing node through the switch module, it is beneficial to improve the scalability of the task scheduling system. Even if the task scheduling system and the server have been put into actual use, new computing nodes of the existing node types or other node types of computing nodes can still be added, and the newly added computing nodes are coupled to the switch module to realize the incorporation of the new computing nodes into the task scheduling system.
[0034] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of another embodiment of the task scheduling system of this application.
[0035] In one embodiment, the first type of computing node described in the foregoing can be equivalent to a near-memory computing node, and the second type of computing node can be a near-storage computing node.
[0036] The task scheduling system can also include a host memory and a storage device.
[0037] The host memory is connected to the switch module and is connected to the host node and the first type of computing node through the switch module.
[0038] The storage device is connected to the second type of computing node and is connected to the switch module and the host node through the second type of computing node.
[0039] Relatively speaking, in this embodiment, the storage device is directly connected to the second type of computing node, while the first type of computing node is indirectly connected to the host memory. The reason for such a design in this application is that considering that the main function of the storage device is storage, storage devices can be attached to each second type of computing node respectively, so as to effectively shorten the physical link between the computing node and the storage device, thereby further improving the data processing efficiency. While collaborating with the first type of computing node, the host memory also needs to implement the relevant computing work of the host node. Therefore, the host memory, the first type of computing node, and the host node are connected to each other through the switch module to balance the working links between the host memory and all parties.
[0040] Optionally, the switch module may further include a first module and a second module. The first module and the second module are connected, and at least one of them is connected to the host node. That is, as Figure 3 exemplified, it may be that the first module is connected to the host node, or the second module is connected to the host node, so as to save the interface resources of the host node; or both the first module and the second module may be connected to the host node, so that both the first type of computing node and the second type of computing node can be connected to the host node through a shorter link, taking into account the task scheduling efficiency of multiple tasks to be executed.
[0041] Specifically, the first module is used to manage the links of the host node, the host memory, and the first type of computing node. That is, the first module can be connected to the host memory and the first type of computing node.
[0042] The first module is used to manage the links of the host node and the first type of computing node. The second module is connected to the second type of computing node.
[0043] It can be seen that in this embodiment, the computing nodes of each node type are respectively connected to a switch module. In other words, the computing nodes of the same node type can be connected to the same switch module, which is beneficial to improving the rationality of the modular architecture of the task scheduling system and can also facilitate the subsequent incorporation of computing nodes of new node types. At the same time, it can also decouple the computing nodes of different node types.
[0044] Please continue to refer to Figure 3 . In an embodiment, the component modules of a first type of computing node are integrated in a field programmable gate array.
[0045] As Figure 3 exemplified, the component modules of the first type of computing node may include a target bus protocol, a target bus protocol host connection module, a heterogeneous device memory controller, an embedded processor, a heterogeneous device memory, and a near-memory computing accelerator. Among them, the target bus protocol may be, for example, a PCIe (peripheral component interconnect express, a high-speed serial computer expansion bus standard) protocol, a CXL (Compute Express Link, an open interconnect standard), etc. Figure 3 Taking the CXL protocol as an example, two protocols, namely.mem and.cache, are exemplified. The.mem can be used to expand the system storage, and the.cache can be used to expand the system cache. That is, Figure 3The constituent modules exemplified herein can be integrated into a Field Programmable Gate Array (FPGA). In this way, the rich resources of the FPGA can be utilized to reduce the need for additional mounted constituent modules, thereby further shortening the communication link within the first type of computing node. For example, by adjusting at the hardware level, the computing efficiency of the first type of computing node itself can be improved.
[0046] Similarly, the constituent modules of a second type of computing node can be integrated into a Field Programmable Gate Array.
[0047] As Figure 3 exemplified herein, the constituent modules of the second type of computing node can include a hard disk controller, a target bus protocol, a target bus protocol host connection module, a heterogeneous device memory controller, an embedded processor, heterogeneous device memory, and a near-memory computing accelerator. Among them, the target bus protocol can be, for example, a PCIe (peripheral component interconnect express, a high-speed serial computer expansion bus standard) protocol, a CXL (Compute Express Link, an open interconnect standard), etc. Figure 3 Taking the CXL protocol as an example as exemplified herein, the target bus protocol of the second type of computing node supports three protocols: .io, .mem, and .cache. .io can be used for discovery, configuration, register access, error reporting, host node physical address lookup, interrupts, etc. .mem can be used to expand system storage, and .cache can be used to expand system cache. That is to say, Figure 3 The constituent modules exemplified herein can be integrated into a Field Programmable Gate Array (FPGA).
[0048] Furthermore, the flash memory storage particles of the storage device (such as flash (flash memory) storage particles) can also be integrated into the Field Programmable Gate Array of the second type of computing node.
[0049] In this way, the rich resources of the FPGA can be utilized to reduce the need for additional mounted constituent modules, thereby further shortening the communication link within the second type of computing node. For example, by adjusting at the hardware level, the computing efficiency of the second type of computing node itself can be improved.
[0050] As exemplified in this embodiment, the target bus protocol may be the CXL protocol, and the switch module may include a CXLSwitch (switch). The target bus protocol communication module may be a CXL Device EP (CXL device), which is used to implement a high-speed and efficient interconnection device between a CPU and accelerators such as a GPU (Graphics Processing Unit) and an FPGA. The target bus protocol included in the first type and / or the second type of computing nodes may also be a CXL Device EP. The target bus protocol host connection module may be a CXL Host Bridge, which is a concept in the CXL protocol. It may represent a group of shared logic CXL Root Ports (root ports) or other pairings, and can be regarded as a software abstraction concept that can be used to identify and manage CXL devices.
[0051] In an alternative embodiment, the computing nodes may also be integrated into a programmable controller such as a CPLD (Complex Programmable Logic Device), which is not limited herein.
[0052] As described above, the node type may also be three, four, etc. For example, it may further include computing nodes of the near-memory GPU type and computing nodes of the near-memory FPGA type. Further, a switch module may be assigned to the computing nodes of each node type, and they are connected to other switch modules or host nodes through the switch module, which is not limited herein.
[0053] The host node, the first type of computing node, and the second type of computing node of the present application will be described in detail below.
[0054] The host node is the core engine of the unified architecture of the task scheduling system of the present application, and it may adopt a CPU and a main memory architecture. The CPU may be responsible for the overall control of the system, task scheduling, and general computing tasks. Through the logical processing ability of the CPU, it can process complex operating systems and application programs. A CXL Host Bridge interface may be integrated in the CPU, and it is connected to the CXL Switch interconnection network through this interface to ensure that the CPU can perform high-speed data interaction and instruction transmission with other nodes in the entire network architecture. The main memory may be connected to the CXL Switch interconnection network through a CXL Device EP chip, enabling the main memory to perform data transmission in an efficient CXL network.
[0055] In Figure 3In the task scheduling system exemplified herein, the first type of computing node can be considered as the near-memory computing node among the first type and the second type.
[0056] The near-memory computing node can integrate the computing logic near the host memory, which can reduce the memory access latency while offloading the CPU workload and improve the data processing speed. Among them, the memory can use high-speed DRAM (Dynamic Random Access Memory) or other emerging memory technologies such as HBM (High Bandwidth Memory) to provide large-capacity and high-bandwidth memory support.
[0057] The near-memory computing node can be integrated near the host memory and use FPGA as the basis. The FPGA contains a CXL Device EP interface, which mainly realizes the interconnection network connection between the near-memory computing node and the CXL Switch. Relying on the ability of the CXL interface to support memory semantic access, it connects to the device memory through the CXL.mem and CXL.cache modules, enabling the host node and the second type of computing node to access the device memory as if accessing local memory, which is beneficial to ensuring the consistency of the device memory in the system. Near-memory computing can adopt the method of an embedded processor + accelerator, and the embedded processor and accelerator access the device memory and the host memory through the CXL Host Bridge.
[0058] The near-memory computing node is used to directly calculate the data in the memory, and can offload a large number of computing tasks of the CPU through high-performance computing capabilities. For example, during the machine learning training process, the near-memory computing unit can perform calculation operations such as matrix multiplication on batch data, reducing the calculation and transmission pressure of the data.
[0059] In Figure 3 In the task scheduling system exemplified herein, the second type of computing node can be considered as the near-storage computing node among the first type and the second type.
[0060] The near-storage computing node means that the computing function can be integrated near the storage device to reduce the transmission latency between the storage and the computing unit. Among them, the storage device can use high-performance storage media such as Flash storage particles and NVMe (Non-Volatile Memory Host Controller Interface Specification) new storage to provide large-capacity data storage capabilities.
[0061] The near-storage computing node can be integrated near the storage device and based on FPGA. The FPGA can include a CXL Device EP interface, which mainly realizes the connection between the near-storage computing unit and the CXL Switch interconnection network, and connects to the SSD (Solid State Disk or Solid State Drive) storage controller through the CXL.io module, so as to realize the end-to-end high-speed data transmission between the storage device, the host node, the near-storage computing node and the near-memory computing node. It can also rely on the ability of the CXL interface to support memory semantic access, connect to the device memory through the CXL.mem and CXL.cache modules, and enable the host computing node and the near-memory computing node to access the device memory as if accessing local memory, ensuring the consistency of the device memory in the system. Near-storage computing adopts the method of embedded processor + accelerator, and the embedded processor and accelerator access the device memory and the host memory through the CXL Host Bridge.
[0062] The near-storage acceleration computing unit is used for data preprocessing, such as data filtering, sorting, compression, etc., to reduce the amount of data that needs to be transmitted to the host computing node. For example, in database applications, the near-storage computing unit can perform preliminary screening of data at the storage end and only transmit the data that meets the conditions to the host computing node for further processing.
[0063] The difference between the target bus protocols of the first type of computing node and the second type of computing node is that the first type of computing omits the CXL.io used to access the storage device.
[0064] As exemplified above, in this embodiment, an interconnection network can be formed through CXL, which is equivalent to using CXL as the communication bridge of the task scheduling system.
[0065] The CXL interconnection network can connect the host computing node, the near-storage computing node and the near-memory computing node. As Figure 3 exemplified, the multi-level CXL Switch constructs and expands the CXL network, which can support the connection of the host node and multiple near-storage computing nodes and near-memory computing nodes. It should be noted that Figure 3 only one first type of computing node and one second type of computing node are exemplified. In the actual task scheduling system, the number of the first type of computing nodes and the second type of computing nodes can both be multiple.
[0066] Through the CXL interconnection network, all the memories in the system are supported to be uniformly addressed. The CXL Switch can quickly forward data according to the destination address of the data packet, realizing high-speed, low-latency, and consistency-supported data transmission under the unified architecture of near-storage and near-memory.
[0067] Through the above unified CXL interface, unified data access to storage devices, host memory, and device memory can be achieved. Relying on the above hardware architecture foundation and integrating the characteristics of near-storage and near-memory computing, task execution under the unified computing architecture is proposed, and through this method, efficient execution of tasks under the unified computing architecture is realized; a different cache management mode that supports the CXL coherence protocol is also proposed, which aims to solve the system performance bottleneck and limited scalability problems faced when supporting CXL cache coherence in a large-scale node environment. The aim is to build an efficient near-storage and near-memory unified computing architecture with data coherence capabilities. Details will be elaborated later.
[0068] As Figure 3 exemplified in
[0069] Specifically, the server selects the current cache management mode from multiple cache management modes. Among them, the cache management modes of the server at least include a global mode and a local mode.
[0070] The global mode means global byte addressing of the server memory for global scheduling of the server memory by the host node; the server memory can include host memory and compute node memory.
[0071] Generally speaking, the global mode can be regarded as global asymmetric cache coherence management, with a centralized control core, that is, the host node is the master node and the compute node is the slave node. The host node can perform global byte addressing on the host memory and compute node memory in the task scheduling system. The host node can achieve cache coherence of global memory on the host side through the CXL Host Bridge and CXL DeviceEP interfaces. At this time, the CXL HostBridge module inside the compute node can be set invalid, and the embedded processor only serves as the computing resource inside the node and cannot access the memory resources of other nodes. The host node is responsible for maintaining data coherence and distributing control instructions to each slave node. The slave node is responsible for storing and providing data, updating the data according to the instructions of the master node, and maintaining data coherence.
[0072] Through the centralized management of the master node, the update and propagation of data can be effectively controlled, unnecessary message interactions can be reduced, it is better in terms of maintenance efficiency and network overhead, and the performance is higher than that of the asymmetric mode. In data-intensive tasks, in pursuit of higher performance, global asymmetric coherence dynamically adjusts the parameters of the cache coherence protocol based on the system load. In high-load situations, the global protocol reduces the frequency of directory updates and reduces communication overhead; in low-load situations, the global protocol increases the frequency of directory updates to improve data coherence.
[0073] The local mode means that the computing node independently controls its computing node memory and synchronizes the data changes of the computing node memory with other computing nodes.
[0074] Generally speaking, the local mode is equivalent to global symmetric cache coherence management. The control core is distributed, and each computing node works in a symmetric mode. The CXL Host Bridge and CXL Device EP interfaces in the computing node are both set to be valid. At this time, the embedded processor not only serves as the computing resource within the node but also as the control node of the node to access the memory resources of other nodes. In this way, computing nodes can share data through the global symmetric cache coherence management mechanism and can be independent of the CPU of the host node. When the data of a cache node (a memory includes multiple cache nodes) changes, an update message can be sent to all other nodes (computing nodes and host nodes) to inform the change of the data. After receiving the message, other nodes can update their cache data according to the message content to ensure the data consistency of all nodes.
[0075] That is to say, the local mode has high flexibility and scalability. Computing nodes can relatively freely join or leave the system, reducing the impact of the risk of single-point failure on the overall architecture.
[0076] Specifically, the computing node can maintain a local cache directory; the host node can maintain a global cache directory; among them, the global cache directory includes the node identifier of the computing node and the local cache directory. In this way, it can be switched between the global mode and the local mode.
[0077] The computing node can monitor the cache state of the cache blocks included in the computing node to update the local cache directory it maintains, and then update the global cache directory; among them, the cache state at least includes a valid state, an invalid state, and a dirty state.
[0078] It should be noted that although the global mode and the local mode are named "global" and "local" in this article, they are actually both global cache management modes. The naming is considered from the range of server memory that can be controlled by the cache control core. The global mode is for the host node to conduct macroscopic scheduling of the server memory, and the local mode is for each computing node to manage its own computing node memory respectively and synchronize relevant information. It is not a special limitation on the cache management of the task scheduling system.
[0079] The working principle and the switchable principle of implementing the global mode and the local mode are elaborated in detail below.
[0080] At the hardware design level, it can include two aspects: node interface integration and interface compatibility design.
[0081] Node interface integration involves integrating interfaces that support heterogeneous cache coherence protocols for all nodes (host nodes and compute nodes). For example, for the CXL protocol, interfaces such as CXL Host Bridge and CXL Device EP can be designed.
[0082] In the motherboard design of the host node, the CXL Host Bridge can be integrated into the motherboard circuit, which is responsible for communicating with other devices that support the CXL protocol, providing a high-speed data transmission channel, and being able to handle complex protocol interactions. For example, during design, a motherboard that supports CXL can be paired, and the CXL Host Bridge can be connected through specific slots and circuits, enabling the host to communicate efficiently with the compute node.
[0083] For the compute node, the CXL Device EP interface can be integrated. The compute node can be an independent board or module. During design, the CXL Device EP chip can be embedded on the board and connected to the storage or memory module inside the compute node through appropriate PCB (Printed Circuit Board) wiring. For example, when designing the second type of compute node, the CXL Device EP chip can be connected to the SSD controller, enabling the data in the SSD to interact with the host or other nodes through the CXL interface.
[0084] The purpose of interface compatibility design is to ensure good compatibility between the CXL Host Bridge and the CXL Device EP in design, and to be able to adapt to two cache management modes.
[0085] In terms of protocol version support, the interface can be controlled to support multiple CXL protocol versions to ensure normal operation in different system environments and application scenarios. For example, support for versions such as CXL 1.0, CXL 2.0, and CXL 3.0 at the same time, so that when the system is upgraded or different devices are used, the interface can automatically adapt to the corresponding protocol version.
[0086] The implementation of the global mode is elaborated in detail below. The implementation of the global mode mainly involves cache coherence management of the host node, response of the compute node, etc.
[0087] Specifically, the cache coherence management of the host node can be broken down into cache directory maintenance and coherence instruction distribution.
[0088] Cache directory maintenance mainly involves that the host node can use the CXL Host Bridge to collect cache status information from each computing node and maintain a global cache directory as described in the previous text. The global cache directory records the status of each data block in the caches of each node (such as valid, invalid, dirty, etc.). For example, the host node can obtain the cache status information by periodically polling or the computing node can actively report it, and store it in the host memory of the host node.
[0089] Consistency instruction distribution mainly involves that when the host node detects a change in the status of a certain data block, it can send a consistency instruction to the corresponding near-memory and near-storage computing nodes through the CXL Host Bridge interface. For example, if the host node finds that a certain data block is modified to a dirty state in the cache of the near-memory computing node, it will send an invalidation instruction to other nodes holding copies of the data block to ensure data consistency.
[0090] The responses of the computing nodes can be divided into status reporting, instruction execution, etc.
[0091] Status reporting means that the computing node regularly reports its cache status information to the host node through the CXL Device EP. For example, each computing node can maintain a cache status table locally, and when the cache status changes, it will send the change information to the host node through the corresponding CXL interface in a timely manner.
[0092] Instruction execution means that when the computing node receives a consistency instruction sent by the host node, it can receive the instruction through the CXL Device EP and update the local cache status according to the requirements of the received instruction. For example, if an invalidation instruction is received, the computing node will mark the corresponding data block as invalid and remove the data block from the local cache.
[0093] The implementation of the local mode is elaborated in detail below. The implementation of the local mode mainly involves symmetric communication between nodes and distributed cache consistency maintenance, etc.
[0094] Specifically, symmetric communication between nodes can involve direct data transmission and message passing, etc.
[0095] Direct data transmission means that each node can achieve direct data transmission through the CXL interface. For example, between computing nodes of the same node type and / or between computing nodes of different node types, data can be directly exchanged without passing through the host node for transfer.
[0096] When message passing is carried out, message passing can be performed between computing nodes through the CXL interface to coordinate cache coherence. For example, when a computing node modifies a certain data block, it can send an update message to other nodes holding copies of the data block through the CXL interface to notify other nodes to update their cache status.
[0097] The design of distributed cache coherence maintenance involves aspects such as local cache management and collaborative coherence algorithms.
[0098] Among them, when performing local cache management, the computing node and the host node can be considered to be at the same level, that is, there is no master-slave relationship between them. The computing node can be responsible for managing its own local cache and maintaining the local cache directory. At the same time, the computing node can update and evict the local cache according to local access requirements and received messages. For example, when a computing node receives an update message from another node, it can check whether there is a copy of the data block in the local cache. If so, it can update the status of the data block copy.
[0099] The collaborative coherence algorithm means that the host node and the computing node can use the same addressing for the server memory and storage, that is, the same address can be used to access the same place. At the same time, a collaborative coherence algorithm (such as a directory-based coherence algorithm or a bus snooping coherence algorithm, etc.) can be adopted to ensure the cache coherence of the entire system. In other words, in this local mode of cache management, the central node for uniformly managing cache coherence can be omitted, and the server nodes can achieve coherence maintenance through mutual cooperation.
[0100] That is to say, the task scheduling system of this application is a relatively general system architecture and can work in different cache management modes, that is, it can work in the global mode or the local mode. Moreover, the hardware architectures of the task scheduling systems working in different cache management modes have relatively small differences, which can improve the generality of the task scheduling system.
[0101] Furthermore, to further improve the generality of the task scheduling system, corresponding designs can also be made for the general data structure in this embodiment.
[0102] Specifically, node information, global cache directory, and cache status can be defined. Among them, the node information is the identification information that can identify the identity information of the computing node, and can also manage the local cache and maintain the local cache directory. The global cache directory represents the global cache directory maintained by the host node in the global mode. Even in the global mode, each computing node can also maintain its own local cache directory, and the local cache directory can be regarded as a kind of local "global" information. The cache status is the possible status of the pre-defined buffer block, such as valid, invalid, dirty, etc., which is not limited here.
[0103] Specifically, the cache states exemplified above can include valid, invalid, and dirty. Three cache state constants can also be defined, namely VALID (valid), INVALID (invalid), and DIRTY (dirty), which are used to represent different cache states of cache blocks respectively. VALID indicates that the data in the cache block is valid and consistent with the data in the main memory; INVALID indicates that the data in the cache block is invalid and cannot be used; DIRTY indicates that the data in the cache block has been modified and is inconsistent with the data in the main memory.
[0104] A compute node class (Node) and a host node class (HostNode) can also be defined.
[0105] The compute node class can include node_id, local_cache, local_cache_directory, update_local_cache, invalidate_local_cache.
[0106] Among them, node_id represents the unique identifier of the node; local_cache is used to store the data blocks and their states in the local cache; local_cache_directory is used to record the nodes holding copies of each data block in the local cache; update_local_cache is used to update the state of the local cache and update the local cache directory at the same time; invalidate_local_cache is used to mark a certain data block in the local cache as invalid and update the local cache directory.
[0107] The host node class can inherit from Node and is applicable to the cache management mode in the global mode. The host node class can include global_cache_directory, collect_cache_status, detect_status_change_and_send_instructions.
[0108] The global cache directory can be denoted as global_cache_directory and is used to record the cache states of each data block on each node.
[0109] Optionally, the host node class can be a nested dictionary used to maintain the global cache directory. The keys of the outer dictionary are the identifiers of the data blocks (such as block1), and the values are another dictionary. The keys of the inner dictionary are the node objects, and the values are the cache states of the corresponding data blocks on these nodes. The specific structure can be exemplified by the following code:
[0110]
[0111] collect_cache_status can collect the cache status information of each node to the global cache directory. detect_status_change_and_send_instructions can detect changes in the data block status. If there is a data block that becomes dirty, it sends invalidation instructions to other nodes holding valid copies of the data block.
[0112] In this way, through the data structure design of the computing node class (Node) and the host node class (HostNode), it can meet the needs of each computing node to independently manage the cache in the local mode and also meet the needs of the host node to uniformly manage the cache in the global mode.
[0113] The following uses an example to elaborate on the working principle of cache management through the global mode.
[0114] Cache management in the global mode can include steps such as initialization, computing nodes updating local caches, host nodes collecting cache status, host nodes detecting status changes and distributing management instructions to adapt to status changes.
[0115] Specifically, when initializing the global mode, a HostNode object and multiple Node objects can be created. Each Node object represents a computing node. The HostNode object represents the host node and is responsible for global cache consistency management. The specific definition code can be as follows:
[0116] “host = HostNode(0)
[0117] node1 = Node(1)
[0118] node2 = Node(2)”
[0119] When a computing node updates its local cache, the computing node can update the local cache status according to its own access requirements, update and maintain the local cache directory. When a computing node updates its local cache, it can call the update_local_cache method. The specific call code can be as follows:
[0120] “node1.update_local_cache('block1', VALID)
[0121] node2.update_local_cache('block1', VALID)”
[0122] Thus, in the update_local_cache method, the computing node can update the status of the data block to the local_cache, and at the same time update the local_cache_directory to record the copy of the data block it holds.
[0123] When the host node collects the cache status, the host node can collect the cache status information of each computing node by means of periodic polling or active reporting by the computing node. It can be implemented by calling the collect_cache_status method, and the specific calling code can be as follows:
[0124] “host.collect_cache_status(node1)
[0125] host.collect_cache_status(node2)”
[0126] Thus, the collect_cache_status method can integrate the local cache status information of the computing node into the global_cache_directory of the host node.
[0127] When the host node detects a status change and distributes instructions, the host node can regularly detect the status change of the data block in the global_cache_directory. If it is found that a certain data block becomes dirty on a certain node, an invalidation instruction can be sent to other nodes that hold a valid copy of the data block. The specific code can be as follows:
[0128] “node1.update_local_cache('block1',DIRTY)
[0129] host.collect_cache_status(node1)
[0130] host.detect_status_change_and_send_instructions()”
[0131] Thus, in the detect_status_change_and_send_instructions method, the host node can traverse the global_cache_directory, find the data blocks in the dirty state, and send invalidation instructions to other nodes that hold valid copies of the data blocks in the dirty state. After receiving the instructions, other nodes can call the invalidate_local_cache method to mark the corresponding data block as invalid.
[0132] The working principle of cache management through the local mode is illustrated by the following example.
[0133] Cache management through the local mode may include steps such as initialization, a computing node accessing a data block, a computing node modifying a data block, and a computing node receiving an update message.
[0134] Specifically, when initializing the local mode, multiple Node objects can be created to enable each computing node to independently manage its internal cache. The specific creation code can be as follows:
[0135] "node1 = Node(1)
[0136] node2 = Node(2)"
[0137] When a computing node accesses a data block, if the data block is not in the local cache, the data block will be loaded into the local cache and marked as valid; if the data block is already in the local cache, it can be directly used. The specific code can be as follows:
[0138] "node1.update_local_cache('block1', VALID)
[0139] node2.update_local_cache('block1', VALID)"
[0140] When a computing node modifies a data block, the status of the data block can be marked as dirty in advance, and an update message can be sent to other nodes holding copies of the data block. The message passing process can be achieved by adding message sending logic to the update_local_cache method, and the message passing process will not be elaborated in detail here. The specific code for sending the update message can be as follows:
[0141] "node1.update_local_cache('block1', DIRTY)"
[0142] When a computing node receives an update message as another node, the computing node can check whether a copy of the data block exists in its local cache. When a copy of the data block exists in the local cache of the computing node, the status of the copy of the data block can be marked as invalid.
[0143] Taking computing node 2 receiving an update message as an example, its code logic can be as follows:
[0144] "node2.invalidate_local_cache('block1')"
[0145] Thus, in this embodiment, in the global mode, the host node can act as the core to achieve the management function. The host node can collect and process the cache status information of each computing node and distribute consistency instructions. In the local mode, each computing node can be considered to be at the same level in terms of cache management, and can maintain cache consistency through mutual communication and cooperation. The computing node can independently manage its internal cache and update the local cache status according to the received messages.
[0146] That is to say, in this embodiment, through the unified design of the hardware interface and the integration of the software control logic, the compatibility of the global mode and the local mode is achieved.
[0147] In the hardware design of the computing node, each computing node can be equipped with an interface that supports a target bus protocol such as CXL, such as the CXL Host Bridge and CXL Device EP interfaces described in the previous text. The interface that supports the target bus protocol can work normally in various cache management modes to provide a hardware foundation for data transmission and consistency maintenance. In the global mode, the above interfaces of the host node achieve cache consistency of the global memory on the host side; in the local mode, each computing node uses the above interfaces to achieve symmetric communication and data sharing between computing nodes.
[0148] At the same time, this embodiment realizes the integration of the software control logic through a general software control logic. By designing a relatively unified cache management module, the cache management module can integrate the algorithms and processes of the global mode and the local mode. When the server or host node is configured in the global mode, the cache management module follows the centralized control logic, and the host node is responsible for maintaining data consistency and distributing instructions to the slave nodes; when the server or host node is configured in the local mode, the cache management module switches to the distributed control logic, and each computing node can independently perform data consistency maintenance and sharing.
[0149] Furthermore, in response to the task scheduling system in this embodiment being able to adapt to various cache management modes, in this application, the cache management mode can be further flexibly switched according to the configuration or runtime requirements. That is to say, it can be switched between multiple cache management modes according to the system state of the task scheduling system.
[0150] Specifically, the configuration switching of the cache management mode can include, in the system initialization stage, setting the currently used cache management mode through a configuration file or system parameters. The global mode or the local mode can be selected as the default current cache management mode according to the characteristics and requirements of the application scenario.
[0151] Generally speaking, based on the system flexibility requirements and data consistency requirements of the server application scenario, one of the global mode and the local mode can be selected as the current cache mode.
[0152] In response to the high requirement for data consistency and relatively low network overhead in the application scenario, the global mode can be selected as the current cache management mode. In response to the application scenario that pays more attention to the flexibility and scalability of the system and has a slightly lower real-time requirement for consistency, the local mode can be selected as the current cache management mode.
[0153] Furthermore, in this embodiment, it is possible to support switching the current cache management mode during the operation of the task scheduling system. That is, in response to the demand to switch the cache management mode as the current cache management mode, the configuration file can be modified and the relevant services can be restarted, so that the system loads the new configuration and switches to the specified cache management mode.
[0154] Specifically, during the initialization of the server and / or the task scheduling system, according to the characteristics and requirements of the application scenario, the configuration file can be edited or the system parameters can be set before the system starts to specify the cache consistency management mode, in the form of configuration file or parameter settings.
[0155] Taking the configuration file as config.yaml as an example, the specific configuration code can be as follows:
[0156] "yaml
[0157] cache_consistency_mode: "global_asymmetric" # or "global_symmetric"
[0158] Among them, global_asymmetric represents the global mode; global_symmetric represents the local mode.
[0159] When the server or the task scheduling system starts, it can read the configuration file or parameters and initialize according to the specified cache management mode to realize the system loading the configuration. The specific loading process code can be as follows:
[0160]
[0161] Among them, the text after "#" represents the code comment.
[0162] When it is necessary to switch the cache management mode during the operation of the server and the task scheduling system, it may be necessary to modify the mode setting in the configuration file.
[0163] For example, change cache_consistency_mode from "global_asymmetric" to "global_symmetric".
[0164] In response to modifying the configuration file, the relevant system services can be restarted to make the new configuration take effect. That is, after the service is restarted, the server and the task scheduling system can re-read the configuration file and initialize it according to the new current cache management mode.
[0165] Optionally, during the operation of the task scheduling system, by monitoring system performance metrics and task execution status, it can be determined whether there is a performance bottleneck in the current cache management mode, so as to automatically determine whether it is necessary to switch the cache consistency management mode. Among them, the system performance metrics can include at least one of cache hit rate, data transmission latency, and network bandwidth utilization.
[0166] In response to determining whether there is a performance bottleneck in the current cache management mode, select another cache management mode as the new current cache management mode. Among them, another cache management mode refers to a cache management mode that has not been used as the current cache management mode.
[0167] For example, as described above, the cache management mode can include a global mode and a local mode. When the global mode is the current cache management mode, if the host node load is too high (for example, the current load exceeds the load threshold) and causes a delay in cache consistency maintenance, it can be dynamically switched to the local mode as the current cache management mode, so as to achieve load dispersion and improve the overall performance of the task scheduling system and the server. That is to say, the above adaptive switching mechanism can optimize cache consistency management according to the real-time state of the task scheduling system, and thus can help improve the stability and efficiency of the system.
[0168] The following gives an example to elaborate on the detailed principle of adaptive switching.
[0169] The performance metrics of the task scheduling system can be monitored. For example, performance metrics such as cache hit rate, data transmission latency, and network bandwidth utilization can be continuously monitored. Monitoring tools can be used or monitoring logic can be added to the code. The following gives an example of the monitoring logic in the code:
[0170]
[0171]
[0172] In this way, according to the monitored performance metrics and the preset performance thresholds, it can be determined whether it is necessary to switch the cache management mode. For example, in the global mode, when the host node load is too high, resulting in the data transmission latency exceeding the threshold and the network bandwidth utilization being too high, switch to the local mode. In this process, the performance metrics can include the current load, data transmission latency, and network bandwidth; correspondingly, the performance thresholds can include the load threshold, latency threshold, and bandwidth threshold, which will not be elaborated here.
[0173] When it is determined that the current cache management mode needs to be switched, the task scheduling system can dynamically switch to a new cache management mode as the current cache management mode. In this process, steps such as reconfiguring nodes and adjusting communication protocols can be carried out. Specifically, in the control code, the mode switch can be achieved by modifying global variables and reinitializing related objects. The specific code can be as follows:
[0174]
[0175]
[0176] Thus, in this embodiment, the cache coherence management mode can be flexibly switched according to different computing requirements and system states, thereby significantly improving the performance and stability of the task scheduling system and the server.
[0177] For the description of the features in the corresponding embodiments of the task scheduling system, reference can be made to the relevant descriptions of the corresponding embodiments of the task scheduling method in the following text, and details will not be repeated here.
[0178] The embodiment of the present application also provides a task scheduling method. Combining with the execution process of the task scheduling method, the detailed working principle of the task scheduling method is illustrated by examples below.
[0179] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of an embodiment of the task scheduling method of the present application.
[0180] S101: The host node of the server obtains the node information of the computing node, identifies the node type represented by the node information, and assigns a type label matching the node type to the computing node; wherein, the node type includes at least a first type and a second type, the computing efficiency of the first type is higher than that of the second type, and the data processing efficiency of the second type is higher than that of the first type.
[0181] In this embodiment, the host node is responsible for the overall control, task scheduling, and general computing tasks of the task scheduling system. The computing node is the node that specifically executes the task to be executed. This embodiment is compatible with computing nodes of multiple node types.
[0182] Among them, the node type includes at least a first type and a second type, the computing efficiency of the first type is higher than that of the second type, and the data processing efficiency of the second type is higher than that of the first type. Generally speaking, the first type of computing node is more proficient in tasks related to computing, and the second type of computing node is more proficient in processing tasks related to data.
[0183] Thus, after the server is powered on / started, the host node can identify the node types of each computing node included in the server, and assign type labels matching their node types to each computing node, which is conducive to improving the acquisition efficiency of computing nodes of the corresponding node type when performing task scheduling for tasks to be executed subsequently.
[0184] S102: The host node obtains a task to be executed and evaluates the task type of the task to be executed.
[0185] In this embodiment, the task to be executed is a task that needs to be scheduled and executed by the host node and the computing nodes. Since this embodiment is compatible with multiple types of computing nodes, the task type of the task to be executed can be evaluated to select a suitable computing node based on the task type of the task to be executed, which is conducive to improving the processing efficiency of the task to be executed.
[0186] S103: The host node uses the node type matching the task type as the target type, selects the computing node carrying the target type label as the target node, and schedules the task to be executed to the target node so that the target node executes the task to be executed.
[0187] In this embodiment, in response to obtaining the task type of the task to be executed, the host node can use the node type suitable for executing the task to be executed as the target type, that is, use the node type matching the task type as the target node. As previously identified and labeled the node types of the computing nodes, the computing node carrying the target type label can be selected as the target node, and the task to be executed is scheduled to the target node so that the target node executes the task to be executed.
[0188] That is to say, as described above, the task scheduling system can be compatible with computing nodes of multiple node types, that is, at least includes the first type of computing node and the second type of computing node, and the computing efficiency and data processing efficiency of the first type and the second type are different. Thus, in this embodiment, the host node can identify and label the node types of the computing nodes. When obtaining a task to be executed, the host node can evaluate the node type suitable for the task to be executed, select the computing node suitable for the task to be executed as the target node, and schedule the task to be executed to the target node for execution. Therefore, the technical problem of poor task execution efficiency can be solved, the technical effect of being compatible with multiple computing nodes to reasonably schedule the computing nodes for executing the task to be executed, and then improving the task execution efficiency can be achieved.
[0189] Optionally, the node information may further include the first link length between the computing node and the storage device, and the second link length between the computing node and the host memory.
[0190] When identifying the node type of a computing node, the host node can obtain the link topologies of the computing node, storage devices, and host memory to obtain and compare the first link length and the second link length of the computing node.
[0191] In response to the first link length being longer than the second link length, the host node determines that the computing node belongs to the first type; in response to the first link length being shorter than the second link length, the host node determines that the computing node belongs to the second type.
[0192] In this way, the node type of the computing node can be further verified, identified, and assisted in determination at the software level, which is beneficial to improving the reliability of node type identification.
[0193] Please refer to Figure 5 , Figure 5 which is a schematic flowchart of an embodiment for the host node to identify the node type of the computing node in this application.
[0194] S201: The host node accesses the base address register of the computing node through the target bus protocol.
[0195] In this embodiment, the target bus protocol represents the communication protocol between the host node and the computing node, that is, both the host node and the computing node support the target bus protocol. Among them, the target bus protocol can be the PCIe protocol, CXL protocol, etc. exemplified in the previous text.
[0196] The base address register (BAR) represents a data register that can be used to store operands or intermediate results to reduce the number of accesses to the memory.
[0197] In this embodiment, the base address register supports the target bus protocol, and an encoding capable of representing the node type of the computing node can be pre-assigned in the base address register of the computing node. After the server is powered on and started, the host node can access the base address register of the computing node through the target bus protocol to obtain the relevant encoded data in the base memory as the node information for evaluating the node type.
[0198] That is to say, in this embodiment, through the hardware-level design, the identification of the node type of the computing node is realized, which is beneficial to reducing the software operation burden of the host node, saving computing power resources, and improving the node type determination efficiency.
[0199] S202: The host node reads the type encoding in the base address register of the computing node.
[0200] In this embodiment, the host node can read the initial value of the reserved type bit in the base address register and use the initial value as the type encoding.
[0201] Specifically, an initial value of the BAR can be reserved for each computing node. For example, the lower 8 bits (bit[7:0]) of the BAR can be used as a type address for storing a type code.
[0202] S203: The host node evaluates the node type of the computing node based on the type code.
[0203] In this embodiment, when the host node determines that the node type of the computing node is the first type, step S204 is executed; when the host node determines that the node type of the computing node is the first type, step S205 is executed.
[0204] The host node can obtain the type code of the computing node and use the type code as node information.
[0205] As exemplified above, the node type of the computing node can include a first type and a second type. Correspondingly, the type code can include a first code and a second code, and the first code is different from the second code.
[0206] In response to the type code being the first code, the host node determines that the computing node belongs to the first type. In response to the type code being the second code, the host node determines that the computing node belongs to the second type.
[0207] For example, when 8’b00000000 (i.e., the type code in the BAR is 00000000) represents a computing node of the first type; 8’b11111111 (i.e., the type code in the BAR is 11111111) represents a computing node of the second type, and other values can be reserved for future expansion device identification.
[0208] S204: The host node assigns a type label of the first type to the computing node.
[0209] In this embodiment, in response to the host node determining that the node type of the computing node is the first type, a type label of the first type is assigned to the computing node.
[0210] S205: The host node assigns a type label of the second type to the computing node.
[0211] In this embodiment, in response to the host node determining that the node type of the computing node is the second type, a type label of the second type is assigned to the computing node.
[0212] That is to say, in this embodiment, the host node can identify the node type and assign a type label matching the node type to the computing node.
[0213] As illustrated by the examples in the foregoing text, the computing nodes of the first type are the near-memory computing nodes among the first type and the second type; the computing nodes of the second type are the near-storage computing nodes among the first type and the second type. Thus, generally speaking, the computing nodes of the second type can be labeled with "near-storage", and the computing nodes of the first type can be labeled with "near-memory".
[0214] S206: The host node registers the computing nodes with the system resource manager.
[0215] In this embodiment, in response to assigning a type label matching the node type to the computing nodes, the host node can register the computing nodes with the system resource manager. Specifically, the device information can be registered in the system resource manager for use by the task scheduler. Among them, the task scheduler is a component module of the host node for implementing the task scheduling method.
[0216] Please refer to Figure 6 , Figure 6 which is a schematic flowchart of another embodiment of the task scheduling method of this application.
[0217] S301: The host node obtains the task to be executed.
[0218] In this embodiment, the host node obtains the task to be scheduled and executed.
[0219] S302: The host node divides the task to be executed into multiple functional subtasks.
[0220] In this embodiment, the host node can divide the task to be executed according to the task function to form multiple functional subtasks.
[0221] For example, when the task to be executed is a database query task, the task to be executed can be divided into functional subtasks such as data preprocessing, database page table connection, data aggregation calculation, and data saving. Thus, according to the dependency relationship between the functional subtasks, the functional subtasks can be controlled to execute in parallel or dependently, thereby improving the execution efficiency of the task to be executed.
[0222] S303: The host node divides the functional subtasks into multiple data subtasks.
[0223] In this embodiment, the host node can further divide at least some of the functional subtasks, and divide one functional subtask into multiple data subtasks. Among them, the multiple data subtasks match in task function, and there are differences in the task processing data.
[0224] In this way, multiple data subtasks of the same functional subtask can be executed in parallel, further improving the execution efficiency of the functional subtask, and thus significantly improving the execution efficiency of the task to be executed. That is to say, in this embodiment, the task to be executed can be segmented at least at two levels, so as to further improve the execution efficiency of a single functional subtask while improving the task execution efficiency based on the task dependency relationship, and make full use of the computing resources of the task scheduling system.
[0225] Optionally, the data requirement factor of the functional subtask can be evaluated. In response to the data requirement factor not reaching the data threshold, it is considered that the amount of data to be processed by the functional subtask is small, and a single computing node can be used for execution, reducing unnecessary consumption of horizontal computing resources while ensuring data consistency. In response to the data requirement factor reaching the data threshold, it is considered that the data to be processed by the functional subtask is large, and multiple computing nodes can be used for distributed execution to improve the execution efficiency of the functional subtask. The data requirement factor will be elaborated in detail later and will not be elaborated here.
[0226] Furthermore, it can also be determined whether the functional subtask is suitable for data segmentation. When it is not suitable for data segmentation, the functional subtask is not segmented into multiple data subtasks.
[0227] In an alternative embodiment, the host node may also not perform data-side partitioning on the functional subtask and use the functional subtask as the smallest task execution unit, which is not limited here.
[0228] S304: The host node evaluates the computing power requirement factor of the task to be executed.
[0229] In this embodiment, the computing power requirement of the task to be executed can be quantitatively analyzed to confirm the node type of the computing node applicable to the task to be executed.
[0230] It should be noted that in this implementation, the task to be executed or the functional subtask or the data subtask can be used as the evaluation subject to evaluate the computing power requirement factor.
[0231] Optionally, the task to be executed can be used as the evaluation subject to reduce the amount of evaluation; or, the data subtask can be used as the evaluation subject to improve the evaluation fineness; or, the functional subtask can be used as the evaluation subject to balance the evaluation amount and the evaluation fineness.
[0232] The computing power requirement factor will be elaborated in detail later and will not be elaborated here.
[0233] S305: The host node evaluates the data requirement factor of the task to be executed.
[0234] In this embodiment, the data processing requirements of the task to be executed can be quantitatively analyzed to confirm the node type of the computing node applicable to the task to be executed.
[0235] As described above, in this embodiment, the task to be executed or the functional subtask or the data subtask can be used as the evaluation subject to evaluate the data requirement factor.
[0236] Optionally, the task to be executed can be used as the evaluation subject to reduce the amount of evaluation; or, the data subtask can be used as the evaluation subject to improve the evaluation fineness; or, the functional subtask can be used as the evaluation subject to balance the amount of evaluation and the evaluation fineness.
[0237] The computing power requirement factor will be elaborated in detail later and will not be repeated here.
[0238] The data requirement factor will be elaborated in detail later and will not be repeated here.
[0239] S306: The host node evaluates the task type of the task to be executed.
[0240] In this embodiment, the host node can comprehensively evaluate the task type of the task to be executed by combining the computing power requirement factor and the data requirement factor. Among them, the task type includes computing power type, memory type, and transmission type.
[0241] S307: Select the computing node carrying the first type of label as the target node.
[0242] In this embodiment, in response to the task type of the task to be executed being the computing type, the host node can use the first type as the target type and select the computing node carrying the first type of label as the target node.
[0243] Furthermore, the host node can preferentially select the computing node that includes a graphics processor and carries the first type of label as the target node.
[0244] Preferential selection means that if there is no suitable computing node of the target type currently, a computing node of other node types can be selected as the target node, which is beneficial to reducing the queuing time of the task to be executed, thereby improving the execution efficiency of the task to be executed and further improving the performance of the task scheduling system. Preferential selection can express a similar meaning later and will not be repeated later.
[0245] S308: Select the computing node carrying the first type of label as the target node.
[0246] In this embodiment, in response to the task type of the task to be executed being the memory type, the host node can use the first type as the target type and select the computing node carrying the first type of label as the target node.
[0247] Furthermore, the host node may preferentially select a computing node carrying a first type of label and including a central processing unit and a programmable controller as the target node.
[0248] S309: Select a computing node carrying a second type of label as the target node.
[0249] In this embodiment, in response to the task type of the task to be executed being a transmission type, the host node may use the second type as the target type and select a computing node carrying a first type of label as the target node.
[0250] Furthermore, the host node may preferentially select a computing node carrying a second type of label and including a central processing unit and a programmable controller as the target node.
[0251] S310: The host node schedules the task to be executed to the target node so that the target node executes the task to be executed.
[0252] In this embodiment, in response to selecting a computing node carrying a target type of label as the target node, the host node may schedule the task to be executed to the target node. After obtaining the task to be executed, the target node may execute the task to be executed.
[0253] Furthermore, as described above, the task to be executed can be split into multiple levels to form multiple data subtasks, and multiple computing nodes can be selected as target nodes to schedule the data subtasks to multiple target nodes, and the multiple target nodes cooperate to execute the task to be executed.
[0254] The data requirement factor and the computing power requirement factor are elaborated in detail below and will not be elaborated here.
[0255] The code of the task to be executed can be statically analyzed based on a code analysis tool to construct an abstract syntax tree (AST) of the code. The code structure can be obtained by analyzing the AST. Among them, the code structure may include the positions and nesting relationships of function definitions, loop statements, conditional statements, etc.
[0256] The function call count, loop count, and various execution ratios are used as statistical key indicators, and each key indicator is obtained.
[0257] The code analysis tool can traverse the AST, count the number of times each function is called, and understand the usage frequency of functions in the code to obtain the function call count.
[0258] Identify loop structures in the code (such as for and while loops) and calculate the number of iterations of the loop. For loops with a fixed number of iterations, the loop count can be directly obtained; for loops with a dynamic number of iterations, take the average value = (maximum loop count + minimum loop count) / 2 for approximate estimation to obtain the loop count.
[0259] Classify the instructions in the code into different types such as computational instructions (e.g., arithmetic operations, logical operations) and memory access instructions (e.g., read and write memory operations), and count the proportion of each type of instruction in the code to obtain the proportion of each type of instruction.
[0260] The key metrics obtained can be used as computing power demand factors and data demand factors, and the key metrics can be analyzed to determine the task type of the task to be executed.
[0261] As described in the previous text, the task types of the tasks to be executed can include computing power types, memory types, and transmission types.
[0262] The computing power type can be considered equivalent to a compute-intensive task. If the loop body in the code to be executed contains a large number of computational instructions and the computational instructions account for a relatively high proportion in the total number of instructions (for example, the set threshold exceeds 60%), then it can be determined that the task to be executed is a compute-intensive task. For example, in a scientific computing program, there are a large number of matrix operations, and the loop contains a large number of multiplication and addition operations, which can be determined to meet the characteristics of a compute-intensive task.
[0263] The memory type can be considered equivalent to a memory-intensive task. If the function frequently accesses memory data structures, the amount of data accessed is large, and the memory access instructions account for a relatively high proportion in the total number of instructions (the set threshold exceeds 50%), then the task to be executed may be a memory-intensive task. For example, in a big data processing program, the function continuously reads and writes large arrays or linked lists.
[0264] The transmission type can be considered equivalent to an I / O (Input / Output) intensive task. If the function frequently accesses storage devices, the code contains a large number of I / O operation statements such as file reading and writing, network requests, etc., and these operations are frequently called during the execution of the task to be executed, the execution time accounts for a relatively large proportion, and the computational operations are relatively few (the set threshold exceeds 50%), then it can be determined that the task to be executed is an I / O intensive task. For example, in a data processing program, there are a large number of open() functions (window opening functions) for file reading and writing, or request libraries for network requests, and the execution time exceeds 50%. In this case, the task belongs to the I / O intensive type.
[0265] Further, the data subtask can be used as the evaluation subject to evaluate the computing power demand factor and the data demand factor. Thus, in this embodiment, the task type of the data subtask can also be represented by a fixed-length variable.
[0266] For example, it can be represented by an 8-bit fixed-length variable. The upper 2 bits represent the subtask type, 2'b00 represents the functional subtask, 2'b01 represents the data subtask, and 2'b11 represents unrecognized. The lower 6 bits represent the characteristic category, 6'b000000 represents unrecognized, 6'b000001 represents the computing power type, 6'b000010 represents the memory type, and 6'b000011 represents the transmission type.
[0267] In this embodiment, further, before selecting the computing node carrying the target type label as the target node, the current state of each computing node can also be scored to select the computing node with a relatively high score as the target node.
[0268] Specifically, the host node evaluates the performance score of the computing node. The performance score is used to represent the computing power performance of the performance modules included in the computing node. The types of performance modules include graphics processors, central processors, and programmable controllers. The host node evaluates the memory score and the transmission score of the computing node. The host node fits the performance score, the memory score, and the transmission score of the computing node to obtain the comprehensive score of the computing node. The host node selects the computing node based on the performance score and / or the comprehensive score.
[0269] That is to say, in this embodiment, resource evaluation of the computing node can be performed to quantitatively evaluate the resource status of the computing node such as computing power, memory capacity, transmission capacity, and comprehensive capacity.
[0270] System performance tools and task queue monitoring tools, etc., can be used to collect the system state data of the computing node in real time. The system state data can include but is not limited to CPU utilization, memory bandwidth, and I / O bandwidth.
[0271] CPU utilization is used to reflect the workload of the computing unit. Memory bandwidth is used to reflect the rate of memory data transmission. The available memory space is used to represent the amount of data that the computing node can carry. I / O bandwidth is used to represent the data transmission efficiency of input / output operations.
[0272] Based on the collected system state data, evaluation scores of the available capabilities are generated for the computing node, namely, the performance score, the memory score, the transmission score, and the comprehensive score.
[0273] The evaluation score can be used as a key indicator to measure the availability and performance level of node resources, and it can relatively intuitively present the real-time state of computing nodes. At the same time, considering that the load of the task scheduling system is in dynamic change, the evaluation score can be dynamically updated according to the real-time fluctuation of the load. By continuously adjusting the evaluation score, the continuous tracking of the system resource state can be realized. During the task execution process, the execution paths of data subtasks / functional subtasks / tasks to be executed are dynamically adjusted according to the real-time updated evaluation score. Data subtasks / functional subtasks / tasks to be executed can be preferentially assigned to nodes with high evaluation scores, abundant resources, and excellent performance, so as to effectively reduce the risk of resource idleness or over-concentration.
[0274] The specific method for evaluating node resources can be described as follows. The scores can be calculated separately from three dimensions: computing power, memory capacity, and I / O capacity, and the comprehensive score is calculated by combining weights. The score for each dimension is based on specific performance indicators and is transformed into a unified score range through standardization processing. Taking the score range of 0-100 as an example, the specific calculation formula is given below.
[0275] The performance score is equivalent to evaluating the computing power of the computing node and can focus on the computing performance of the computing node. The specific evaluation formula can be as follows:
[0276] CPUscore = (Actual CPU value - Minimum CPU value) / (Maximum CPU value - Minimum CPU value) × 100
[0277] GPUscore = (Actual GPU value - Minimum GPU value) / (Maximum GPU value - Minimum GPU value) × 100
[0278] Among them, CPUscore represents the CPU score; the actual CPU value represents the measured performance value of the current CPU of the task scheduling system; the minimum CPU value represents the calibrated minimum performance value of the CPU; the maximum CPU value represents the calibrated maximum performance value of the CPU. GPUscore represents the GPU score; the actual GPU value represents the measured performance value of the current GPU of the task scheduling system; the minimum GPU value represents the calibrated minimum performance value of the GPU; the maximum GPU value represents the calibrated maximum performance value of the GPU.
[0279] FPGAscore = (Number of used logic units / Total number of logic units) * 100
[0280] Among them, FPGAscore represents the performance score of the FPGA.
[0281] For example, the current performance modes of the computing node include CPU and FPGA, and the calculation formula for the performance score can be as follows:
[0282] Performance score = i * CPU score + j * FPGA score
[0283] Among them, Performance score represents the performance score; i and j represent the weight factors of each performance module, and the sum of the two can be 1. That is, the scores of each performance module of the computing node can be weighted and fused to obtain the performance score of the computing node.
[0284] The memory score can be obtained through the embedded processor to get the total physical memory and available physical memory size of the computing node, and use kernel tools to test the read and write bandwidth of the computing node's memory. The specific evaluation formula can be as follows:
[0285] Memory score = k * available physical memory / total physical memory + l * actual memory bandwidth / reference memory bandwidth
[0286] Among them, Memory score represents the memory score; k and l are weight coefficients, and k + l = 1.
[0287] The transmission score can be obtained through the embedded processor to get the disk I / O performance of the computing node, test the disk read and write performance to obtain the disk read and write speed. At the same time, read the disk size and available disk space. The specific evaluation formula can be as follows:
[0288] IO score = q * disk read and write speed / reference disk read and write speed + r * available disk space / disk size;
[0289] Among them, IO score represents the transmission score; q and r are weight coefficients, and q + r = 1.
[0290] The performance score, memory score, and transmission score can be weighted and fused. Specifically, weights can be assigned to each dimension according to the requirements of the application scenario, and the weights can be adjusted according to specific needs.
[0291] For example, when the task to be executed is the preprocessing of database data tables, 40% of the weight can be assigned to the computing power, 30% of the weight can be assigned to the memory capacity, and 30% of the weight can be assigned to the I / O capacity. The specific evaluation formula for the comprehensive score can be: Comprehensive score = (Performance score * 0.4) + (Memory score * 0.3) + (IO score * 0.3).
[0292] That is to say, if the task type of the task to be executed is a transmission type, the weight of the I / O capacity can be increased; if the task type of the task to be executed is a computing power type, the weight of the computing power can be increased, etc.
[0293] That is to say, during the execution of the task to be executed, the task execution path can be planned according to the task characteristics and resource status.
[0294] During the process of planning the task execution path, the input information may include array task information, including the priority, waiting duration, required memory space, data processing volume, sub-task type, sub-task characteristic category, and dependent task array number of the sub-tasks; various node resource evaluation information, including computing power score, memory capacity score, I / O capacity score, and comprehensive capacity score. Among them, the sub-tasks can represent functional sub-tasks or data sub-tasks.
[0295] Based on the input information, the execution path of the task to be executed is planned. The output information in response to the completion of the planning may include the sub-task execution path, specifically including the execution order and data transmission path of the sub-tasks between the first type and the second type of computing nodes.
[0296] Generally speaking, during task scheduling in the present application, the tasks can be sorted according to the task priorities, and the execution of high-priority tasks is preferentially satisfied. If other tasks need to be depended on, the dependent tasks are preferentially executed to ensure that the key tasks are completed in time. Then, during the scheduling of tasks with the same priority, the tasks with a longer waiting time are preferentially allocated. If other tasks need to be depended on, the dependent tasks are preferentially executed.
[0297] After confirming the sub-tasks to be executed, the acceleration computing nodes are scheduled according to the task type.
[0298] If the task type is a compute-intensive task, it is preferentially scheduled to a near-memory GPU computing node with a high computing resource score and a low load. The score is updated according to the GPU computing performance scoring rules, and the computing node with the highest score is selected. This can ensure that the task can be executed quickly. For example, for tasks such as regular expressions and sorting in a database.
[0299] If the task type is memory-intensive, it is allocated to a node with a high memory resource score, such as a near-memory node or a near-storage computing node equipped with a large amount of memory, to ensure that the task has sufficient memory space to process data. For example, for data caching and preprocessing tasks in big data analysis, they are preferentially arranged to nodes with sufficient memory.
[0300] If the task type is an I / O-intensive task, it is allocated to a near-storage computing node with good storage I / O performance, taking advantage of its proximity to the storage device to reduce data transmission latency. For example, for database query tasks, they are scheduled to a near-storage computing node with high-speed SSD storage and sufficient I / O bandwidth. For memory or I / O-intensive tasks, the CPU and FPGA nodes are preferred. The evaluation scores are calculated according to the rules, the required weight ratios are set, and the comprehensive scores are obtained. The node with the highest score is selected.
[0301] In summary, this application utilizes the cache coherence, memory semantics, and low-latency communication features of the target bus protocol to provide a task scheduling system architecture for computing nodes that is compatible with multiple node types, and also provides a corresponding task scheduling method. HIA also covers different cache management modes that support the coherence of the target bus protocol. For example, by integrating near-storage computing and near-memory computing, the synergistic advantages of both can be achieved. On the premise of ensuring support for heterogeneous cache coherence, computing tasks can be efficiently offloaded to acceleration devices near storage and memory, enabling tasks to be executed nearby.
[0302] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0303] An embodiment of this application also provides an electronic device.
[0304] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of an embodiment of the electronic device of this application.
[0305] In one embodiment, the electronic device may include a memory and a processor.
[0306] The memory stores a computer program.
[0307] The processor is configured to run the computer program to execute the steps in any of the above-described embodiments of the task scheduling method.
[0308] An embodiment of this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, wherein the computer program is configured to execute the steps in any of the above-described embodiments of the task scheduling method when running.
[0309] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROMs), random access memories (RAMs), external hard drives, magnetic disks, or optical discs, etc., all of which can store computer programs.
[0310] An embodiment of this application also provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the task scheduling method.
[0311] An embodiment of the present application also provides another computer program product. The computer program product includes a non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in any of the above-described embodiments of the task scheduling method are implemented.
[0312] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0313] The above has introduced in detail a task scheduling method, a task scheduling system, an electronic device, and a computer-readable storage medium provided by the present application. Specific examples are used herein to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A task scheduling method, characterized in that, The task scheduling method includes: The host node of the server obtains the node information of the computing nodes, identifies the node types represented by the node information, and assigns a type label matching the node type to the computing nodes; wherein, the node types at least include a first type and a second type, the computing efficiency of the first type is higher than that of the second type, and the data processing efficiency of the second type is higher than that of the first type; the computing nodes of the first type are closer to the host memory than the computing nodes of the second type, and their first link length is longer than their second link length; the first link length represents the link length between the computing node and the storage device, and the second link length represents the link length between the computing node and the host memory; the computing nodes of the second type are closer to the storage device than the computing nodes of the first type, and their first link length is shorter than their second link length; The host node obtains the task to be executed and evaluates the task type of the task to be executed; The host node uses the node type matching the task type as the target type, selects the computing nodes carrying the target type label as the target nodes, and schedules the task to be executed to the target nodes so that the target nodes execute the task to be executed.
2. The task scheduling method according to claim 1, wherein, The node information includes a type code; the identifying the node types represented by the node information includes: The host node obtains the type code of the computing node; In response to the type code being the first code, the host node determines that the computing node belongs to the first type; In response to the type code being the second code, the host node determines that the computing node belongs to the second type.
3. The task scheduling method according to claim 2, wherein The host node obtaining the node information of the computing nodes connected thereto includes: The host node accesses the base address register of the computing node through the target bus protocol; The host node reads the initial value of the reserved type bit in the base address register and uses the initial value as the type code.
4. The task scheduling method according to claim 1, wherein The evaluating the task type of the task to be executed includes: The host node evaluates the computing power demand factor and the data demand factor of the task to be executed; The host node comprehensively evaluates the computing power demand factor and the data demand factor to evaluate the task type of the task to be executed; wherein, the task types include a computing type, a memory type, and a transmission type.
5. The task scheduling method according to claim 4, wherein The host node using the node type matching the task type as the target type and selecting the computing nodes carrying the target type label as the target nodes includes: In response to the task type of the task to be executed being a computing type, the host node uses the first type as the target type; preferentially selects the computing nodes including a graphics processor and carrying the first type label as the target nodes; and / or, In response to the task type of the task to be executed being a memory type, the host node uses the first type as the target type; preferentially selects the computing nodes carrying the first type label and including a central processing unit and a programmable controller as the target nodes; and / or, In response to the task type of the to-be-executed task being a transmission type, the host node takes the second type as the target type; and preferentially selects a computing node carrying a second type tag and including a central processing unit and a programmable controller as the target node.
6. The task scheduling method according to claim 1, wherein Scheduling the to-be-executed task to the target node includes: The host node divides the to-be-executed task according to task functions to form multiple functional subtasks; and divides the functional subtasks to form multiple data subtasks; wherein, the task functions of the multiple data subtasks match, and the task processing data is different. The host node schedules the data subtasks to multiple target nodes.
7. The task scheduling method according to claim 1 or 6, characterized in that, Before selecting a computing node carrying a target type tag as the target node, it includes: The host node evaluates the performance score of the computing node; wherein, the performance score is used to represent the computing power performance of the performance modules included in the computing node, and the types of performance modules include a graphics processing unit, a central processing unit, and a programmable controller. The host node evaluates the memory score and transmission score of the computing node. The host node fits the performance score, memory score, and transmission score of the computing node to obtain the comprehensive score of the computing node. The host node selects the computing node based on the performance score and / or the comprehensive score.
8. The task scheduling method according to claim 1, wherein The task scheduling method includes: The server selects the current cache management mode from multiple cache management modes. Among them, the cache management modes of the server at least include a global mode and a local mode; the global mode means global byte addressing of the server memory for the host node to perform global scheduling on the server memory; the server memory includes a host memory and a computing node memory; the local mode means that the computing node independently controls its computing node memory and synchronizes the data changes of the computing node memory with other computing nodes.
9. The task scheduling method according to claim 8, characterized in that Before the server selects the current cache management mode from multiple cache management modes, it includes: The computing node maintains a local cache directory; the host node maintains a global cache directory; wherein, the global cache directory includes the node identifier of the computing node and the local cache directory. The computing node monitors the cache status of the cache blocks included in the computing node; wherein, the cache status at least includes a valid state, an invalid state, and a dirty state.
10. A task scheduling system, characterized in that, The task scheduling system is used to implement the task scheduling method according to any one of claims 1 to 9. The task scheduling system includes: A host node; Computing nodes, including computing nodes of multiple node types; the node types at least include a first type and a second type, the computing efficiency of the first type is higher than that of the second type, and the data processing efficiency of the second type is higher than that of the first type. A switch module, respectively connected to the host node and the computing nodes, and the host node and the computing nodes communicate with each other through the switch module.
11. The task scheduling system according to claim 10, wherein The task scheduling system further includes a host memory and a storage device. The host memory is connected to the switch module and is connected to the host node and the first type of computing node through the switch module; The storage device is connected to the second type of computing node and is connected to the switch module and the host node through the second type of computing node.
12. The task scheduling system according to claim 10 or 11, characterized in that, The switch module includes a first module and a second module; the first module and the second module are connected, and at least one of the two is connected to the host node; The first module is connected to the host memory and the first type of computing node; The second module is connected to the second type of computing node.
13. The task scheduling system according to claim 10, wherein The constituent modules of one of the first type of computing nodes are integrated in a field programmable gate array; the constituent modules of one of the second type of computing nodes are integrated in a field programmable gate array.
14. An electronic device, characterized in that, Comprising: A memory for storing a computer program; A processor for implementing the steps of the task scheduling method according to any one of claims 1 to 9 when executing the computer program.
15. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the task scheduling method according to any one of claims 1 to 9 when executed by a processor.
Citation Information
Patent Citations
Task processing method, device, system, equipment and medium
CN114443236A