Resource scheduling method, electronic equipment, medium and product

By dynamically adjusting transmission paths and utilizing the CXL shared memory pool in heterogeneous computing systems, the problem of fragmented cross-network congestion control in multi-host and multi-GPU hybrid architectures is solved, achieving efficient resource scheduling and utilization, and improving the performance of AI training and inference tasks.

CN121029431AActive Publication Date: 2025-11-28INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202511556699.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-29
Publication Date
2025-11-28
Estimated Expiration
2045-10-29

AI Technical Summary

Technical Problem

In a multi-host and multi-GPU hybrid architecture, the lack of an effective cross-network data scheduling mechanism leads to phenomena such as idle GPU computing resources, lack of memory resources, and bandwidth congestion of switches. Cross-network congestion control is fragmented, and the allocation of video memory and computing power is uneven, resulting in low resource utilization.

Method used

By employing a heterogeneous computing system, the bandwidth utilization of switching devices is obtained through the host, a backup switching device is found based on a round-robin query strategy, the transmission path is dynamically adjusted, and data storage is performed using the CXL shared memory pool, thereby achieving cross-network traffic coordination and resource optimization.

Benefits of technology

It solves the communication bottleneck caused by the fragmentation of cross-network congestion control in AI systems, achieving high throughput, low latency and high resource utilization, and improving the efficiency of large-scale AI training and inference tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121029431A_ABST
    Figure CN121029431A_ABST
Patent Text Reader

Abstract

The invention discloses a resource scheduling method, an electronic device, a medium and a product, and relates to the technical field of AI intelligent computing, and the method comprises the steps: obtaining the bandwidth utilization rate of a main switching device in a plurality of switching devices through a host, and if the bandwidth utilization rate of the main switching device is greater than a first preset utilization rate, executing the first preset utilization rate; if yes, whether standby switching equipment meeting a preset standby condition exists in the remaining switching equipment except the main switching equipment in the multiple pieces of switching equipment or not is queried based on a first circular query strategy, and if not, a target switching equipment meeting a preset transmission condition is determined from the multiple first switching equipment, the to-be-calculated data is stored in the shared memory pool through the target switch, and low congestion, high throughput and high resource utilization rate under large-scale AI training and reasoning tasks are realized by introducing multi-level popularity perception, dynamic threshold adjustment, standby path switching and cross-network data unloading mechanisms into a hybrid interconnection architecture of the shared memory pool.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of AI (Artificial Intelligence) intelligent computing, in particular to a resource scheduling method, an electronic device, a medium and a product. BACKGROUND

[0002] With the wide application of deep learning models in natural language processing, computer vision and other fields, the demand for computing power, memory capacity and data exchange bandwidth for AI training and inference has increased dramatically. In a multi-host and multi-GPU hybrid architecture, Scale up networks and Scale out networks coexist, and the data flow path is complex. When congestion occurs in a certain interconnection subnet or resources are not fully utilized, there is a lack of effective cross-network data scheduling mechanism, which may result in idle GPU computing resources, lack of memory resources, and switch bandwidth congestion.

[0003] In related technologies, multiple focuses are on single congestion control, memory or performance improvement, which may cause cross-network congestion control fragmentation, uneven memory and computing power allocation, thereby limiting the computing power of large-scale AI systems and resource utilization efficiency, and there is an urgent need to solve this problem. SUMMARY

[0004] The present application provides a resource scheduling method, an electronic device, a medium and a product to at least solve the problems of communication bottleneck, global resource unable to be optimized cooperatively and low resource utilization rate caused by cross-network congestion control fragmentation in AI systems in related technologies.

[0005] The present application provides a resource scheduling method, which is applied to a heterogeneous computing system including a host, a first switch and a shared memory pool, the host and the shared memory pool are connected with the first switch, and the host includes multiple switching devices. The method includes the following steps: acquiring bandwidth utilization of a main switching device in the multiple switching devices through the host; if the bandwidth utilization of the main switching device is greater than a first preset utilization rate, querying whether a standby switching device meeting a preset device condition exists in the remaining switching devices except the main switching device in the multiple switching devices based on a first cyclic query strategy; if no standby switching device meeting the preset device condition exists, determining a target switch meeting a preset transmission condition from the multiple first switches, and storing to-be-computed data to the shared memory pool through the target switch.

[0006] The application further provides a resource scheduling device, which is applied to a heterogeneous computing system, the heterogeneous computing system comprising a host, a first switch and a shared memory pool, the host and the shared memory pool being connected with the first switch, the host comprising a plurality of switching devices, wherein the device comprises: an acquisition module, configured to acquire a bandwidth utilization rate of a master switching device in the plurality of switching devices through the host; a query module, configured to, if the bandwidth utilization rate of the master switching device is greater than a first preset utilization rate, query whether a standby switching device meeting a preset device condition exists in the remaining switching devices except the master switching device in the plurality of switching devices based on a first cyclic query strategy; a storage module, configured to, if the standby switching device meeting the preset device condition does not exist, determine a target switch meeting a preset transmission condition from the plurality of first switches, and store to-be-computed data to the shared memory pool through the target switch.

[0007] The application further provides an electronic device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the steps of the resource scheduling method according to any one of the above.

[0008] The application further provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the resource scheduling method according to any one of the above.

[0009] The application further provides a computer program product, wherein the computer program is executed by a processor to implement the steps of the resource scheduling method according to any one of the above.

[0010] This application obtains the bandwidth utilization rate of the main switching device among multiple switching devices through a host. If the bandwidth utilization rate of the main switching device is greater than a first preset utilization rate, a first loop query strategy is used to query whether there is a backup switching device that meets the preset backup conditions among the remaining switching devices other than the main switching device. If not, a target switch with preset transmission conditions is determined from multiple first switches, and the data to be computed is stored in a shared memory pool through the target switch. This solves the technical problems of communication bottlenecks, lack of global resource optimization, and low resource utilization caused by the fragmentation of cross-network congestion control in AI systems. By introducing multi-level heat perception, dynamic threshold adjustment, backup path switching, and cross-network data offloading mechanisms in the hybrid interconnection architecture of Scale-up+Scale-out+CXL (ComputeExpress Link) shared memory pool, low congestion, high throughput, and high resource utilization are achieved under large-scale AI training and inference tasks. Attached Figure Description

[0011] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A flowchart illustrating a resource scheduling method provided in an embodiment of this application; Figure 2 This is an overall block diagram of a system implementation scheme according to an embodiment of the present invention; Figure 3 This is a schematic diagram of a multi-path backup switch activation scheme according to an embodiment of the present invention; Figure 4 This is a schematic diagram of a dynamic collaborative scheme between video memory and computing power according to an embodiment of the present invention; Figure 5 This is a block diagram of a resource scheduling device according to an embodiment of the present invention; Figure 6 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0014] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0015] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0016] The embodiments of this application provide a heterogeneous computing system, and the system is described in detail below in conjunction with the heterogeneous computing system.

[0017] Specifically, before introducing the embodiments of this application, let's first introduce the hardware architecture used in artificial intelligence servers in related technologies. This mainly includes: a multi-GPU (Graphics Processing Unit) scale-up architecture based on PCIe (Peripheral Component Interconnect Express) or NVLink. This architecture enables direct data exchange between GPUs through high-speed interconnect buses and switching chips. Some systems support P2P (Peer-to-Peer) communication between GPUs to reduce cross-device access latency. However, to support larger-scale AI model training / inference or more accumulated context and token amounts, this architecture can only increase the number of GPUs in the network or offload data to host DRAM. This approach offers limited memory improvement and significantly increases costs, and also increases the probability of congestion on hot links, leading to increased data transmission latency. Furthermore, scale-up architectures based on RDMA (Remote Direct Memory Access) protocols... The out architecture uses the RDMA protocol to interconnect multiple host nodes, allowing one host node to directly read and write the memory of another host node without the intervention of the other's CPU or the operating system kernel, thereby reducing the latency of cross-node communication to the microsecond level.

[0018] Congestion control mechanisms typically fall into the following categories when congestion occurs: (1) Congestion control based on explicit feedback: When the switch detects that the queue length exceeds the threshold, it marks the ECN (Explicit Congestion Notification) bit in the data packet. After receiving the mark, the receiver sends feedback to the sender. The sender reduces the sending rate according to the feedback information, thereby preventing the queue from growing continuously. (2) Rate-limited flow control: The sending end dynamically adjusts the data sending rate based on the link bandwidth utilization and historical packet loss to avoid link overload; (3) Priority-based queue scheduling: RDMA flows are classified according to priority, and different priority flows occupy different queue resources in the switch to ensure that critical task flows can still obtain low latency and stable bandwidth under congestion conditions; (4) Priority-based queue scheduling: RDMA flows are classified according to priority, and different priority flows occupy different queue resources in the switch to ensure that critical task flows can still obtain low latency and stable bandwidth under congestion conditions; While the above methods can effectively alleviate link congestion in a single network environment, in a multi-GPU + multi-host hybrid architecture coexisting with PCIe networks, RDMA congestion control and PCIe congestion control are independent of each other, lacking a cross-network global traffic scheduling mechanism. This makes it difficult to achieve optimal resource utilization across the entire system, and therefore, the main problems are as follows: (1) There is a lack of comprehensive optimization mechanism for large-scale models. Most existing solutions make local improvements for single problems (such as memory expansion, single network congestion control or node performance optimization). At the same time, RDMA has high latency in small data block transmission. The overall solution lacks global cross-network optimization capability for large-scale AI model training / inference, and it is difficult to meet the requirements of low latency and high throughput at the same time. (2) The congestion control across networks is fragmented. The congestion detection and scheduling mechanisms of PCIe / NVLink networks and RDMA networks are independent of each other and lack global collaborative optimization. When a bottleneck occurs on one side of the network, the link or computing resources on the other side may be idle, resulting in a decrease in the overall computing power utilization. (3) Uneven allocation of GPU memory and computing power. In the existing architecture, the allocation strategy of GPU memory and computing power has not achieved real-time dynamic coordination, which can easily lead to some GPU memory overflow, some GPU memory idle, or GPU waiting for data transmission and running idle, affecting training and inference efficiency. (4) Uneven utilization of multiple switching chip links. In large-scale training tasks, data flow is prone to forming hot links when passing through multiple switching chips. The existing mechanism lacks dynamic traffic rerouting function based on real-time link utilization, which leads to local congestion limiting overall throughput. This reduces the transmission rate of low-priority transactions and also prevents the transmission performance of hot links from reaching its best.

[0019] Therefore, based on the above-mentioned problems, the embodiments of this application mainly solve the following: (1) Hybrid architecture optimization for large-scale AI tasks, constructing a multi-GPU + multi-host + CXL shared memory pool architecture adapted to large-scale AI tasks, replacing the RDMA protocol with the CXL shared memory pool to complete data interaction between hosts, and simultaneously ensuring high throughput, low latency and high resource utilization in scenarios such as large-scale model training, inference and multimodal processing; (2) Cross-network congestion control mechanism, establishing a global congestion detection and scheduling method that simultaneously covers scaleup and scaleout networks, realizing dynamic coordination of traffic between different network domains, and avoiding the impact of single-sided bottlenecks on global performance; (3) Dynamic collaborative scheduling of GPU memory and computing power, dynamically optimizing the allocation of GPU memory and computing resources based on GPU real-time memory utilization, computing power utilization and CXL shared memory pool availability, reducing memory overflow and computing power idleness; (4) Global optimization of multi-switching chip link bandwidth, proposing a new multi-switching chip cross-network traffic rerouting strategy based on link bandwidth utilization, alleviating hot link congestion and improving overall network throughput.

[0020] Specifically, Figure 1 This is a flowchart illustrating a resource scheduling method provided in an embodiment of the present invention. The method is applied to a heterogeneous computing system, which includes a host, a first switch, and a shared memory pool. Both the host and the shared memory pool are connected to the first switch, and the host includes multiple switching devices.

[0021] First, let's introduce the heterogeneous computing system and its architecture involved in this application.

[0022] Specifically, to address the aforementioned problems, the resource scheduling method of this application is applied to a heterogeneous computing system, such as an efficient hybrid network congestion control and resource scheduling system for large-scale AI training and inference. In the heterogeneous computing system, it mainly includes a host, a first switch, and a shared memory pool (e.g., a CXL shared memory pool). The host and the shared memory pool are both connected to the first switch. The host includes multiple switching devices. At the same time, this heterogeneous computing system is suitable for training and inference of large-scale AI models and adopts a hybrid scale-up and scale-out network architecture with multiple hosts and multiple GPUs.

[0023] Among them, such as Figure 2As shown, the first switch is a high-speed serial computer expansion bus standard switch (e.g., a CXL switch), and the switching devices are compute fast connection switches (e.g., PCIe switches). The host is configured to sequentially query the bandwidth utilization of the switching devices based on a first round-robin query strategy, starting from the main PCIe switch and querying the bandwidth utilization of the PCIe switches in a fixed order, ensuring the fairness and predictability of the query. The host includes multiple switching devices, and among these devices are multiple GPUs that communicate with at least some of the switching devices. The GPUs within each host are interconnected via a scale-up network. Each host is equipped with 8 PCIe switches, and each PCIe switch connects all 8 GPUs, thereby achieving high-speed data exchange between GPUs. Simultaneously, the multiple hosts are interconnected via a scale-out network. The shared memory pool includes multiple (e.g., 8) compute fast connection controllers (e.g., CXL controllers), meaning they share a CXL shared memory pool composed of 8 CXL controllers. Each CXL controller supports multiple remote direct memory access channels; for example, each CXL controller supports 5 CXL DRAM channels. Each host is connected to all CXL switches. In this embodiment, all CXL switches... All switches are connected to all CXL controllers, thus forming a high-bandwidth, low-latency cross-node memory access channel.

[0024] Furthermore, in this heterogeneous computing system, each GPU monitors its own memory utilization and computing power utilization in real time, while the PCIe switch and CXL switch monitor the bandwidth utilization of each port in real time, thereby forming fine-grained link load monitoring data to provide a basis for subsequent congestion detection and scheduling.

[0025] Therefore, this application constructs a multi-GPU + multi-host + CXL shared memory pool architecture adapted to large-scale AI tasks. The CXL shared memory pool replaces the RDMA protocol to complete data interaction between multiple hosts, while ensuring high throughput, low latency and high resource utilization in scenarios such as large-scale model training, inference and multimodal processing.

[0026] The resource scheduling method of this invention includes the following steps: In step S101, the bandwidth utilization rate of the main switching device among multiple switching devices is obtained through the host.

[0027] According to one embodiment of the present invention, before obtaining the bandwidth utilization rate of the main switching device among multiple switching devices through the host, the method further includes: determining whether there is a data transmission requirement; if there is a data transmission requirement, determining a target acceleration unit based on the data transmission requirement, and calculating a first transmission path between the target acceleration unit and at least some of the switching devices; and, based on the first transmission path between the target acceleration unit and at least some of the switching devices, designating the switching device with the shortest first transmission path between the target acceleration unit and the target acceleration unit among the at least some switching devices as the main switching device.

[0028] Specifically, such as Figure 2 As shown, in a scale-up network, after a transaction is transmitted from the GPU to the PCIe switch, it is then routed by the PCIe switch to the receiving GPU. Data transmission between GPUs first selects the PCIe switch closest to the receiving GPU as the primary PCIe switch for data transmission. When the bandwidth utilization of the primary PCIe switch exceeds a preset threshold, the system allows transactions transmitted to the receiving GPU to switch to other PCIe switches to alleviate the pressure on hot links.

[0029] Specifically, such as Figure 3 As shown, firstly, it is determined whether there is a data transmission requirement. If there is a data transmission requirement, the target acceleration unit of this application embodiment is determined based on the data transmission requirement. The target acceleration unit is the GPU used for transaction reception. Then, the first transmission path between the target acceleration unit and at least some switching devices is calculated, and multiple first transmission paths can be obtained. Among the multiple first transmission paths, the switching device with the shortest first transmission path is determined as the main switching device. That is, the PCIe switch closest to the GPU used for transaction reception is used as the main PCIe switch for data transmission.

[0030] Therefore, by selecting the PCIe switch closest to the GPU used for transaction reception as the primary PCIe switch for data transmission, the physical latency of data transmission can be directly and effectively reduced. At the same time, load balancing can be achieved, routing decisions can be simplified, and the system can quickly respond to each new transaction request, thereby improving the overall scheduling efficiency and real-time performance.

[0031] In step S102, if the bandwidth utilization of the main switching device is greater than the first preset utilization, then based on the first cyclic query strategy, it is queried whether there is a backup switching device that meets the preset backup conditions among the remaining switching devices other than the main switching device.

[0032] According to one embodiment of the present invention, based on a first cyclic query strategy, querying whether there is a backup switching device that meets preset backup conditions among the remaining switching devices other than the primary switching device among a plurality of switching devices includes: determining a target switching device based on the primary switching device according to the first cyclic query strategy, and obtaining the bandwidth utilization rate and the heat value of the target switching device; if the heat value of the target switching device is less than a first preset heat value, then determining whether the bandwidth utilization rate of the target switching device is less than a second preset utilization rate; if the bandwidth utilization rate of the target switching device is less than the second preset utilization rate, then determining that there is a backup switching device that meets the preset backup conditions.

[0033] According to one embodiment of the present invention, after determining that there is a backup switching device that meets the preset backup conditions, the method further includes: using the target switching device as the backup switching device and determining the data transmission path between the backup switching device and the target acceleration unit; and transmitting the data to be calculated to the target acceleration unit through the backup switching device based on the data transmission path.

[0034] The first preset utilization rate, the second preset utilization rate, the first preset heat value, and the preset backup conditions can all be set by those skilled in the art according to actual usage needs, or they can be obtained through a limited number of simulations, and are not specifically limited here.

[0035] Specifically, such as Figure 3 As shown, in one implementation method, the main PCIe switch first performs PCIe switch heat query control. The main PCIe switch maintains a PCIe switch heat table. Initially, the heat of all PCIe switches is 0. The host starts from the main PCIe switch and queries the bandwidth utilization of each PCIe switch to the right in turn. That is, the bandwidth utilization of the main PCIe switch is queried first. If the bandwidth utilization of the main PCIe switch is greater than or equal to the first preset utilization (for example, the bandwidth utilization exceeds 70%), then based on the first loop query strategy, the remaining switching devices other than the main switching device are queried, and it is determined whether there is a backup switching device that meets the preset backup conditions among the remaining switching devices. If so, the target switching device is determined based on the main switching device, and the bandwidth utilization and heat value of the target switching device are obtained.

[0036] Specifically, when the bandwidth utilization of the primary PCIe switch exceeds a first preset utilization rate, a backup path query mechanism is triggered. Based on a first loop query strategy, the next PCIe switch is queried sequentially to the right, and this next PCIe switch is designated as the target PCIe switch. The bandwidth utilization and heat value of the target PCIe switch are then obtained. Next, it is determined whether the heat value of the target PCIe switch is less than a first preset heat value (e.g., 5). In other words, after obtaining the bandwidth utilization and heat value of the target switch, if the heat value of the target PCIe switch is less than 5, the heat value of the target PCIe switch is incremented by 1. That is, during the query process, each time a PCIe switch is traversed and its heat value is less than 5, the backup path query mechanism is triggered. The heat value of the PCIe switch is increased by 1. Then, it is further determined whether the bandwidth utilization of the next switching device is less than the second preset utilization. If the bandwidth utilization of the next switching device is less than the second preset utilization, it means that there is a backup switching device that meets the preset backup conditions. Then, the next switching device is used as the target switching device, which is the backup switching device. Then, the data transmission path between the backup switching device and the target acceleration unit is determined. The target acceleration unit is the GPU. So, based on the data transmission path, the data to be computed can be transmitted to the target acceleration unit through the backup switching device. In other words, the system will instruct the GPU that receives the transaction to send it and enable the target PCIe switch as the backup PCIe switch for data transmission.

[0037] It should be noted that if a PCIe Switch does not respond during the query process, its bandwidth utilization will be considered to be 100%. After querying the rightmost PCIe Switch, the host will continue to query the leftmost PCIe Switch in a loop, forming a circular query mechanism.

[0038] Therefore, by using the core indicator of heat value, multiple functions such as path selection, congestion prevention, and load balancing are organically integrated to achieve adaptive data scheduling, while maintaining the smoothness and efficiency of circular queries, quickly shortening the discovery time of alternative paths, and reducing the overall communication latency caused by path selection delays.

[0039] According to one embodiment of the present invention, before obtaining the bandwidth utilization rate and the heat value of the target switching device, the method further includes: determining a first side and a second side of the main switching device based on a first cyclic query strategy, and obtaining the heat value of the first side and the heat value of the second side; updating a first preset heat value according to the heat value of the first side and the heat value of the second side.

[0040] According to one embodiment of the present invention, determining a first side and a second side of a primary switching device based on a first cyclic query strategy includes: determining a query direction according to the first cyclic query strategy; and, based on the primary switching device, designating the side opposite to the query direction as the first side and the side with the same query direction as the second side.

[0041] According to one embodiment of the present invention, updating a first preset heat value based on the heat value of a first side and a heat value of a second side includes: determining whether the heat value of the first side is greater than the heat value of the second side; if the heat value of the first side is greater than the heat value of the second side, then updating the first preset bandwidth utilization rate to a third preset bandwidth utilization rate; otherwise, updating the first preset bandwidth utilization rate to a fourth preset bandwidth utilization rate; wherein, the third preset bandwidth utilization rate is less than the first preset bandwidth utilization rate, and the fourth preset bandwidth utilization rate is greater than the first preset bandwidth utilization rate.

[0042] Specifically, in this embodiment of the application, before obtaining the bandwidth utilization and heat value of the target switching device, a threshold adjustment needs to be made based on the measured heat value of the main switching device to prevent the hardware link from becoming congested due to excessively high heat values ​​on one side.

[0043] Specifically, in this embodiment, firstly, based on a first cyclic query strategy, i.e., based on the query direction of the main switching device sequentially querying the next PCIe switch to the right, the side opposite to the query direction is designated as the first side, i.e., the first side of the main switching device is the left side, and the side in the same direction as the query direction is designated as the second side, i.e., the second side of the main switching device is the right side, and the heat values ​​of the first side and the second side are obtained; then, it is determined whether the heat value of the first side is greater than the heat value of the second side. If the heat value of the first side is greater than the heat value of the second side, it indicates that the link on the first side is more congested, and the heat value of the first side needs to be reduced. The bandwidth utilization rate of the first side allows the main switching equipment to continue to accommodate more load. Therefore, when the heat value of the first side is greater than that of the second side, the first preset bandwidth utilization rate (e.g., 70%) needs to be updated to the third preset bandwidth utilization rate (e.g., 50%). Conversely, if the heat value of the first side is less than that of the second side, it means that the link on the second side is more congested, and the bandwidth utilization rate of the first side can continue to accommodate more load. In this case, the bandwidth utilization rate of the first side can be increased, and the first preset bandwidth utilization rate can be updated to the fourth preset bandwidth utilization rate (e.g., 90%) to improve the link utilization rate.

[0044] In other words, in terms of dynamic threshold adjustment in this application embodiment, the first preset utilization rate can be set to 70% of the PCIe switch bandwidth utilization rate. If the heat value of the PCIe switch on the left side of the main PCIe switch is higher than that on the right side, it indicates that the left link is more congested. At this time, the first preset utilization rate can be reduced to 50% to accommodate more load. Conversely, if the heat value on the left side is lower than that on the right side, the first preset utilization rate is increased to 90% to improve the link utilization rate.

[0045] Therefore, by introducing a "threshold dynamic adjustment based on adjacent link heat comparison" mechanism, the embodiments of this application enable the system to perceive the difference in "heat value" between the links on both sides of the main switch, thereby dynamically adjusting the congestion judgment threshold of the primary path. This gives the system an intelligent control system with environmental awareness and self-adjustment capabilities, significantly improving the intelligence level of the algorithm. At the same time, it enables proactive avoidance of local hotspots, preventing traffic from continuing to flood into the already overloaded left link, and enhancing the system's robustness in dealing with unbalanced loads.

[0046] According to one embodiment of the present invention, after obtaining the bandwidth utilization rate and the heat value of the target switching device, the method further includes: if the heat value of the target switching device is greater than or equal to a first preset heat value, then reducing the heat value of the target switching device according to a preset heat reduction strategy, and determining the next switching device based on the primary switching device according to a first cyclic query strategy; taking the next switching device as the target switching device, and re-executing the steps of obtaining the bandwidth utilization rate and the heat value of the target switching device, until it is determined that the heat values ​​of at least some switching devices are greater than or equal to the first preset heat value or the bandwidth utilization rate of at least some switching devices are greater than or equal to the second preset utilization rate, and determining that there is no backup switching device that meets the preset backup conditions.

[0047] Specifically, such as Figure 3As shown, as another implementation, when the bandwidth utilization of the primary PCIe switch is greater than a first preset utilization rate, a backup path query mechanism is triggered. Based on the first loop query strategy, the next PCIe switch is queried to the right sequentially, and the next PCIe switch is taken as the target PCIe switch. The bandwidth utilization and heat value of the target PCIe switch are obtained. Then, it is determined whether the heat value of the target PCIe switch is greater than or equal to a first preset heat value, for example, whether the heat value of the target PCIe switch is greater than or equal to 5. If the heat value of the target PCIe switch is greater than or equal to 5, in order to prevent system instability and jitter and to achieve intelligent scheduling, this embodiment of the application needs to reduce the heat value of the target switching device according to a preset heat reduction strategy. That is, when encountering a target PCIe switch with a heat value greater than or equal to 5, the heat value of the target switching device needs to be reduced. At 5 o'clock, the query for the target PCIe switch is skipped. This means that even if its current bandwidth utilization is lower than the first preset utilization, the system will temporarily ignore it and decrement its popularity value by 1. During this period, the system will continue to determine the next switching device based on the first loop query strategy, that is, continue to query the next PCIe switch as the backup path in turn to the right, and re-execute the steps of obtaining the bandwidth utilization and popularity value of the target switching device until it is determined that the popularity value of at least some switching devices is greater than or equal to the first preset popularity value or the bandwidth utilization of at least some switching devices is greater than or equal to the second preset utilization. It can be determined that there is no backup switching device that meets the preset backup conditions. In other words, the backup switching device cannot be activated at this time, so it is necessary to forcibly distribute the traffic away from the popular path.

[0048] Therefore, by proactively reducing the heat value of the target switching device, traffic can be redirected and cooled down before the "hot" path truly becomes a bottleneck (by skipping queries and reducing heat by 1), thus significantly reducing the probability of actual congestion and improving network stability and data transmission smoothness.

[0049] According to an embodiment of the present invention, after determining whether the bandwidth utilization rate of the target switching device is less than the second preset utilization rate, the method further includes: if the bandwidth utilization rate of the target switching device is greater than or equal to the second preset utilization rate, then based on the first cyclic query strategy, determining the next switching device according to the primary switching device; taking the next switching device as the target switching device, and re-executing the steps of obtaining the bandwidth utilization rate and the heat value of the target switching device, until it is determined that the heat values ​​of at least some switching devices are greater than or equal to the first preset heat value or the bandwidth utilization rate of at least some switching devices are greater than or equal to the second preset utilization rate, and determining that there is no backup switching device that meets the preset backup conditions.

[0050] Specifically, if the bandwidth utilization of the target switching device is greater than or equal to the second preset utilization, it indicates that the data transmission has exceeded the limit and cannot be forwarded in time, which will increase the data processing speed. Therefore, it is necessary to further query the next PCIe switch to the right, and then take the next PCIe switch as the target switching device, and repeat the steps of obtaining the bandwidth utilization and heat value of the target switching device until it is determined that the heat value of at least some switching devices is greater than or equal to the first preset heat value or the bandwidth utilization of at least some switching devices is greater than or equal to the second preset utilization. It is then determined that there is no backup switching device that meets the preset backup conditions. This step has been explained above and will not be discussed in detail here to avoid redundancy.

[0051] Therefore, when the bandwidth utilization of the target switching device is greater than or equal to the second preset utilization, querying the next PCIe switch to the right ensures that the query process will not stall or jump randomly. At the same time, the system can check each possible PCIe switch in sequence and completely, avoiding the unfair phenomenon that some switches are ignored for a long time due to random selection, while other switches are over-queried, thus ensuring that all potential available paths have the opportunity to be discovered and utilized.

[0052] According to one embodiment of the present invention, after transmitting the data to be computed to the target acceleration unit through a backup switching device based on the data transmission path, the method further includes: obtaining the current transmission progress of the data to be computed to the target acceleration unit transmitted by the backup switching device; if the current transmission progress is complete, then reducing the heat value of the target switching device according to a preset heat reduction strategy.

[0053] Specifically, after the bandwidth utilization of the target switching device is greater than or equal to the second preset utilization and the next PCIe switch is queried to the right, the current transmission progress of the backup switching device transmitting the data to be calculated to the target acceleration unit is obtained. If all switching devices have completed the query, that is, the current transmission progress is completed, and the heat value of all switching devices is greater than 5, the scale-up network is considered to be in a fully loaded state. After the data transmission is completed, the heat value of the relevant link PCIe switch needs to be reduced by 1 according to the preset heat reduction strategy, so as to ensure that the heat table dynamically reflects the link load.

[0054] Therefore, when the scale-up network is fully loaded, it indicates that the effort to find alternative paths among all available switches has failed, and all paths are "overheated" due to frequent use. This means that the elastic scheduling capability of the scale-up network itself has been exhausted, and the overall load has reached the processing limit of its architecture. As a result, the heat value of the PCIe switches is reduced, and the system is no longer limited to queries between PCIe switches. Instead, a new, high-bandwidth escape channel is opened, namely the CXL shared memory pool. This is equivalent to providing elastic scheduling capability across network domains for data flow, transferring the pressure from the saturated scale-up network to the relatively idle scale-out network (CXL section below).

[0055] Specifically, after determining that the Scale-up network is fully loaded, the system obtains the bandwidth utilization and heat value of the high-speed serial computer expansion bus standard switch closest to the host. When the heat value of all PCIe switches in the Scale-up network is 5, the system determines that the network is fully loaded. At this time, the GPU data transaction offload channel is automatically opened, and the GPU will use the Memory-scale-up network to store the transaction data to be transmitted in the CXL controller of the shared memory pool (CXL shared memory pool) through the host.

[0056] Furthermore, the data transmission of the Memory-scale-up network adopts a similar heat query mechanism to the Scale-up network to monitor the bandwidth utilization of the CXL switch. The host maintains the CXL switch heat table, and the query process is similar to that of the PCIe switch mentioned above. By dynamically enabling the backup CXL switch to transmit transactions, single-path congestion can be effectively avoided. The CXL switch heat value is also limited to a maximum value of 5. When the limit is reached, the query is skipped and the CXL switch heat value is reduced by 1.

[0057] Furthermore, after placing the transaction data to be transmitted to the CXL controller, the GPU is sent to continue subsequent computing tasks. The host continuously sends the data location to the GPU's main PCIe switch until it is received. When the data transmission occupying the bandwidth is completed, the PCIe switch first determines whether there is a transaction to be transmitted. If there is a transaction to be transmitted, the PCIe switch will prioritize receiving the transaction to be transmitted stored in the CXL controller.

[0058] Therefore, when the network becomes fully loaded, the data to be transmitted is stored in a shared memory pool, which actively avoids potential congestion points, significantly reduces data transmission latency and packet loss risk, improves the reliability and predictability of communication, and avoids a sharp drop in performance due to congestion.

[0059] In step S103, if there is no backup switching device that meets the preset backup conditions, the target switch with the preset transmission conditions is determined from multiple first switches, and the data to be calculated is stored in the shared memory pool through the target switch.

[0060] According to one embodiment of the present invention, determining a target switch with preset transmission conditions from a plurality of first switches includes: determining a second transmission path between a host and at least a portion of the first switches; based on the second transmission path between the host and at least a portion of the first switches, selecting the first switch with the shortest second transmission path to the host among the at least a portion of the first switches as a candidate switch; obtaining the bandwidth utilization rate and initial switch heat value of the candidate switch, and determining whether the heat value of the candidate switch is less than a second preset heat value; if the heat value of the candidate switch is less than the second preset heat value, determining whether the bandwidth utilization rate of the candidate switch is less than a third preset utilization rate; if the bandwidth utilization rate of the candidate switch is less than the third preset utilization rate, selecting the candidate switch as the target switch.

[0061] According to one embodiment of the present invention, after determining whether the heat value of the pending switch is less than a second preset heat value, the method further includes: if the heat value of the pending switch is greater than or equal to the second preset heat value, then according to a preset heat reduction strategy, the heat value of the pending switch is reduced, and based on a second loop query strategy, the next switch is determined according to the host; the next switch is designated as the pending switch, and the steps of obtaining the bandwidth utilization rate and heat value of the pending switch are re-executed until it is determined that the heat value of the pending switch is less than the second preset heat value.

[0062] According to one embodiment of the present invention, after storing the data to be computed into the shared memory pool through the target switch, the method further includes: determining whether the bandwidth utilization rate and heat value of the main switching device both meet the corresponding conditions; if the bandwidth utilization rate and heat value of the main switching device both meet the corresponding conditions, then receiving the transaction data to be transmitted through the main switching device and transmitting the transaction data to be transmitted to the acceleration unit for transaction reception.

[0063] The preset transmission conditions can be set by those skilled in the art based on actual network usage needs, and are not specifically limited here.

[0064] Specifically, among the bandwidth utilization and heat value of the target switching device calculated above, if there is no backup switching device that meets the preset backup conditions among the remaining switching devices other than the main switching device, then based on the first loop query strategy, the target switching device that meets the preset transmission conditions is queried among the multiple first switches. That is, the second transmission path between the host and at least some of the first switches is determined. Then, among the multiple second transmission paths, the first switch with the shortest second transmission path is selected as the pending switch, and the above process of obtaining the bandwidth utilization and heat value of the pending switch is continued. It is then determined whether the heat value of the pending switch is less than the second preset heat value (e.g., 5). If the heat value of the pending switch is less than the second preset heat value, the heat value of the pending switch is incremented by 1, and it is determined whether the bandwidth utilization of the pending switch is less than the third preset utilization. If the bandwidth utilization of the pending switch is less than the third preset utilization, the pending switch is selected as the target switch. The determination process of the target switch is the same as the determination process of the main switching device. To avoid redundancy, it will not be described in detail here.

[0065] Furthermore, if the heat value of the pending switch is greater than or equal to the second preset heat value, in order to prevent system instability and jitter and to achieve intelligent scheduling, it is necessary to continue to reduce the heat value of the pending switch according to the preset heat reduction strategy. Based on the second loop query strategy, the next switch is determined according to the host, that is, the next PCIe switch is queried sequentially to the right as the backup path, and the next switch is designated as the pending switch. The steps of obtaining the bandwidth utilization and heat value of the pending switch are repeated until it is determined that the heat value of the pending switch is less than the second preset heat value. Then, it is further determined whether the bandwidth utilization of the pending switch is less than the third preset utilization. If the bandwidth utilization of the pending switch is less than the third preset utilization, it means that there is a backup switching device that meets the preset backup conditions. The next switching device is designated as the target switching device, which is the backup switching device. Then, the data transmission path between the backup switching device and the target acceleration unit is determined, so that the data to be calculated can be transmitted to the target acceleration unit through the backup switching device based on the data transmission path.

[0066] Therefore, by using the core indicator of heat value, multiple functions such as path selection, congestion prevention, and load balancing are organically integrated to achieve adaptive data scheduling, while maintaining the smoothness and efficiency of circular queries, quickly shortening the discovery time of alternative paths, and reducing the overall communication latency caused by path selection delays.

[0067] According to one embodiment of the present invention, after determining whether the bandwidth utilization rate of the pending switch is less than a third preset utilization rate, the method further includes: if the bandwidth utilization rate of the pending switch is greater than or equal to the third preset utilization rate, then based on the second loop query strategy, the next switch is determined according to the host; the next switch is designated as the pending switch, and the steps of obtaining the bandwidth utilization rate and the heat value of the pending switch are re-executed until it is determined that the bandwidth utilization rate of the pending switch is less than the third preset utilization rate, and the pending switch is designated as the target switch.

[0068] Specifically, if the bandwidth utilization of the pending switch is greater than or equal to the third preset utilization rate, it also indicates that the data transmission has exceeded the limit and cannot be forwarded in time, which will increase the data processing speed. Therefore, it is necessary to use the second loop query strategy, that is, to query the next PCIe switch to the right in turn, and then use the next PCIe switch as the pending switch, and re-execute the steps of obtaining the bandwidth utilization and heat value of the pending switch until it is determined that the bandwidth utilization of the pending switch is less than the third preset utilization rate, and then use the pending switch as the target switch. The target switching device is the backup switching device. Then, based on the data transmission path between the backup switching device and the target acceleration unit, the data to be calculated is transmitted to the target acceleration unit through the backup switching device.

[0069] Therefore, when the bandwidth utilization of the target switching device is greater than or equal to the third preset utilization, querying the next PCIe switch to the right ensures that the query process will not be stalled or randomly skipped. At the same time, the system can check each possible PCIe switch in sequence and completely, avoiding the unfair phenomenon that some switches are ignored for a long time due to random selection, while other switches are over-queried, thus ensuring that all potential available paths have the opportunity to be discovered and utilized.

[0070] According to one embodiment of the present invention, the method further includes: obtaining the current video memory utilization rate and the current computing power utilization rate of the acceleration unit; determining whether the current video memory utilization rate is greater than a third preset utilization rate; if the current video memory utilization rate is greater than the third preset utilization rate, determining whether the current computing power utilization rate is less than a fourth preset utilization rate; if the current computing power utilization rate is less than the fourth preset utilization rate, opening the data offloading channel and storing the data of the acceleration unit in the shared memory pool in the form of memory interleaving.

[0071] The third and fourth preset utilization rates can be set by those skilled in the art based on actual usage needs, or they can be obtained through a limited number of simulations, and are not specifically limited here.

[0072] Specifically, such as Figure 4As shown, since the GPU used for transaction reception can dynamically adjust its data offloading strategy based on its real-time calculated current computing power utilization and current memory utilization, this embodiment of the application can obtain the GPU's current memory utilization and current computing power utilization. If the current memory utilization is greater than a third preset utilization (e.g., the current memory utilization is greater than 90%), it further determines whether the current computing power utilization is less than a fourth preset utilization. If the current computing power utilization is less than the fourth preset utilization (e.g., the current computing power utilization is less than 70%), it indicates that the system is already under significant memory pressure and is highly susceptible to triggering [a certain event / problem]. A memory overflow error can cause task interruption or a sharp drop in performance (due to frequent memory swapping). However, since the current computing power utilization is less than 70%, it means that the GPU is waiting for data or a synchronization signal. At this time, the GPU's computing units are idle, but the I / O channels may have idle bandwidth. Therefore, the I / O bandwidth in these waiting cycles can be used to perform data offloading, minimizing the impact on core computing tasks. Thus, the data offloading channel is enabled at this time, and the GPU's data is stored in the CXL shared memory pool in a memory interleaved form, releasing the GPU's memory pressure and improving the overall system efficiency.

[0073] According to an embodiment of the present invention, after determining whether the current video memory utilization rate is greater than a third preset utilization rate, the method further includes: if the current video memory utilization rate is less than or equal to the third preset utilization rate, then re-execute the step of obtaining the current video memory utilization rate and the current computing power utilization rate of the acceleration unit.

[0074] According to one embodiment of the present invention, after storing the data of the acceleration unit in the shared memory pool in the form of memory interleaving, the method further includes: re-acquiring the new computing power utilization rate of the acceleration unit; if the re-acquired computing power utilization rate is greater than a fifth preset utilization rate, then closing the data offloading channel.

[0075] The fifth preset utilization rate can be set by those skilled in the art based on actual usage needs, or it can be obtained through a limited number of simulations, and is not specifically limited here.

[0076] Specifically, such as Figure 4 As shown, if the current memory utilization is less than or equal to 90%, it means that the system memory still has space and has not reached saturation. Therefore, the steps of obtaining the current memory utilization and current computing power utilization of the GPU can be repeated. At this time, the new computing power utilization of the GPU can be obtained again. If the new computing power utilization is greater than the fifth preset utilization, that is, when the current computing power utilization is greater than 90%, since the GPU computing power is close to saturation, performing the unloading operation at this time may make matters worse and further slow down the core computing tasks. Therefore, the data unloading can be paused first to ensure computing performance.

[0077] Therefore, when the current memory utilization is less than or equal to 90%, by continuously executing the steps of obtaining the current memory utilization and current computing power utilization of the GPU, a never-ending, dynamic resource status monitoring closed loop can be achieved. This allows for the capture of instantaneous resource fluctuations, ensuring the timeliness and accuracy of scheduling decisions, while also enabling it to adapt to complex and ever-changing AI workloads. Furthermore, when the current computing power utilization is greater than 90%, to avoid further slowing down core computing tasks by continuing to perform unloading operations, data unloading is paused to ensure computing performance.

[0078] It should be noted that the scale-out network uses the CXL shared memory pool to transmit data. The sending host stores the data in the CXL shared memory pool in a memory-interleaved manner and uses the RDMA path to notify the data storage address. The host receives the data storage address notification through the RDMA protocol and then passes it to the corresponding GPU. The GPU can directly access the data in the CXL memory for computation.

[0079] In summary, the embodiments of this application combine multi-level heat query, dynamic threshold adjustment, multi-path backup switch activation and data offloading mechanism to form a set of efficient hybrid network congestion control and resource scheduling scheme for large-scale AI training and inference. This significantly improves the link bandwidth utilization and computing resource coordination efficiency in multi-GPU and multi-host systems, ensuring low latency and high throughput performance of the system under large-scale tasks.

[0080] The main features are: (1) Multi-level heat perception and backup switch selection mechanism. A heat table mechanism is introduced in PCIe Switch and CXLSwitch to monitor bandwidth utilization in real time. The timing of backup path activation is controlled by ring sequential query and heat value limit to avoid path jitter caused by frequent switching. At the same time, the backup path activation threshold (50%~90%) is adaptively adjusted according to the difference in link heat on both sides of the primary switch to optimize link utilization and traffic allocation under different congestion distribution conditions; (2) Cross-network transaction data offloading mechanism and computing power-GPU memory joint scheduling strategy. When the Scale-up network enters a full-load state, GPU transaction data is written to the CXL shared memory pool through the Memory-scale-up network to alleviate system congestion pressure and form cross-domain bandwidth elastic scheduling capability. Meanwhile, based on the real-time calculation results of GPU computing power utilization and video memory utilization, it is dynamically determined whether to unload data to CXL memory in a memory interleaving manner to achieve bidirectional optimization of video memory and computing power resources; (3) Multi-GPU multi-Host global traffic and resource collaborative control framework, combined with CXL shared memory pool and RDMA protocol, sends the host to write data to CXL memory in a memory interleaving manner and sends address notification through RDMA, and the GPU receiving the host can directly access the remote data, eliminating CPU relay delay. Link congestion control, path switching, bandwidth threshold adjustment, transaction offloading, computing power / video memory scheduling and cross-host direct access are integrated into a unified control logic to achieve overall optimization of multiple network domains and multiple computing resources.

[0081] Therefore, compared with existing AI servers that rely solely on a single Scale-up or Scale-out network architecture, this application introduces multi-level heat perception, dynamic threshold adjustment, alternative path switching, and cross-network data offloading mechanisms into a hybrid interconnect architecture of Scale-up + Scale-out + CXL shared memory pool. This achieves low congestion, high throughput, and high resource utilization under large-scale AI training and inference tasks, resulting in the following beneficial effects: (1) To achieve efficient adaptation of large-scale AI hybrid architecture, this application's embodiments introduce a CXL shared memory pool in a multi-GPU, multi-host system to replace traditional RDMA cross-node communication, thereby achieving high-bandwidth, low-latency data interaction between hosts, while taking into account the advantages of both scale-up and scale-out architectures. In large-scale model training, inference, and multimodal processing tasks, it significantly reduces cross-node data access latency and improves overall resource utilization.

[0082] (2) Achieving global congestion awareness and dynamic scheduling across networks: By deploying a unified link heat table and bandwidth utilization monitoring mechanism in both PCIe Switch and CXL Switch networks, this invention can achieve collaborative congestion control between scale-up and scale-out networks. Data flows can be dynamically distributed between different network domains, avoiding global performance bottlenecks caused by local link saturation and ensuring continuous high throughput under multi-network collaboration.

[0083] (3) To achieve dynamic collaborative utilization of video memory and computing power, this application embodiment dynamically adjusts the data offloading and backhaul strategy based on the monitoring results of GPU real-time video memory utilization and computing power utilization, as well as the capacity of the CXL shared memory pool. When video memory is scarce and computing power utilization is low, data is automatically interleaved and stored in CXL memory to release video memory pressure; when computing power demand is high, offloading is paused to prioritize computing performance, thereby reducing video memory overflow and computing power idleness.

[0084] (4) To improve the bandwidth utilization and overall throughput of multi-switching chip links, this application adopts a multi-path traffic rerouting strategy and combines it with a dynamic adjustment mechanism for hot thresholds. When link congestion occurs, it can quickly switch to backup switching paths to balance the bandwidth load of multiple switching chips. This mechanism effectively alleviates hot link congestion and improves the overall network throughput and data transmission stability.

[0085] According to the resource scheduling method proposed in this embodiment of the invention, the bandwidth utilization rate of the main switching device among multiple switching devices is obtained by the host. If the bandwidth utilization rate of the main switching device is greater than a first preset utilization rate, a first cyclic query strategy is used to query whether there is a backup switching device among the remaining switching devices other than the main switching device that meets the preset backup conditions. If not, a target switching device with preset transmission conditions is determined from multiple first switches, and the data to be computed is stored in the shared memory pool through the target switching device. This solves the technical problems of communication bottlenecks and low resource utilization caused by the fragmentation of cross-network congestion control in AI systems. By introducing multi-level heat perception, dynamic threshold adjustment, backup path switching, and cross-network data offloading mechanisms in the hybrid interconnection architecture of Scale-up+Scale-out+CXL shared memory pool, low congestion, high throughput, and high resource utilization are achieved under large-scale AI training and inference tasks.

[0086] Figure 5 This is a block diagram of a resource scheduling device according to an embodiment of this application.

[0087] like Figure 5As shown, the resource scheduling device 10 is applied to a heterogeneous computing system. The heterogeneous computing system includes a host, a first switch, and a shared memory pool. Both the host and the shared memory pool are connected to the first switch. The host includes multiple switching devices, including: an acquisition module 100, a query module 200, and a storage module 300.

[0088] Among them, the acquisition module 100 is used to acquire the bandwidth utilization rate of the main switching device among multiple switching devices through the host; The query module 200 is used to query, based on the first cyclic query strategy, whether there is a backup switching device that meets the preset backup conditions among the remaining switching devices other than the main switching device if the bandwidth utilization rate of the main switching device is greater than the first preset utilization rate. The storage module 300 is used to determine the target switch with preset transmission conditions from a plurality of first switches if there is no backup switching device that meets the preset backup conditions, and to store the data to be calculated into a shared memory pool through the target switch.

[0089] For a description of the features in the embodiment corresponding to the resource scheduling device, please refer to the relevant description in the embodiment corresponding to the resource scheduling method, which will not be repeated here.

[0090] Figure 6 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. The electronic device may include: The memory 601, the processor 602, and the computer program stored on the memory 601 and capable of running on the processor 602.

[0091] When the processor 602 executes the program, it implements the resource scheduling method provided in the above embodiments.

[0092] Furthermore, electronic devices also include: Communication interface 603 is used for communication between memory 601 and processor 602.

[0093] The memory 601 is used to store computer programs that can run on the processor 602.

[0094] The memory 601 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0095] If the memory 601, processor 602, and communication interface 603 are implemented independently, then the communication interface 603, memory 601, and processor 602 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 6 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0096] Optionally, in a specific implementation, if the memory 601, processor 602, and communication interface 603 are integrated on a single chip, then the memory 601, processor 602, and communication interface 603 can communicate with each other through an internal interface.

[0097] Processor 602 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.

[0098] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described resource scheduling method embodiments at runtime.

[0099] Embodiments of the present invention also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described resource scheduling method embodiments.

[0100] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can implement the described functions using different methods for at least one specific application, but such implementation should not be considered beyond the scope of the invention.

[0101] The reinforcement learning computational simulation method provided by this invention has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this invention. It should be noted that those skilled in the art can make several improvements and modifications to this invention without departing from the principles of this invention, and these improvements and modifications also fall within the protection scope of the claims of this invention.

Claims

1. A resource scheduling method, characterized in that, The method is applied to a heterogeneous computing system, which includes a host, a first switch, and a shared memory pool. Both the host and the shared memory pool are connected to the first switch. The host includes multiple switching devices. The method includes the following steps: The bandwidth utilization rate of the main switching device among the plurality of switching devices is obtained through the host. If the bandwidth utilization of the main switching device is greater than the first preset utilization, then based on the first cyclic query strategy, query whether there is a backup switching device that meets the preset backup conditions among the remaining switching devices other than the main switching device. If no backup switching device meets the preset backup conditions, a target switch with preset transmission conditions is determined from among the multiple first switches, and the data to be calculated is stored in the shared memory pool through the target switch.

2. The method according to claim 1, characterized in that, The step of querying whether there are any backup switching devices that meet preset backup conditions among the remaining switching devices other than the primary switching device, based on the first loop query strategy, includes: Based on the first loop query strategy, the target switching device is determined according to the main switching device, and the bandwidth utilization and heat value of the target switching device are obtained. If the heat value of the target switching device is less than the first preset heat value, then it is determined whether the bandwidth utilization rate of the target switching device is less than the second preset utilization rate. If the bandwidth utilization of the target switching device is less than the second preset utilization, it is determined that there is a backup switching device that meets the preset backup conditions.

3. The method according to claim 2, characterized in that, After obtaining the bandwidth utilization and heat value of the target switching device, the following is also included: If the heat value of the target switching device is greater than or equal to the first preset heat value, then the heat value of the target switching device is reduced according to the preset heat reduction strategy, and the next switching device is determined based on the first loop query strategy according to the main switching device. The next switching device is taken as the target switching device, and the steps of obtaining the bandwidth utilization rate and the heat value of the target switching device are executed again until it is determined that the heat value of at least some switching devices is greater than or equal to the first preset heat value or the bandwidth utilization rate of at least some switching devices is greater than or equal to the second preset utilization rate. Then it is determined that there is no backup switching device that meets the preset backup conditions.

4. The method according to claim 2, characterized in that, After determining whether the bandwidth utilization rate of the target switching device is less than the second preset utilization rate, the method further includes: If the bandwidth utilization of the target switching device is greater than or equal to the second preset utilization, then the next switching device is determined based on the first cyclic query strategy and the main switching device. The next switching device is taken as the target switching device, and the steps of obtaining the bandwidth utilization rate and the heat value of the target switching device are executed again until it is determined that the heat value of at least some switching devices is greater than or equal to the first preset heat value or the bandwidth utilization rate of at least some switching devices is greater than or equal to the second preset utilization rate. Then it is determined that there is no backup switching device that meets the preset backup conditions.

5. The method according to claim 2, characterized in that, After determining that a backup switching device exists that meets the preset backup conditions, the process further includes: The target switching device is used as a backup switching device, and the data transmission path between the backup switching device and the target acceleration unit is determined. Based on the data transmission path, the data to be computed is transmitted to the target acceleration unit via the backup switching device.

6. The method according to claim 1, characterized in that, Before obtaining the bandwidth utilization of the primary switching device among the plurality of switching devices through the host, the method further includes: Determine if there is a data transmission requirement; If the data transmission requirement exists, a target acceleration unit is determined based on the data transmission requirement, and a first transmission path is calculated between the target acceleration unit and at least some of the switching devices. Based on the first transmission path between the target acceleration unit and at least some of the switching devices, the switching device with the shortest first transmission path between the target acceleration unit and the at least some of the switching devices is designated as the main switching device.

7. The method according to claim 5, characterized in that, After transmitting the data to be computed to the target acceleration unit via the backup switching device based on the data transmission path, the method further includes: Obtain the current transmission progress of the data to be computed from the backup switching device to the target acceleration unit; If the current transmission progress indicates that the transmission is complete, then the heat value of the target switching device is reduced according to the preset heat reduction strategy.

8. The method according to claim 1, characterized in that, The step of determining the target switch with preset transmission conditions from a plurality of the first switches includes: Determine a second transmission path between the host and at least a portion of the first switch; Based on the second transmission path between the host and at least a portion of the first switches, the first switch with the shortest second transmission path to the host among the at least a portion of the first switches is designated as the undetermined switch. Obtain the bandwidth utilization rate of the switch to be determined and the heat value of the initial switch, and determine whether the heat value of the switch to be determined is less than the second preset heat value. If the heat value of the pending switch is less than the second preset heat value, then it is determined whether the bandwidth utilization rate of the pending switch is less than the third preset utilization rate. If the bandwidth utilization rate of the pending switch is less than the third preset utilization rate, then the pending switch will be designated as the target switch.

9. The method according to claim 8, characterized in that, After determining whether the heat value of the switch to be determined is less than the second preset heat value, the process further includes: If the heat value of the pending switch is greater than or equal to the second preset heat value, then the heat value of the pending switch is reduced according to the preset heat reduction strategy, and the next switch is determined based on the host according to the second loop query strategy. The next switch is designated as the pending switch, and the steps of obtaining the bandwidth utilization and popularity value of the pending switch are repeated until it is determined that the popularity value of the pending switch is less than the second preset popularity value.

10. The method according to claim 8, characterized in that, After determining whether the bandwidth utilization rate of the switch to be determined is less than the third preset utilization rate, the process further includes: If the bandwidth utilization rate of the pending switch is greater than or equal to the third preset utilization rate, then the next switch is determined based on the host according to the second loop query strategy. The next switch is designated as the pending switch, and the steps of obtaining the bandwidth utilization rate and heat value of the pending switch are repeated until it is determined that the bandwidth utilization rate of the pending switch is less than the third preset utilization rate, and the pending switch is designated as the target switch.

11. The method according to claim 2, characterized in that, Before obtaining the bandwidth utilization and heat value of the target switching device, the following steps are also included: Based on the first loop query strategy, the first side and the second side of the main switching device are determined, and the heat value of the first side and the heat value of the second side are obtained. The first preset heat value is updated based on the heat value of the first side and the heat value of the second side.

12. The method according to claim 11, characterized in that, The step of determining the first side and the second side of the main switching device based on the first cyclic query strategy includes: The query direction is determined based on the first loop query strategy; Based on the main switching device, the side opposite to the query direction is designated as the first side, and the side in the same direction as the query direction is designated as the second side.

13. The method according to claim 11 or 12, characterized in that, The step of updating the first preset heat value based on the heat value of the first side and the heat value of the second side includes: Determine whether the heat value of the first side is greater than the heat value of the second side; If the heat value of the first side is greater than the heat value of the second side, then the first preset bandwidth utilization rate is updated to the third preset bandwidth utilization rate; otherwise, the first preset bandwidth utilization rate is updated to the fourth preset bandwidth utilization rate. The third preset bandwidth utilization rate is less than the first preset bandwidth utilization rate, and the fourth preset bandwidth utilization rate is greater than the first preset bandwidth utilization rate.

14. The method according to claim 1, characterized in that, After storing the data to be computed into the shared memory pool via the target switch, the process further includes: Determine whether the bandwidth utilization and heat value of the main switching device both meet the corresponding conditions; If the bandwidth utilization and heat value of the main switching device both meet the corresponding conditions, the main switching device receives the transaction data to be transmitted and transmits the transaction data to the acceleration unit for transaction reception.

15. The method according to claim 1, characterized in that, Also includes: Get the current memory utilization and current computing power utilization of the acceleration unit; Determine whether the current video memory utilization rate is greater than the third preset utilization rate; If the current video memory utilization rate is greater than the third preset utilization rate, then determine whether the current computing power utilization rate is less than the fourth preset utilization rate; If the current computing power utilization rate is less than the fourth preset utilization rate, the data offloading channel is activated, and the data of the acceleration unit is stored in the shared memory pool in a memory interleaving form.

16. The method according to claim 15, characterized in that, After storing the data of the acceleration unit into the shared memory pool in a memory-interleaved manner, the method further includes: Reacquire the computing power utilization rate of the acceleration unit; If the newly acquired computing power utilization rate is greater than the fifth preset utilization rate, then the data unloading channel will be closed.

17. The method according to claim 15, characterized in that, After determining whether the current video memory utilization rate is greater than the third preset utilization rate, the method further includes: If the current memory utilization rate is less than or equal to the third preset utilization rate, then the step of obtaining the current memory utilization rate and current computing power utilization rate of the acceleration unit is executed again.

18. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the steps of the resource scheduling method as described in any one of claims 1 to 17.

19. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the resource scheduling method as described in any one of claims 1 to 17.

20. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the resource scheduling method as described in any one of claims 1 to 17.

Citation Information

Patent Citations

  • Storage system, cluster node, system, storage resource scheduling method and device

    CN119045750A

  • System, server and method for sharing memory resource pool, medium and program product

    CN120448141A

  • Programmable switching engine with storage, analytic and processing capabilities

    US20150063349A1

  • Network having secure fast packet switching and guaranteed quality of service

    US5485455A

Cited By

  • Host, interconnection system, data processing method, equipment, medium and product

    CN122268806A