Method and apparatus for dynamically scheduling computing power in a high-bandwidth network of an intelligent computing center cloud platform

By introducing a computing job awareness module and a link resource mapping engine into the intelligent computing center, fine-grained resource isolation and dynamic bandwidth scheduling for computing jobs are achieved, solving the on-demand problem of computing bandwidth scheduling in the intelligent computing center and improving network stability and resource utilization.

CN120343086BActive Publication Date: 2025-10-28DATACANVAS LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510803700.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-10-28
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

In existing intelligent computing centers, bandwidth scheduling of computing power faces difficulties in on-demand scheduling, which affects network stability and the stability of access to critical computing resources.

Method used

Metadata is collected by the computing power job perception module. Combined with the computing power bandwidth scheduling unit and the link resource mapping engine, it can achieve accurate perception and fine-grained resource isolation of job priority, bandwidth requirements and network topology, and dynamically adjust physical link resources to ensure the priority and bandwidth allocation of computing power job data packets in the virtual channel.

Benefits of technology

It improves the utilization rate of computing resources and the accuracy of on-demand scheduling, solves the problems of multi-tenant bandwidth contention and communication topology mismatch, and realizes a high-bandwidth network environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343086B_ABST
    Figure CN120343086B_ABST
Patent Text Reader

Abstract

This invention provides a method and apparatus for dynamically scheduling computing power in a high-bandwidth network within an intelligent computing center cloud platform, relating to the field of computing infrastructure technology. The method includes: a computing power main control management unit sending computing job metadata pushed by a computing job perception module to a computing power bandwidth scheduling unit; the computing power bandwidth scheduling unit determining first scheduling data for the computing job identifier based on the bandwidth requirement information and the priority information; the computing power bandwidth scheduling unit sending the first scheduling data to the GPU server indicated by the network topology information; and a link resource mapping engine sending second scheduling data to the switch indicated by the network topology information. This achieves dynamic adaptation between computing power running tasks and computing power bandwidth resources, effectively realizing a high-bandwidth network for computing power.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent computing centers, smart computing centers and computing infrastructure technology, and specifically to a method and apparatus for dynamically scheduling computing power in a high-bandwidth network of an intelligent computing center cloud platform. Background Technology

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "smart computing centers" have emerged.

[0003] An "intelligent computing center" refers to a facility that provides the necessary computing power, data, and algorithms for artificial intelligence applications (such as the development, training, and inference of deep learning models) by utilizing large-scale heterogeneous computing resources, including general-purpose and intelligent computing power. Intelligent computing centers encompass facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enablement.

[0004] "Intelligent computing center" includes, but is not limited to, "intelligent computing center".

[0005] "Intelligent computing center" or artificial intelligence computing center is a type of computing infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications, based on artificial intelligence theory and adopting artificial intelligence computing architecture.

[0006] "Computing power" is the core of "intelligent computing center" and "smart computing center". It is the ability of computer equipment or computing / data center to process information. It is the ability of computer hardware and software to work together to perform a certain computing requirement. It is the computing power to achieve the target output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity and data storage capacity. It mainly provides services to society through computing power infrastructure.

[0007] In scenarios such as Artificial Intelligence (AI) training, scientific computing, and large-scale data processing, higher demands are placed on high-bandwidth, low-latency data exchange capabilities. Traditional Network Interface Cards (NICs) typically offer bandwidths of 10Gbps / 25Gbps / 40Gbps, while some high-end data centers have deployed SmartNICs with specifications of 100Gbps or higher to support more complex virtualization and acceleration capabilities. In typical SmartNIC solutions, such as those based on the Mellanox ConnectX series, network performance is improved through built-in protocol stack acceleration mechanisms such as remote direct memory access (RDMA) and the Intel Data Plane Development Kit (DPDK). Their structure generally includes components such as a main controller, DMA engine, network protocol processing module, PCIe interface, and SR-IOV virtual channel, working with drivers to achieve efficient interaction with the host operating system. However, as AI models grow in scale and inter-cluster communication becomes more frequent, existing SmartNIC solutions are still prone to bandwidth contention, affecting the stability of the network used by important tenants of the intelligent computing center cloud platform, as well as the stability of access to some critical computing resources.

[0008] It is evident that since the emergence of intelligent computing centers, how to achieve on-demand scheduling of computing power bandwidth to realize high-bandwidth networks has become an urgent problem to be solved. Summary of the Invention

[0009] This invention provides a method and apparatus for dynamically scheduling computing power in a high-bandwidth network on a cloud platform of an intelligent computing center, in order to solve the problem of how to achieve on-demand scheduling of computing power bandwidth to achieve a high-bandwidth network since the emergence of intelligent computing centers.

[0010] To solve the above problems, the present invention is implemented as follows:

[0011] In a first aspect, embodiments of the present invention provide a method for dynamically scheduling computing power in a high-bandwidth network of an intelligent computing center cloud platform, the method comprising:

[0012] Step S1: The computing power main control management unit sends the computing power job metadata pushed by the computing power job perception module to the computing power bandwidth scheduling unit. The computing power job metadata includes the computing power job identifier corresponding to the computing power job data packet, the bandwidth requirement information corresponding to the computing power job data packet, the priority information corresponding to the computing power job data packet, and the network topology information. The computing power job perception module is used to subscribe to the computing power job metadata of the GPU server.

[0013] Step S2: The computing power bandwidth scheduling unit determines the first scheduling data of the computing power job identifier based on the bandwidth requirement information and the priority information. The first scheduling data includes the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel and the bandwidth information of the virtual channel. Different virtual channels correspond to different priorities.

[0014] Step S3: The computing power bandwidth scheduling unit sends the first scheduling data to the GPU server indicated by the network topology information. The GPU server is used to perform bandwidth scheduling according to the first scheduling data.

[0015] Step S4: The link resource mapping engine sends the second scheduling data to the switch indicated by the network topology information. The switch is used to perform bandwidth scheduling according to the second scheduling data. The second scheduling data is the mapping data of the first scheduling data generated by the link resource mapping engine based on the first scheduling data.

[0016] In one embodiment, after step S3 and before step S4, the method further includes:

[0017] Step S5: The link resource mapping engine obtains link layer discovery protocol information from the switch, and the link layer discovery protocol information is used for link resource mapping.

[0018] Step S6: The link resource mapping engine generates the second scheduling data based on the link layer discovery protocol information and the first scheduling data sent by the computing power bandwidth scheduling unit.

[0019] In one embodiment, step S2 includes:

[0020] Step S21: The computing power bandwidth scheduling unit sets a differential service code point for the job data indicated by the computing power job identifier according to the priority information. The differential service code point is used to determine the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel.

[0021] Step S22: When the differential service code point indicates that the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel is the first channel, the computing power bandwidth scheduling unit determines the first bandwidth information of the first channel according to the bandwidth demand information. The first scheduling data includes the differential service code point and the first bandwidth information.

[0022] In one embodiment, step S2 further includes:

[0023] Step S23: When the differential service code point indicates that the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel is the second channel, the computing power bandwidth scheduling unit generates a scheduling instruction for the second channel. The scheduling instruction is used to instruct the computing power job data packet to be allocated to the second channel. The second bandwidth information of the second channel is a fixed bandwidth value. The first scheduling data includes the differential service code point and the scheduling instruction. The priority of the second channel is greater than the priority of the first channel.

[0024] In one embodiment, step S22 includes:

[0025] Step S221: The computing power bandwidth scheduling unit determines the first bandwidth information of the first channel based on the proportion of the bandwidth value corresponding to the bandwidth demand information to the total bandwidth value corresponding to the total bandwidth demand information.

[0026] In one embodiment, after step S3, the method further includes:

[0027] Step S7: Upon receiving the first scheduling data, the GPU server calls the multi-channel transceiver module to generate the first virtual channel. Each virtual channel in the first virtual channel has an independent caching and queue management mechanism.

[0028] Step S8: The GPU server performs bandwidth scheduling based on the first virtual channel and the first scheduling data.

[0029] Secondly, embodiments of the present invention also provide a device for dynamically scheduling computing power in a high-bandwidth network of an intelligent computing center cloud platform, the device comprising a computing power main control management unit, a computing power bandwidth scheduling unit, a link resource mapping engine, and a computing power job perception module;

[0030] The computing power master control management unit is used to send the computing power job metadata pushed by the computing power job perception module to the computing power bandwidth scheduling unit. The computing power job metadata includes the computing power job identifier corresponding to the computing power job data packet, the bandwidth requirement information corresponding to the computing power job data packet, the priority information corresponding to the computing power job data packet, and the network topology information. The computing power job perception module is used to subscribe to the computing power job metadata of the GPU server.

[0031] The computing power bandwidth scheduling unit is used to determine the first scheduling data of the computing power job identifier according to the bandwidth demand information and the priority information. The first scheduling data includes the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel and the bandwidth information of the virtual channel. Different virtual channels correspond to different priorities.

[0032] The computing power bandwidth scheduling unit is further configured to send the first scheduling data to the GPU server indicated by the network topology information, and the GPU server is configured to perform bandwidth scheduling according to the first scheduling data;

[0033] The link resource mapping engine is used to send the second scheduling data to the switch indicated by the network topology information. The switch is used to perform bandwidth scheduling according to the second scheduling data. The second scheduling data is the mapping data of the first scheduling data generated by the link resource mapping engine based on the first scheduling data.

[0034] Thirdly, the present invention also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the method for dynamically scheduling computing power in a high-bandwidth network of an intelligent computing center cloud platform as described in the first aspect above.

[0035] Fourthly, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the method for dynamically scheduling computing power in a high-bandwidth network of an intelligent computing center cloud platform as described in the first aspect above.

[0036] Fifthly, the present invention also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the method for dynamically scheduling computing power in a high-bandwidth network of an intelligent computing center cloud platform as described in the first aspect above.

[0037] In this embodiment of the invention, the computing power job perception module collects computing power job metadata and transmits it to the computing power bandwidth scheduling unit, thereby achieving accurate perception of job priority, bandwidth requirements, and network topology. The computing power master control management unit sends the computing power job metadata to the computing power bandwidth scheduling unit, which allocates virtual channels of different priorities and configures corresponding bandwidth information for computing power job data packets indicated by different computing power job identifiers. This solves the problem of coarse granularity in traditional scheduling and achieves fine-grained resource isolation and priority guarantee. Then, the computing power bandwidth scheduling unit sends the first scheduling data to the GPU server for execution, enabling the GPU server to mark traffic and bind queues according to policies, thus achieving on-demand bandwidth scheduling. Furthermore, the link resource mapping engine maps virtual channels to bandwidth allocation rules for physical ports of switches, and combines topology information to achieve end-to-end collaborative scheduling, dynamically adjusting physical link resources. This effectively solves problems such as multi-tenant bandwidth contention conflicts and mismatch between communication topology and computing power running task behavior, improves the utilization rate of computing power bandwidth resources and the accuracy of on-demand scheduling, and achieves dynamic adaptation between computing power running tasks and computing power bandwidth resources, effectively realizing a high-bandwidth network for computing power. Attached Figure Description

[0038] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0039] Figure 1 This is one of the flowcharts of a method for dynamically scheduling computing power in a high-bandwidth network of an intelligent computing center cloud platform according to an embodiment of the present invention;

[0040] Figure 2 This is the second flowchart of a method for dynamically scheduling computing power in a high-bandwidth network of an intelligent computing center cloud platform, provided in an embodiment of the present invention.

[0041] Figure 3 This is one of the structural diagrams of a device for dynamically scheduling computing power in a high-bandwidth network of an intelligent computing center cloud platform, provided in an embodiment of the present invention.

[0042] Figure 4 This is the second structural diagram of a device for dynamically scheduling computing power in a high-bandwidth network of an intelligent computing center cloud platform, provided in an embodiment of the present invention.

[0043] Figure 5 This is a structural diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0045] The “computing power” mentioned in this invention refers to: the ability of computer equipment or computing / data center to process information; the ability of computer hardware and software to work together to perform a certain computing requirement; the computing power to achieve the target result output by processing information data; and a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity, mainly providing services to society through computing power infrastructure.

[0046] The "computational power" (CP) described in this invention refers to the ability of a data center server to process data and output results. It is a comprehensive indicator of a data center's computing power, encompassing general computing power, supercomputing power, and intelligent computing power. The commonly used unit of measurement is floating-point operations per second (FLOPS, 1 EFLOPS = 10^18 FLOPS), with higher values ​​indicating stronger overall computing power. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe-2A supercomputers, 500,000 mainstream server CPUs, or 2 million mainstream laptops. The calculation formula is: CP = CP 通用 +CP 智能 +CP 超级 .

[0047] The "Network Power" (NP) mentioned in this invention refers to the performance of data transmission capability of computing facilities, which includes comprehensive capabilities such as network architecture, network bandwidth, transmission latency, intelligent management and scheduling, and involves network transmission within and between data centers. It is a comprehensive indicator for measuring network transmission scheduling capability.

[0048] The "Storage Power" (SP) described in this invention refers to the comprehensive capabilities of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon operation. It is a comprehensive indicator for measuring the data storage capacity of a data center, including external storage devices such as storage arrays and internal storage devices within servers. The commonly used unit of measurement for storage capacity is exabytes (EB, 1EB = 2^60 bytes), while the commonly used unit of measurement for performance is the number of read / write operations per second (IOPS / TB). Disaster recovery ratio is an important indicator of security and reliability.

[0049] The "computing infrastructure" mentioned in this invention refers to a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, enabling centralized computing, storage, transmission, and application of information.

[0050] The "new information infrastructure" mentioned in this invention refers to network infrastructure such as 5G networks, fiber optic broadband networks, backbone networks, international communication networks, and satellite internet; computing infrastructure such as data centers, general computing centers, intelligent computing centers, and supercomputing centers; and new technology facilities such as artificial intelligence, blockchain, and quantum computing.

[0051] The “computing power” mentioned in this invention includes: general computing power, intelligent computing power, and supercomputing power.

[0052] The "general computing power" mentioned in this invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0053] The "intelligent computing power" mentioned in this invention refers to: a computing platform deployed on a large scale based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit) for various artificial intelligence innovative applications, such as natural language processing and machine vision.

[0054] The “supercomputing power” mentioned in this invention refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, and gene analysis.

[0055] The "intelligent computing center" described in this invention refers to a facility that, through the use of large-scale heterogeneous computing resources, including general-purpose computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.), primarily provides the necessary computing power, data, and algorithms for artificial intelligence applications (such as the development, training, and inference of deep learning models). The intelligent computing center encompasses facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enablement.

[0056] The "intelligent computing center cloud platform" described in this invention refers to a cloud computing platform that integrates hardware and software resources of an intelligent computing center.

[0057] The "intelligent computing center" mentioned in this invention includes, but is not limited to, "smart computing center".

[0058] The "intelligent computing center" mentioned in this invention, also known as an artificial intelligence computing center, is a type of computing infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications, based on artificial intelligence theory and adopting an artificial intelligence computing architecture.

[0059] The "computing center" mentioned in this invention refers to a facility that is mainly composed of infrastructure such as wind, thermal, hydro, and electricity, and IT hardware and software equipment, and has computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0060] The "supercomputing center" mentioned in this invention refers to a supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters. It can provide large-scale computing, storage and network services and is widely used in aerospace, defense, oil exploration, climate modeling and genome sequencing and other application scenarios.

[0061] The “computing resources” mentioned in this invention refer to the technologies and facilities required for the development of the digital society that have the ability to compute, transmit, store and apply information, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and supporting and guaranteeing resources such as wind, fire, water and electricity.

[0062] The "model" mentioned in this invention includes, but is not limited to, "large language model" and "multimodal large model".

[0063] The "large language model" mentioned in this invention refers to a large-scale language model (LLM), which is a language model with a large number of parameters. It is designed to understand and generate human language, and is trained with a large amount of text data. It can perform a wide range of tasks, including text summarization, translation, and sentiment analysis.

[0064] The “Multimodal Large Models” mentioned in this invention refer to models that combine multimodal information such as text, images, videos, and audio for training, including but not limited to multimodal large language models.

[0065] The “computing power running task” mentioned in this invention refers to a specific workload or job executed on computing power resources that requires a certain amount of computing power support, usually involving complex data processing, numerical calculation, model training or simulation scenarios.

[0066] The "computing node" mentioned in this invention refers to the computing resources of a server / container capable of processing computing tasks.

[0067] Please see Figure 1 , Figure 1 This is one of the flowcharts of a method for dynamically scheduling computing power in a high-bandwidth network of an intelligent computing center cloud platform according to an embodiment of the present invention, such as... Figure 1 As shown, it includes the following steps:

[0068] Step S1: The computing power main control management unit sends the computing power job metadata pushed by the computing power job perception module to the computing power bandwidth scheduling unit. The computing power job metadata includes the computing power job identifier corresponding to the computing power job data packet, the bandwidth requirement information corresponding to the computing power job data packet, the priority information corresponding to the computing power job data packet, and the network topology information. The computing power job perception module is used to subscribe to the computing power job metadata of the GPU server.

[0069] In this step, such as Figure 2 As shown, a GPU cluster can include multiple GPU servers, each of which can have a compute job awareness module. The compute job awareness module obtains compute job metadata either by directly reading compute task data (such as Job ID and communication graph) issued by the scheduler from the GPU server's memory via a high-speed serial computer extension (Peripheral Component Interconnect Express, PCIe) bus, or by subscribing to the control plane and pulling information such as job topology and bandwidth requirements from the compute scheduler (such as Kubernetes) through a RESTful interface or message queue (such as Kafka). After the compute job awareness module pushes the compute job metadata to the compute master control management unit, the compute master control management unit, acting as the control plane hub, can forward the compute job metadata to the compute bandwidth scheduling unit through an internal interface (such as gRPC) for further dynamic bandwidth scheduling in subsequent steps.

[0070] The metadata for computing jobs can include a Job ID, bandwidth requirement information, priority information, and network topology (such as communication graphs and computing task topology). The Job ID serves as a unique identifier for the computing job data packet, facilitating association with subsequent scheduling operations (such as virtual channel allocation and bandwidth allocation). The bandwidth requirement information is used to calculate the bandwidth resource quota for virtual channels. The priority information guides the generation of subsequent Quality of Service (QoS) policies. The network topology information is used to determine the data transmission path for computing job data packets corresponding to different Job IDs.

[0071] Step S2: The computing power bandwidth scheduling unit determines the first scheduling data of the computing power job identifier based on the bandwidth requirement information and the priority information. The first scheduling data includes the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel and the bandwidth information of the virtual channel. Different virtual channels correspond to different priorities.

[0072] In this step, the computing power bandwidth scheduling unit allocates a dedicated virtual network interface card (VNIC) to each computing job data packet corresponding to a computing job identifier from the first virtual channel of the network interface card (NIC) based on job priority and bandwidth requirements. The first virtual channel can be multiple VNICs virtualized using SR-IOV technology, with each VNIC corresponding to an independent PCIeFunction. Different VNICs can correspond to different priorities. For example, if the priority of the computing job data packet corresponding to Job ID1 is higher than that of the computing job data packet corresponding to Job ID2, then the virtual channel corresponding to the computing job data packet indicated by Job ID1 in the first virtual channel could be VNIC1, and the virtual channel corresponding to the computing job data packet indicated by Job ID2 in the first virtual channel could be VNIC2, with VNIC1 having a higher priority than VNIC2. Furthermore, the computing power bandwidth scheduling unit determines the bandwidth information of each virtual channel according to the bandwidth requirements of the computing job data packets corresponding to different computing job identifiers. For example, if Job ID1 indicates a computing job data packet with a bandwidth requirement of 40Gbps, and Job ID2 indicates a computing job data packet with a bandwidth requirement of 20Gbps, and the virtual channel corresponding to Job ID1 can be VNIC1 and the virtual channel corresponding to Job ID2 can be VNIC2, then when determining the bandwidth information of VNIC1 and VNIC2, the ratio between the bandwidth information corresponding to VNIC1 and the bandwidth information corresponding to VNIC2 can be determined to be 2:1 based on the bandwidth requirement corresponding to Job ID1 (40Gbps) and the bandwidth requirement corresponding to Job ID2 (20Gbps). By dynamically adjusting the bandwidth allocation ratio according to the bandwidth requirement of the job data indicated by each computing job identifier, on-demand scheduling of job-level bandwidth is achieved, improving the utilization rate of cluster resources.

[0073] Step S3: The computing power bandwidth scheduling unit sends the first scheduling data to the GPU server indicated by the network topology information. The GPU server is used to perform bandwidth scheduling according to the first scheduling data.

[0074] In this step, the computing power bandwidth scheduling unit sends the first scheduling data to the GPU server indicated by the network topology information, enabling the GPU server to perform bandwidth scheduling based on the first scheduling data and allocate a dedicated receive / transmit queue to each VNIC. By performing bandwidth scheduling based on the first scheduling data within the GPU server, bandwidth resources are allocated on a job-by-job basis, rather than the traditional allocation based on virtual machines or containers, thus improving the accuracy of resource allocation.

[0075] Step S4: The link resource mapping engine sends the second scheduling data to the switch indicated by the network topology information. The switch is used to perform bandwidth scheduling according to the second scheduling data. The second scheduling data is the mapping data of the first scheduling data generated by the link resource mapping engine based on the first scheduling data.

[0076] In this step, the link resource mapping engine can work with ToR / Leaf switches to obtain topology information through LLDP and dynamically adjust port bandwidth mapping. This enables fine-grained bandwidth scheduling and end-to-end communication optimization at the job level, reduces resource contention, and avoids problems such as coarse-grained scheduling.

[0077] In this embodiment, the computing power job perception module collects computing power job metadata and transmits it to the computing power bandwidth scheduling unit, achieving accurate perception of job priority, bandwidth requirements, and network topology. The computing power main control management unit sends the computing power job metadata to the computing power bandwidth scheduling unit, which allocates virtual channels of different priorities and configures corresponding bandwidth information for computing power job data packets indicated by different computing power job identifiers. This solves the problem of coarse granularity in traditional scheduling and achieves fine-grained resource isolation and priority guarantee. Then, the computing power bandwidth scheduling unit sends the first scheduling data to the GPU server for execution, enabling the GPU server to mark traffic and bind queues according to policies, achieving on-demand bandwidth scheduling. Furthermore, the link resource mapping engine maps virtual channels to bandwidth allocation rules for physical ports of switches, and combines topology information to achieve end-to-end collaborative scheduling, dynamically adjusting physical link resources. This effectively solves problems such as multi-tenant bandwidth contention conflicts and mismatch between communication topology and computing power running task behavior, improving the utilization rate of computing power bandwidth resources and the accuracy of on-demand scheduling. It achieves dynamic adaptation between computing power running tasks and computing power bandwidth resources, effectively realizing a high-bandwidth network for computing power.

[0078] In one embodiment, after step S3 and before step S4, the method further includes:

[0079] Step S5: The link resource mapping engine obtains link layer discovery protocol information from the switch, and the link layer discovery protocol information is used for link resource mapping.

[0080] Step S6: The link resource mapping engine generates the second scheduling data based on the link layer discovery protocol information and the first scheduling data sent by the computing power bandwidth scheduling unit.

[0081] In this embodiment, the link resource mapping engine obtains Link Layer Discovery Protocol (LLDP) link information from the ToR switch, such as real-time information on the physical connection topology and port capabilities between the switch and the server, providing a foundation for subsequent mapping of virtual resources to physical links. For example, SNMP or some ToR switch management interfaces can be used to periodically or trigger the acquisition of LLDP link information from the switch. Then, the link resource mapping engine combines the LLDP link information with the first scheduling data to accurately map virtual channels to physical ports, generating switch flow table configuration instructions to achieve collaborative optimization of logical scheduling strategies and physical network resources. In this way, the bandwidth scheduling strategy can fully consider physical link characteristics (such as bandwidth limits and link quality), improving the rationality of end-to-end path planning; by dynamically discovering network changes through LLDP, the mapping strategy can be adjusted in real time, enhancing the system's adaptability to network topology changes; at the same time, the binding of physical link status with virtual channels ensures that high-priority job traffic can be guided to the optimal path, reducing congestion risks and improving the communication quality and bandwidth guarantee capability of critical services.

[0082] In one embodiment, step S2 includes:

[0083] Step S21: The computing power bandwidth scheduling unit sets a differential service code point for the job data indicated by the computing power job identifier according to the priority information. The differential service code point is used to determine the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel.

[0084] Step S22: When the differential service code point indicates that the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel is the first channel, the computing power bandwidth scheduling unit determines the first bandwidth information of the first channel according to the bandwidth demand information. The first scheduling data includes the differential service code point and the first bandwidth information.

[0085] In this embodiment, by setting different Differentiated Services Code Points (DSCPs) for data packets indicating computing power jobs with different job identifiers, job priorities are directly mapped to network layer QoS tags. This allows GPU servers and switches to perform differentiated processing of traffic with different priorities based on DSCP values ​​(e.g., prioritizing the forwarding of traffic with high DSCP values), achieving fine-grained priority scheduling. After determining the DSCP value corresponding to a virtual channel, dedicated bandwidth information is configured for that channel according to bandwidth requirements (e.g., the first bandwidth information for the first channel), binding logical priorities with the bandwidth resources of computing power. This ensures low-latency transmission for high-priority jobs and avoids resource contention conflicts in multi-tenant scenarios through bandwidth quotas. This enables GPU servers to accurately identify the virtual channel to which traffic belongs and perform bandwidth scheduling based on DSCP values. Combined with the switch's priority processing of DSCP values, an end-to-end priority awareness and bandwidth guarantee link is formed, improving the stability of the network used by important tenants of the intelligent computing center cloud platform and the stability of access to some critical computing resources.

[0086] In one embodiment, step S2 further includes:

[0087] Step S23: When the differential service code point indicates that the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel is the second channel, the computing power bandwidth scheduling unit generates a scheduling instruction for the second channel. The scheduling instruction is used to instruct the computing power job data packet to be allocated to the second channel. The second bandwidth information of the second channel is a fixed bandwidth value. The first scheduling data includes the differential service code point and the scheduling instruction. The priority of the second channel is greater than the priority of the first channel.

[0088] In this embodiment, a specific virtual channel can be set for the computing job data packet indicated by the highest priority computing job identifier. Specifically, the computing bandwidth scheduling unit sets a Differential Service Code Point (DSCP) value (e.g., DSCP=46) for the job data indicated by the computing job identifier (e.g., Job ID0) based on priority information, indicating that the job data indicated by Job ID0 has the highest priority. Therefore, DSCP=46 indicates that the virtual channel corresponding to the computing job data packet indicated by Job ID0 in the first virtual channel is the second channel, and the second channel can be the virtual channel with the highest priority. In this way, by allocating a second channel to high-priority computing job data packets and setting a fixed bandwidth value, a differentiated bandwidth guarantee mechanism is constructed.

[0089] Specifically, when the Differential Service Code Point (DSCP) indicates a high-priority job, a dedicated scheduling instruction is generated to direct its traffic to a dedicated second channel. This ensures that critical services receive stable bandwidth resources and avoids bandwidth contention caused by bursts of low-priority traffic. Simultaneously, by setting a fixed bandwidth value (e.g., 40Gbps) rather than dynamically allocating bandwidth proportionally, deterministic network services are provided for high-priority jobs, meeting the stringent bandwidth requirements of latency-sensitive operations such as gradient synchronization in AI training. The design of the second channel having a higher priority than the first channel allows the switch to prioritize forwarding high-priority traffic during congestion, further reducing packet loss and latency for critical services. This hierarchical scheduling of multi-priority jobs ensures the stable operation of high-priority computing tasks while allowing low-priority computing tasks to flexibly utilize remaining bandwidth, improving the overall system performance and resource utilization in heterogeneous workload deployment scenarios.

[0090] In one embodiment, step S22 includes:

[0091] Step S221: The computing power bandwidth scheduling unit determines the first bandwidth information of the first channel based on the proportion of the bandwidth value corresponding to the bandwidth demand information to the total bandwidth value corresponding to the total bandwidth demand information.

[0092] In this embodiment, taking Job A with a bandwidth requirement of 40Gbps, Job B with a bandwidth requirement of 30Gbps, and Job C with a bandwidth requirement of 30Gbps as examples, and the total bandwidth can be 100Gbps, of which 20Gbps is a fixed bandwidth value, the total allocable bandwidth is 80Gbps. First, the computing power bandwidth scheduling unit calculates the proportion of the bandwidth value corresponding to the bandwidth requirement information of each Job ID within the total bandwidth value corresponding to the total bandwidth requirement information: Job A accounts for 40Gbps / (40+30+30)Gbps=40%; Job B accounts for 30Gbps / 100Gbps=30%; Job C accounts for 30Gbps / 100Gbps=30%. Based on this, it can be determined that: the first bandwidth information of the first channel corresponding to Job A is 80Gbps×40%=32Gbps; the first bandwidth information of the first channel corresponding to Job B is 80Gbps×30%=24Gbps; and the first bandwidth information of the first channel corresponding to Job C is 80Gbps×30%=24Gbps.

[0093] If Job C releases 10Gbps of demand midway through (i.e., the new demand is 20Gbps), the total demand becomes 90Gbps, and the proportions can be recalculated:

[0094] Job A accounts for approximately 44.4% of the total, and the bandwidth information of the first channel corresponding to Job A is approximately 35.5Gbps (80Gbps × 44.4%).

[0095] Job B accounts for approximately 33.3% of the total, and the bandwidth information of the first channel corresponding to Job B is approximately 80Gbps × 33.3% ≈ 26.6Gbps.

[0096] Job C accounts for approximately 22.2% of the total, and the bandwidth information of the first channel corresponding to Job C is approximately 80Gbps × 22.2% ≈ 17.7Gbps.

[0097] This approach allows for proportional compression of bandwidth for individual jobs when total bandwidth is insufficient, preventing excessive preemption by some jobs and starvation of others, thus improving overall resource utilization. It ensures bandwidth allocation is directly proportional to actual job demand, guaranteeing higher bandwidth requirements for jobs and allocating less bandwidth to lower-priority jobs as needed, reducing resource waste. Furthermore, the proportional allocation mechanism provides relatively fair bandwidth guarantees for jobs of different priorities, ensuring the basic needs of high-priority jobs while allowing lower-priority jobs to use resources reasonably. When total bandwidth or job demand changes, the proportion is automatically recalculated and the allocation adjusted without manual intervention, enhancing system flexibility. It is suitable for multi-tenant shared cluster scenarios, dynamically balancing bandwidth allocation based on real-time job demands. It achieves dynamic adaptation between computing power and the bandwidth resources of the computing power, effectively realizing a high-bandwidth network for the computing power.

[0098] In one embodiment, after step S3, the method further includes:

[0099] Step S7: Upon receiving the first scheduling data, the GPU server calls the multi-channel transceiver module to generate the first virtual channel. Each virtual channel in the first virtual channel has an independent caching and queue management mechanism.

[0100] Step S8: The GPU server performs bandwidth scheduling based on the first virtual channel and the first scheduling data.

[0101] In this embodiment, a first virtual channel with independent caching and queue management mechanisms is generated on the GPU server side through a multi-channel transceiver module, transforming the logical scheduling strategy into physically executable network resources; then, bandwidth scheduling is performed based on this virtual channel. The independent caching of each virtual channel avoids cache contention between jobs of different priorities, ensuring that high-priority job traffic is not blocked by low-priority traffic, thus improving the stability of network usage for important tenants of the intelligent computing center cloud platform, as well as the stability of access to some critical computing resources. Furthermore, it allows for the configuration of dedicated QoS policies (such as PFC / DCQCN) for different virtual channels, meeting the low-latency requirements such as gradient synchronization in AI training, and enabling the allocation of dedicated virtual channels for different job types (such as training / inference). For example, low-latency queues can be allocated for real-time inference traffic, while high-bandwidth queues can be allocated for batch training, improving overall efficiency in mixed load scenarios.

[0102] Furthermore, by incorporating hardware offloading technologies such as Smart NICs, the creation and management of virtual channels can be directly executed by the network card firmware, reducing CPU overhead and accelerating data paths. When congestion occurs in a virtual channel, an independent caching mechanism can prevent it from affecting other channels, improving system fault tolerance. In this way, through the visualization and fine-grained management of virtual channels on the server side, bandwidth scheduling strategies are transformed into physically executable flow control behaviors, improving end-to-end network performance and resource utilization.

[0103] In some embodiments, such as Figure 3 As shown, the method for dynamically scheduling computing power on a high-bandwidth network in an intelligent computing center cloud platform may include the following steps:

[0104] The computing job awareness module subscribes to computing job metadata to achieve accurate awareness of job priority, bandwidth requirements, and network topology;

[0105] The computing power control and management unit analyzes computing power job metadata. The computing power control and management unit has the ability to interface with the computing power scheduler, and receives policies such as computing power running task topology and bandwidth requirements. It can send computing power job metadata to the computing power bandwidth scheduling unit.

[0106] The computing power bandwidth scheduling unit adjusts the QoS policy according to job priority, and sends different job data packets into different dscp values ​​according to the priority of computing power running tasks. That is, they enter the VNIC channel allocated by the multi-channel transceiver module to the GPU server in the GPU cluster. The bandwidth of each VNIC channel is dynamically allocated according to the bandwidth requirement of each job, which solves the problem of coarse scheduling granularity in traditional scheduling and realizes fine-grained resource isolation and priority guarantee.

[0107] The link resource mapping engine coordinates dynamic bandwidth allocation with the TOR / Leaf switch. By dynamically adjusting physical link resources by allocating bandwidth to each VSP channel, it can effectively solve problems such as multi-tenant bandwidth contention conflicts and mismatch between communication topology and computing power operation task behavior.

[0108] The multi-channel transceiver module supports concurrent communication of multiple independent virtual channels (VNIC and VSP). Each channel has an independent buffer and queue management mechanism. The independent buffer of each virtual channel avoids buffer contention between jobs of different priorities.

[0109] The RDMA enhancement engine optimizes communication performance for large models and supports the RoCEv2 protocol.

[0110] This enables on-demand bandwidth scheduling, allowing for precise end-to-end scheduling from logical resources to physical links. This improves the network efficiency of heterogeneous computing power clusters, enhances the stability of the network used by important tenants of the intelligent computing center cloud platform, and improves the stability of access to some critical computing resources.

[0111] Please see Figure 4 , Figure 4 This is a second structural diagram of a device for dynamically scheduling computing power in a high-bandwidth network of an intelligent computing center cloud platform, as provided in an embodiment of the present invention. Figure 4 As shown, the device 400 for dynamically scheduling computing power in a high-bandwidth network of an intelligent computing center cloud platform includes: a computing power main control management unit 401, a computing power bandwidth scheduling unit 402, a link resource mapping engine 403, and a computing power job perception module.

[0112] The computing power master control management unit 401 is used to send the computing power job metadata pushed by the computing power job perception module to the computing power bandwidth scheduling unit. The computing power job metadata includes the computing power job identifier corresponding to the computing power job data packet, the bandwidth requirement information corresponding to the computing power job data packet, the priority information corresponding to the computing power job data packet, and the network topology information. The computing power job perception module is used to subscribe to the computing power job metadata of the GPU server.

[0113] The computing power bandwidth scheduling unit 402 is used to determine the first scheduling data of the computing power job identifier according to the bandwidth requirement information and the priority information. The first scheduling data includes the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel and the bandwidth information of the virtual channel. Different virtual channels correspond to different priorities.

[0114] The computing power bandwidth scheduling unit 402 is further configured to send the first scheduling data to the GPU server indicated by the network topology information, and the GPU server is configured to perform bandwidth scheduling according to the first scheduling data;

[0115] The link resource mapping engine 403 is used to send the second scheduling data to the switch indicated by the network topology information. The switch is used to perform bandwidth scheduling according to the second scheduling data. The second scheduling data is the mapping data of the first scheduling data generated by the link resource mapping engine according to the first scheduling data.

[0116] In one embodiment, the link resource mapping engine is further configured to obtain link layer discovery protocol information from the switch, and the link layer discovery protocol information is used for link resource mapping.

[0117] The link resource mapping engine is further configured to generate the second scheduling data based on the link layer discovery protocol information and the first scheduling data sent by the computing power bandwidth scheduling unit.

[0118] In one embodiment, the computing power bandwidth scheduling unit 402 is specifically used for:

[0119] The computing power bandwidth scheduling unit sets a differential service code point for the job data indicated by the computing power job identifier according to the priority information. The differential service code point is used to determine the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel.

[0120] When the differential service code point indicates that the virtual channel corresponding to the computing job data packet indicated by the computing job identifier in the first virtual channel is the first channel, the computing bandwidth scheduling unit determines the first bandwidth information of the first channel according to the bandwidth demand information. The first scheduling data includes the differential service code point and the first bandwidth information.

[0121] In one embodiment, the computing power bandwidth scheduling unit 402 is further specifically used for:

[0122] When the differential service code point indicates that the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel is the second channel, the computing power bandwidth scheduling unit generates a scheduling instruction for the second channel. The scheduling instruction is used to instruct the computing power job data packet to be allocated to the second channel. The second bandwidth information of the second channel is a fixed bandwidth value. The first scheduling data includes the differential service code point and the scheduling instruction. The priority of the second channel is greater than the priority of the first channel.

[0123] In one embodiment, the computing power bandwidth scheduling unit 402 is further specifically used for:

[0124] The computing power bandwidth scheduling unit determines the first bandwidth information of the first channel based on the proportion of the bandwidth value corresponding to the bandwidth demand information to the total bandwidth value corresponding to the total bandwidth demand information.

[0125] In one embodiment, the apparatus further includes a GPU server:

[0126] Upon receiving the first scheduling data, the GPU server invokes the multi-channel transceiver module to generate the first virtual channel. Each virtual channel in the first virtual channel has an independent caching and queue management mechanism.

[0127] The GPU server performs bandwidth scheduling based on the first virtual channel and the first scheduling data.

[0128] The apparatus for dynamically scheduling computing power in a high-bandwidth network of an intelligent computing center cloud platform provided in this embodiment of the invention is a process of each embodiment of the method for dynamically scheduling computing power in a high-bandwidth network of an intelligent computing center cloud platform. The technical features are one-to-one and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0129] It should be noted that the device for dynamically scheduling computing power in the intelligent computing center cloud platform in the embodiments of the present invention can be a device, or it can be a component, integrated circuit, or chip in an electronic device.

[0130] This invention also provides an electronic device, see [link to relevant documentation]. Figure 5 , Figure 5 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. The electronic device includes a memory 501, a processor 502, and a program or instructions stored in the memory 501 that run on the processor 502. When the program or instructions are executed by the processor 502, they can achieve the following: Figure 1 The steps in the corresponding intelligent computing center cloud platform's method for dynamically scheduling computing power in a high-bandwidth network, and achieving the same beneficial effects, will not be elaborated here.

[0131] The processor 502 can be a CPU, ASIC, FPGA, or GPU.

[0132] Those skilled in the art will understand that all or part of the steps of the above-described method embodiment for dynamically scheduling computing power in a high-bandwidth network for intelligent computing center cloud platforms can be implemented by hardware related to program instructions, and the program can be stored in a readable medium.

[0133] This invention also provides a readable storage medium storing a computer program, which, when executed by a processor, can perform the above-described functions. Figure 1Any step in the corresponding method embodiment for dynamically scheduling computing power in a high-bandwidth network of an intelligent computing center cloud platform can achieve the same technical effect, and will not be described again here to avoid repetition. The storage medium mentioned includes, for example, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0134] The present invention also provides a computer program product, including computer instructions that, when executed by a processor, implement the above-described... Figure 1 The various processes of the corresponding intelligent computing center cloud platform's method for dynamically scheduling computing power in a high-bandwidth network are described in the embodiment, and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0135] In the embodiments of this invention, the terms "first," "second," etc., are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices. Additionally, the use of "and / or" in this application indicates at least one of the connected objects, such as A and / or B and / or C, representing seven possibilities: A alone, B alone, C alone, both A and B present, both B and C present, both A and C present, and A, B, and C present.

[0136] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0137] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or second terminal device, etc.) to execute the methods of the various embodiments of this application.

[0138] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A method for dynamically scheduling computing power in a high-bandwidth network of an intelligent computing center cloud platform, characterized in that, The method includes: Step S1: The computing power main control management unit sends the computing power job metadata pushed by the computing power job perception module to the computing power bandwidth scheduling unit. The computing power job metadata includes the computing power job identifier corresponding to the computing power job data packet, the bandwidth requirement information corresponding to the computing power job data packet, the priority information corresponding to the computing power job data packet, and the network topology information. The computing power job perception module is used to subscribe to the computing power job metadata of the GPU server. Step S2: The computing power bandwidth scheduling unit determines the first scheduling data of the computing power job identifier based on the bandwidth requirement information and the priority information. The first scheduling data includes the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel and the bandwidth information of the virtual channel. Different virtual channels correspond to different priorities. The computing power bandwidth scheduling unit determines the bandwidth information of each virtual channel according to the bandwidth requirement of the computing power job data packet indicated by different computing power job identifiers. Step S2 includes: Step S21: The computing power bandwidth scheduling unit sets a differential service code point for the job data indicated by the computing power job identifier according to the priority information. The differential service code point is used to determine the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel. Step S22: When the differential service code point indicates that the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel is the first channel, the computing power bandwidth scheduling unit determines the first bandwidth information of the first channel according to the bandwidth demand information. The first scheduling data includes the differential service code point and the first bandwidth information. Step S3: The computing power bandwidth scheduling unit sends the first scheduling data to the GPU server indicated by the network topology information. The GPU server is used to perform bandwidth scheduling according to the first scheduling data. Step S4: The link resource mapping engine sends the second scheduling data to the switch indicated by the network topology information. The switch is used to perform bandwidth scheduling according to the second scheduling data. The second scheduling data is the mapping data of the first scheduling data generated by the link resource mapping engine based on the first scheduling data.

2. The method as described in claim 1, characterized in that, After step S3 and before step S4, the method further includes: Step S5: The link resource mapping engine obtains link layer discovery protocol information from the switch, and the link layer discovery protocol information is used for link resource mapping. Step S6: The link resource mapping engine generates the second scheduling data based on the link layer discovery protocol information and the first scheduling data sent by the computing power bandwidth scheduling unit.

3. The method as described in claim 1, characterized in that, Step S2 further includes: Step S23: When the differential service code point indicates that the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel is the second channel, the computing power bandwidth scheduling unit generates a scheduling instruction for the second channel. The scheduling instruction is used to instruct the computing power job data packet to be allocated to the second channel. The second bandwidth information of the second channel is a fixed bandwidth value. The first scheduling data includes the differential service code point and the scheduling instruction. The priority of the second channel is greater than the priority of the first channel.

4. The method as described in claim 1, characterized in that, Step S22 includes: Step S221: The computing power bandwidth scheduling unit determines the first bandwidth information of the first channel based on the proportion of the bandwidth value corresponding to the bandwidth demand information to the total bandwidth value corresponding to the total bandwidth demand information.

5. The method according to any one of claims 1 to 4, characterized in that, After step S3, the method further includes: Step S7: Upon receiving the first scheduling data, the GPU server calls the multi-channel transceiver module to generate the first virtual channel. Each virtual channel in the first virtual channel has an independent caching and queue management mechanism. Step S8: The GPU server performs bandwidth scheduling based on the first virtual channel and the first scheduling data.

6. A device for dynamically scheduling computing power in a high-bandwidth network of an intelligent computing center cloud platform, characterized in that, The device includes a computing power main control management unit, a computing power bandwidth scheduling unit, a link resource mapping engine, and a computing power job perception module; The computing power master control management unit is used to send the computing power job metadata pushed by the computing power job perception module to the computing power bandwidth scheduling unit. The computing power job metadata includes the computing power job identifier corresponding to the computing power job data packet, the bandwidth requirement information corresponding to the computing power job data packet, the priority information corresponding to the computing power job data packet, and the network topology information. The computing power job perception module is used to subscribe to the computing power job metadata of the GPU server. The computing power bandwidth scheduling unit is used to determine first scheduling data of the computing power job identifier based on the bandwidth requirement information and the priority information. The first scheduling data includes the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel and the bandwidth information of the virtual channel. Different virtual channels correspond to different priorities. The computing power bandwidth scheduling unit determines the bandwidth information of each virtual channel according to the bandwidth requirement of the computing power job data packet indicated by different computing power job identifiers. The computing power bandwidth scheduling unit is specifically used for: The computing power bandwidth scheduling unit sets a differential service code point for the job data indicated by the computing power job identifier according to the priority information. The differential service code point is used to determine the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel. When the differential service code point indicates that the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel is the first channel, the computing power bandwidth scheduling unit determines the first bandwidth information of the first channel according to the bandwidth requirement information. The first scheduling data includes the differential service code point and the first bandwidth information. The computing power bandwidth scheduling unit is further configured to send the first scheduling data to the GPU server indicated by the network topology information, and the GPU server is configured to perform bandwidth scheduling according to the first scheduling data; The link resource mapping engine is used to send the second scheduling data to the switch indicated by the network topology information. The switch is used to perform bandwidth scheduling according to the second scheduling data. The second scheduling data is the mapping data of the first scheduling data generated by the link resource mapping engine based on the first scheduling data.

7. An electronic device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of a method for dynamically scheduling computing power in a high-bandwidth network of an intelligent computing center cloud platform as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method for dynamically scheduling computing power in a high-bandwidth network of an intelligent computing center cloud platform as described in any one of claims 1 to 5.

9. A computer program product, characterized in that, The method includes computer instructions that, when executed by a processor, implement the steps of a method for dynamically scheduling computing power in a high-bandwidth network of an intelligent computing center cloud platform as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Service transmission method and device and storage medium

    CN115734291A

  • Calculation power scheduling method and system

    CN119127473A