Method and device for dynamically scheduling computing power high-bandwidth network by intelligent computing center cloud platform

Through the dynamic scheduling computing power method of the intelligent computing center cloud platform, the computing power main control management unit, computing power bandwidth scheduling unit and link resource mapping engine are used to achieve fine-grained resource isolation and priority guarantee, solving the on-demand problem of computing power bandwidth scheduling in the intelligent computing center, and improving the stability and resource utilization of the network.

CN120343086AActive Publication Date: 2025-07-18DATACANVAS LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510803700.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-07-18
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

In the existing intelligent computing center, it is difficult to achieve on-demand scheduling of computing power, resulting in a competition for network bandwidth, affecting the stability of tenants' network usage and the access stability of key computing resources.

Method used

Through the coordinated work of the computing power main control management unit, the computing power bandwidth scheduling unit and the link resource mapping engine, the bandwidth resources of computing power operations are dynamically scheduled, fine-grained resource isolation and priority guarantee are achieved, and end-to-end collaborative scheduling is carried out in combination with network topology information.

Benefits of technology

It improves the bandwidth utilization rate of computing power resources and the accuracy of on-demand scheduling, solves the problems of multi-tenant bandwidth preemption conflict and communication topology mismatch, and ensures the stability of high-bandwidth networks and the access stability of key computing power resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343086A_ABST
    Figure CN120343086A_ABST
Patent Text Reader

Abstract

The invention provides a method and a device for dynamically scheduling a computing power high-bandwidth network by an intelligent computing center cloud platform, and relates to the technical field of computing power infrastructure, the method comprises the following steps: a computing power master control management unit sends computing power operation metadata pushed by a computing power operation sensing module to a computing power bandwidth scheduling unit; the computing power bandwidth scheduling unit determines first scheduling data of the computing power job identifier according to the bandwidth demand information and the priority information; the computing power bandwidth scheduling unit sends the first scheduling data to a GPU server indicated by the network topology information; and the link resource mapping engine sends the second scheduling data to the switch indicated by the network topology information. Dynamic adaptation between the computing power operation task and the bandwidth resource of the computing power is realized, and a high-bandwidth network of the computing power is effectively realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers and computing power infrastructure, and particularly relates to a method and device for a high-bandwidth network for dynamically scheduling computing power in an intelligent computing center cloud platform. Background Art

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.

[0003] An "intelligent computing center" refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power and intelligent computing power, and mainly provides the required computing power, data and algorithms for artificial intelligence applications (such as scenarios of artificial intelligence deep learning model development, model training and model inference, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.

[0004] The "intelligent computing center" includes but is not limited to the "intelligent computing center".

[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure based on artificial intelligence theory, adopting an artificial intelligence computing architecture, and providing computing power services, data services and algorithm services required for artificial intelligence applications.

[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers". It is the ability of computer devices or computing / data centers to process information. It is the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement. It is the computing ability to achieve the output of the target result by processing information data. It is a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.

[0007] In scenarios such as artificial intelligence (AI) training, scientific computing, and large-scale data processing, higher requirements are placed on high-bandwidth and low-latency data exchange capabilities. Traditional network interface cards (NICs) generally have a bandwidth of 10 Gbps / 25 Gbps / 40 Gbps. In some high-end data centers, intelligent network interface cards (SmartNICs) with a specification of 100 Gbps or higher have been deployed to support more complex virtualization and acceleration capabilities. In a typical SmartNIC solution, such as a card based on the Mellanox ConnectX series, network performance is enhanced through built-in protocol stack acceleration mechanisms such as remote direct memory access (RDMA) and the Intel Data Plane Development Kit (DPDK). Its structure generally includes components such as a main controller, a DMA engine, a network protocol processing module, a PCIe interface, and SR-IOV virtual channels, and cooperates with a driver to achieve efficient interaction with the host operating system. However, as the scale of AI models expands and inter-cluster communication becomes frequent, the existing SmartNIC solutions are still prone to bandwidth contention, affecting the stability of network usage by important tenants in the cloud platform of the intelligent computing center and the stability of access to some key computing power resources.

[0008] It can be seen that since the emergence of intelligent computing centers, how to achieve on-demand scheduling of computing power bandwidth to realize a high-bandwidth network has become an urgent problem to be solved. Summary of the Invention

[0009] Embodiments of the present invention provide a method and device for a high-bandwidth network for dynamically scheduling computing power in a cloud platform of an intelligent computing center to solve the problem of how to achieve on-demand scheduling of computing power bandwidth to realize a high-bandwidth network since the emergence of intelligent computing centers.

[0010] To solve the above problems, the present invention is implemented as follows: In a first aspect, an embodiment of the present invention provides a method for a high-bandwidth network for dynamically scheduling computing power in a cloud platform of an intelligent computing center, and the method includes: Step S1, a computing power master management unit sends the computing power job metadata pushed by a computing power job perception module to a computing power bandwidth scheduling unit. The computing power job metadata includes a computing power job identifier corresponding to a computing power job data packet, bandwidth requirement information corresponding to the computing power job data packet, priority information corresponding to the computing power job data packet, and network topology information. The computing power job perception module is used to subscribe to the computing power job metadata of a GPU server; Step S2: The computing power bandwidth scheduling unit determines first scheduling data of the computing power job identifier according to the bandwidth demand information and the priority information. The first scheduling data includes the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel and the bandwidth information of the virtual channel, and different virtual channels correspond to different priorities. Step S3: The computing power bandwidth scheduling unit sends the first scheduling data to the GPU server indicated by the network topology information, and the GPU server is used to perform bandwidth scheduling according to the first scheduling data. Step S4: The link resource mapping engine sends second scheduling data to the switch indicated by the network topology information, and the switch is used to perform bandwidth scheduling according to the second scheduling data. The second scheduling data is the mapping data of the first scheduling data generated by the link resource mapping engine according to the first scheduling data.

[0011] In one embodiment, after the step S3 and before the step S4, the method further includes: Step S5: The link resource mapping engine obtains link layer discovery protocol information from the switch, and the link layer discovery protocol information is used for link resource mapping. Step S6: The link resource mapping engine generates the second scheduling data according to the link layer discovery protocol information and the first scheduling data sent by the computing power bandwidth scheduling unit.

[0012] In one embodiment, the step S2 includes: Step S21: The computing power bandwidth scheduling unit sets a differentiated services code point for the job data indicated by the computing power job identifier according to the priority information, and the differentiated services code point is used to determine the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel. Step S22: When the differentiated services code point indicates that the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel is the first channel, the computing power bandwidth scheduling unit determines the first bandwidth information of the first channel according to the bandwidth demand information, and the first scheduling data includes the differentiated services code point and the first bandwidth information.

[0013] In one embodiment, the step S2 further includes: Step S23: When the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the differential service code point is the second channel in the first virtual channel, the computing power bandwidth scheduling unit generates a scheduling instruction for the second channel. The scheduling instruction is used to indicate that the computing power job data packet is allocated to the second channel. The second bandwidth information of the second channel is a fixed bandwidth value. The first scheduling data includes the differential service code point and the scheduling instruction. The priority of the second channel is higher than that of the first channel.

[0014] In one embodiment, step S22 includes: Step S221: The computing power bandwidth scheduling unit determines the first bandwidth information of the first channel according to the ratio of the total available bandwidth and the bandwidth value corresponding to the bandwidth requirement information in the total bandwidth value corresponding to the total bandwidth requirement information.

[0015] In one embodiment, after step S3, the method further includes: Step S7: When the GPU server receives the first scheduling data, it calls the multi-channel transceiver module to generate the first virtual channel. Each virtual channel in the first virtual channel has an independent cache and queue management mechanism; Step S8: The GPU server performs bandwidth scheduling based on the first virtual channel according to the first scheduling data.

[0016] In a second aspect, an embodiment of the present invention further provides a high-bandwidth network device for dynamically scheduling computing power in an intelligent computing center cloud platform. The device includes a computing power main control management unit, a computing power bandwidth scheduling unit, a link resource mapping engine, and a computing power job perception module; The computing power main control management unit is configured to send the computing power job metadata pushed by the computing power job perception module to the computing power bandwidth scheduling unit. The computing power job metadata includes the computing power job identifier corresponding to the computing power job data packet, the bandwidth requirement information corresponding to the computing power job data packet, the priority information corresponding to the computing power job data packet, and the network topology information. The computing power job perception module is used to subscribe to the computing power job metadata of the GPU server; The computing power bandwidth scheduling unit is configured to determine the first scheduling data of the computing power job identifier according to the bandwidth requirement information and the priority information. The first scheduling data includes the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel and the bandwidth information of the virtual channel. Different virtual channels have different priorities; The computing power bandwidth scheduling unit is further configured to send the first scheduling data to the GPU server indicated by the network topology information, and the GPU server is configured to perform bandwidth scheduling according to the first scheduling data; The link resource mapping engine is configured to send second scheduling data to the switch indicated by the network topology information, and the switch is configured to perform bandwidth scheduling according to the second scheduling data, where the second scheduling data is mapping data of the first scheduling data generated by the link resource mapping engine based on the first scheduling data.

[0017] In a third aspect, the present invention further provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps in the method for a high-bandwidth network for dynamically scheduling computing power of the intelligent computing center cloud platform as described in the first aspect above are implemented.

[0018] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the method for a high-bandwidth network for dynamically scheduling computing power of the intelligent computing center cloud platform as described in the first aspect above are implemented.

[0019] In a fifth aspect, the present invention further provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, the steps in the method for a high-bandwidth network for dynamically scheduling computing power of the intelligent computing center cloud platform as described in the first aspect above are implemented.

[0020] In the embodiments of the present invention, the computing power job perception module collects computing power job metadata and transmits it to the computing power bandwidth scheduling unit to achieve precise perception of job priority, bandwidth requirements, and network topology; the computing power main control management unit sends the computing power job metadata to the computing power bandwidth scheduling unit, and the computing power bandwidth scheduling unit assigns virtual channels with different priorities to the computing power job data packets indicated by different computing power job identifiers and configures corresponding bandwidth information, solving the problem of rough traditional scheduling granularity and achieving fine-grained resource isolation and priority guarantee; then, the computing power bandwidth scheduling unit sends the first scheduling data to the GPU server for execution, enabling the GPU server side to mark traffic and bind queues according to the policy, achieving on-demand scheduling of bandwidth; and the link resource mapping engine maps the virtual channels to the bandwidth allocation rules of the switch physical ports, combines the topology information to achieve end-to-end collaborative scheduling, and dynamically adjusts the physical link resources, which can effectively solve problems such as multi-tenant bandwidth preemption conflicts and mismatches between communication topologies and computing power operation task behaviors, improving the utilization rate of computing power bandwidth resources and the accuracy of on-demand scheduling, realizing the dynamic adaptation of computing power operation tasks and computing power bandwidth resources, and effectively realizing a high-bandwidth network for computing power. Brief Description of the Drawings

[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the following described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0022] Figure 1 It is one of the flowcharts of a method for a high - bandwidth network for dynamic scheduling of computing power in an intelligent computing center cloud platform provided by an embodiment of the present invention; Figure 2 It is the second flowchart of a method for a high - bandwidth network for dynamic scheduling of computing power in an intelligent computing center cloud platform provided by an embodiment of the present invention; Figure 3 It is one of the structural diagrams of a device for a high - bandwidth network for dynamic scheduling of computing power in an intelligent computing center cloud platform provided by an embodiment of the present invention; Figure 4 It is the second structural diagram of a device for a high - bandwidth network for dynamic scheduling of computing power in an intelligent computing center cloud platform provided by an embodiment of the present invention; Figure 5 It is the structural diagram of an electronic device provided by an embodiment of the present invention. Detailed Embodiments

[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention fall within the scope of protection of the present invention.

[0024] The "computing power" described in the present invention refers to: the ability of a computer device or a computing / data center to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to output a target result by processing information data, a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly providing services to society through computing power infrastructure.

[0025] The "Computational Power (CP)" described in the present invention refers to: the ability of a data center server to process data and output results, which is a comprehensive indicator for measuring the computing power of a data center and includes general computing power, supercomputing power, and intelligent computing power. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1 EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing power. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP 通用 + CP 智能 + CP 超级 。

[0026] The "Network Power (NP)" described in the present invention refers to: the performance of the data transmission ability of computing power facilities, which is a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., and involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling ability.

[0027] The "Storage Power (SP)" described in the present invention refers to: the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon, which is a comprehensive indicator for measuring the data storage ability of a data center and includes external storage devices such as storage arrays and server internal storage devices. The commonly used measurement unit for storage capacity is exabyte (EB, 1 EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read and write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.

[0028] The "computing power infrastructure" described in the present invention refers to: a new type of information infrastructure that integrates information computing power, network carrying power, and data storage power, and can realize the centralized computing, storage, transmission, and application of information.

[0029] The "new type of information infrastructure" described in the present invention mainly includes network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, and satellite Internet, computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, and supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.

[0030] The "computing power" described in the present invention includes: general computing power, intelligent computing power, and supercomputing power.

[0031] The "general computing power" described in the present invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0032] The "intelligent computing power" described in the present invention refers to: for various artificial intelligence innovation applications, a computing platform deployed on a large scale based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit), such as natural language processing, machine vision, and so on.

[0033] The "super computing power" described in the present invention refers to: mainly the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.

[0034] The "intelligent computing center" described in the present invention refers to: a facility that provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.

[0035] The "intelligent computing center cloud platform" described in the present invention refers to: a cloud computing platform that comprehensively serves based on the hardware resources and software resources of the intelligent computing center.

[0036] The "intelligent computing center" described in the present invention includes but is not limited to the "intelligent computing center".

[0037] The "intelligent computing center" described in the present invention, that is, the artificial intelligence computing center, is a type of computing power infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications based on artificial intelligence theory and using an artificial intelligence computing architecture.

[0038] The "computing power center" described in the present invention refers to: a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, which has computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0039] The "supercomputing center" referred to in the present invention means: namely, a supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters, capable of providing functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.

[0040] The "computing power resources" referred to in the present invention means: technologies and facilities with information computing, transmission, storage, and application capabilities required for the development of the digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.

[0041] The "models" referred to in the present invention include but are not limited to "large language models" and "multimodal large models".

[0042] The "large language model" referred to in the present invention means a large language model (LLM), which is a language model with a relatively large number of parameters, aiming to understand and generate human language, trained with a large amount of text data, and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.

[0043] The "multimodal large models" (Multimodal Large Models) referred to in the present invention means: models trained by combining multimodal information such as text, images, videos, and audio, including but not limited to multimodal large language models.

[0044] The "computing power operation task" referred to in the present invention means: specific workloads or jobs that are executed on computing power resources and require a certain amount of computing power support, usually involving scenarios such as complex data processing, numerical calculations, model training, or simulation.

[0045] The "computing power node" referred to in the present invention means: the computing resources of servers / containers that can process computing tasks.

[0046] Please refer to Figure 1 , Figure 1 is one of the flowcharts of a method for a high-bandwidth network for dynamically scheduling computing power in an intelligent computing center cloud platform provided by an embodiment of the present invention. As Figure 1 shown, it includes the following steps: Step S1: The computing power master control unit sends the computing power job metadata pushed by the computing power job perception module to the computing power bandwidth scheduling unit. The computing power job metadata includes the computing power job identifier corresponding to the computing power job data packet, the bandwidth requirement information corresponding to the computing power job data packet, the priority information corresponding to the computing power job data packet, and the network topology information. The computing power job perception module is used to subscribe to the computing power job metadata of the GPU server. In this step, as Figure 2 shown, the GPU cluster may include multiple GPU servers, and each GPU server may be provided with a computing power job perception module. The computing power job perception module can obtain the computing power job metadata by directly reading the computing power operation task data (such as Job ID, communication graph) issued by the scheduler in the GPU server memory through the Peripheral Component Interconnect Express (PCIe) bus monitoring, or by subscribing based on the control plane and pulling job topology, bandwidth requirements and other information from the computing power scheduler (such as Kubernetes) through the RESTful interface or message queue (such as Kafka). After the computing power job perception module pushes the computing power job metadata to the computing power master control unit, the computing power master control unit, as the control plane hub, can forward the computing power job metadata to the computing power bandwidth scheduling unit through the internal interface (such as gRPC), and further perform dynamic bandwidth scheduling through subsequent steps.

[0047] Among them, the computing power job metadata may include the computing power job identifier (Job ID), bandwidth requirement information, priority information, and network topology (such as communication graph and computing power operation task topology, etc.). The Job ID, as the unique identifier of the computing power job data packet, is convenient for associating subsequent scheduling operations (such as virtual channel allocation, bandwidth allocation); the bandwidth requirement information is used to calculate the bandwidth resource quota of the computing power of the virtual channel; the priority information is used to guide the generation of subsequent Quality of Service (QoS) policies; the network topology information is used to determine the data transmission path of the computing power job data packet corresponding to different Job IDs.

[0048] Step S2: The computing power bandwidth scheduling unit determines the first scheduling data of the computing power job identifier according to the bandwidth requirement information and the priority information. The first scheduling data includes the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel and the bandwidth information of the virtual channel. Different virtual channels correspond to different priorities. In this step, according to the job priority and bandwidth requirements, the computing power bandwidth scheduling unit can allocate exclusive virtual network interface cards (VNICs) for the computing power job data packets corresponding to each computing power job identifier from the first virtual channel of the network interface card (Network Interface Card, NIC). The first virtual channel can be multiple VNICs virtualized from a physical NIC using the SR-IOV technology, and each VNIC corresponds to an independent PCIe Function. Different VNICs can correspond to different priorities. For example, if the priority of the computing power job data packet corresponding to Job ID1 is higher than that of the computing power job data packet corresponding to Job ID2, then the virtual channel corresponding to the computing power job data packet indicated by Job ID1 in the first virtual channel can be VNIC1, and the virtual channel corresponding to the computing power job data packet indicated by Job ID2 in the first virtual channel can be VNIC2, and the priority of VNIC1 is higher than that of VNIC2. Moreover, the computing power bandwidth scheduling unit determines the bandwidth information of each virtual channel according to the bandwidth requirements of the computing power job data packets indicated by different computing power job identifiers. For example, if the bandwidth requirement of the computing power job data packet indicated by Job ID1 is 40 Gbps, the bandwidth requirement of the computing power job data packet indicated by Job ID2 is 20 Gbps, and the virtual channel corresponding to Job ID1 can be VNIC1, and the virtual channel corresponding to Job ID2 can be VNIC2, then when determining the bandwidth information of VNIC1 and VNIC2, the ratio between the bandwidth information corresponding to VNIC1 and the bandwidth information corresponding to VNIC2 can be determined according to the bandwidth requirement corresponding to Job ID1 (40 Gbps) and the bandwidth requirement corresponding to Job ID2 (20 Gbps) as 2:1. Dynamically adjusting the bandwidth allocation ratio according to the bandwidth requirements of the job data indicated by each computing power job identifier realizes on-demand scheduling of job-level bandwidth and improves the utilization rate of cluster resources.

[0049] Step S3, the computing power bandwidth scheduling unit sends the first scheduling data to the GPU server indicated by the network topology information, and the GPU server is used to perform bandwidth scheduling according to the first scheduling data; In this step, the computing power bandwidth scheduling unit sends the first scheduling data to the GPU server indicated by the network topology information, so that the GPU server can perform bandwidth scheduling according to the first scheduling data and allocate exclusive receive / transmit queues for each VNIC. Performing bandwidth scheduling according to the first scheduling data in the GPU server realizes the division of computing power bandwidth resources in units of jobs, rather than the traditional division in units of virtual machines or containers, and improves the accuracy of resource allocation.

[0050] Step S4: The link resource mapping engine sends the second scheduling data to the switch indicated by the network topology information. The switch is used to perform bandwidth scheduling according to the second scheduling data, and the second scheduling data is the mapping data of the first scheduling data generated by the link resource mapping engine based on the first scheduling data.

[0051] In this step, the link resource mapping engine can cooperate with the ToR / Leaf switch, obtain topology information through LLDP, dynamically adjust the port bandwidth mapping, realize fine-grained bandwidth scheduling at the job level and end-to-end communication optimization, reduce the situation of resource preemption, and avoid problems such as coarse-grained scheduling.

[0052] In this embodiment, the computing power job perception module collects computing power job metadata and transfers it to the computing power bandwidth scheduling unit to achieve accurate perception of job priority, bandwidth requirements, and network topology; the computing power main control management unit sends the computing power job metadata to the computing power bandwidth scheduling unit, and the computing power bandwidth scheduling unit allocates different-priority virtual channels and configures corresponding bandwidth information for the computing power job data packets indicated by different computing power job identifiers, solving the problem of rough traditional scheduling granularity and realizing fine-grained resource isolation and priority guarantee; then, the computing power bandwidth scheduling unit sends the first scheduling data to the GPU server for execution, enabling the GPU server side to mark traffic and bind queues according to the policy, realizing on-demand scheduling of bandwidth; and the link resource mapping engine maps the virtual channels to the bandwidth allocation rules of the switch physical ports, combines the topology information to achieve end-to-end collaborative scheduling, and dynamically adjusts the physical link resources, which can effectively solve problems such as multi-tenant bandwidth preemption conflicts and mismatches between communication topologies and computing power operation task behaviors, improve the utilization rate of computing power bandwidth resources and the accuracy of on-demand scheduling, realize the dynamic adaptation of computing power operation tasks and computing power bandwidth resources, and effectively realize a high-bandwidth network for computing power.

[0053] In one embodiment, after step S3 and before step S4, the method further includes: Step S5: The link resource mapping engine obtains link layer discovery protocol information from the switch, and the link layer discovery protocol information is used for link resource mapping. Step S6: The link resource mapping engine generates the second scheduling data according to the link layer discovery protocol information and the first scheduling data sent by the computing power bandwidth scheduling unit.

[0054] In this embodiment, the link resource mapping engine obtains Link Layer Discovery Protocol (LLDP) link information from a Top-of-Rack (ToR) switch, such as real-time information like the physical connection topology between the switch and the server, port capabilities, etc., providing a basis for subsequent mapping of virtual resources to physical links. Exemplarily, SNMP or some ToR switch management interfaces can be used to obtain the LLDP link information of the switch periodically or triggeringly; then, the link resource mapping engine combines the LLDP link information and the first scheduling data to accurately map the virtual channels to the physical ports, generating switch flow table configuration instructions to achieve the collaborative optimization of the logical scheduling policy and the physical network resources. In this way, the bandwidth scheduling policy can fully consider the physical link characteristics (such as bandwidth upper limit, link quality), improving the rationality of the end-to-end path planning; by dynamically discovering network changes through LLDP, the mapping policy can be adjusted in real time, enhancing the adaptability of the system to network topology changes; at the same time, the binding of the physical link state and the virtual channel ensures that high-priority job traffic can be directed to the optimal path, reducing the congestion risk and improving the communication quality and bandwidth guarantee ability of critical services.

[0055] In one embodiment, step S2 includes: Step S21, the computing power bandwidth scheduling unit sets a Differentiated Services Code Point for the job data indicated by the computing power job identifier according to the priority information, and the Differentiated Services Code Point is used to determine the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel; Step S22, when the Differentiated Services Code Point indicates that the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel is the first channel, the computing power bandwidth scheduling unit determines the first bandwidth information of the first channel according to the bandwidth requirement information, and the first scheduling data includes the Differentiated Services Code Point and the first bandwidth information.

[0056] In this embodiment, by setting different Differentiated Services Code Points (DSCPs) for the computing power job data packets indicated by different computing power job identifiers, the job priorities are directly mapped to network layer QoS markings, enabling the GPU server and the switch to perform differential processing on different priority traffic based on the DSCP values (such as preferentially forwarding traffic with high DSCP values), thus achieving fine-grained priority scheduling. After determining the DSCP value corresponding to the virtual channel, exclusive bandwidth information (such as the first bandwidth information of the first channel) is configured for this channel according to the bandwidth requirement, binding the logical priority to the bandwidth resources of the computing power. This not only ensures low-latency transmission of high-priority jobs but also avoids resource preemption conflicts in a multi-tenant scenario through bandwidth quotas. The GPU server can accurately identify the virtual channel to which the traffic belongs based on the DSCP value and perform bandwidth scheduling. Combining with the priority processing of the DSCP value by the switch, an end-to-end priority awareness and bandwidth guarantee link is formed, improving the stability of network usage by important tenants of the intelligent computing center cloud platform and the stability of access to some key computing power resources.

[0057] In one embodiment, step S2 further includes: Step S23: When the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the differentiated services code point is the second channel in the first virtual channel, the computing power bandwidth scheduling unit generates a scheduling instruction for the second channel, where the scheduling instruction is used to indicate that the computing power job data packet is allocated to the second channel. The second bandwidth information of the second channel is a fixed bandwidth value. The first scheduling data includes the differentiated services code point and the scheduling instruction, and the priority of the second channel is higher than that of the first channel.

[0058] In this embodiment, a specific virtual channel can be set for the computing power job data packet indicated by the computing power job identifier with the highest priority. That is, the computing power bandwidth scheduling unit sets the differentiated services code point DSCP (such as DSCP = 46) value for the job data indicated by the computing power job identifier (such as Job ID0) according to the priority information, indicating that the job data indicated by Job ID0 has the highest priority. Therefore, it is indicated by DSCP = 46 that the virtual channel corresponding to the computing power job data packet indicated by Job ID0 in the first virtual channel is the second channel, and the second channel can be the virtual channel with the highest priority. In this way, by allocating the second channel and setting a fixed bandwidth value for the high-priority computing power job data packet, a differentiated bandwidth guarantee mechanism is constructed.

[0059] Specifically, when the Differentiated Services Code Point, i.e., the DSCP value, indicates a high-priority job, a dedicated scheduling instruction is generated to direct its traffic to an exclusive second channel, ensuring that critical services obtain stable computing power and bandwidth resources and avoiding bandwidth preemption caused by sudden bursts of low-priority traffic. At the same time, by setting a fixed bandwidth value (such as 40 Gbps) instead of dynamic proportional allocation, deterministic network services are provided for high-priority jobs to meet the strict bandwidth requirements of latency-sensitive operations such as gradient synchronization in AI training. The design that the priority of the second channel is higher than that of the first channel enables the switch to preferentially forward high-priority traffic during congestion, further reducing the packet loss rate and latency of critical services. Hierarchical scheduling of multi-priority jobs is achieved, which not only ensures the stable operation of high-priority computing power running tasks but also allows low-priority computing power running tasks to flexibly use the remaining bandwidth, improving the overall system efficiency and resource utilization rate in the heterogeneous workload hybrid deployment scenario.

[0060] In one embodiment, step S22 includes: Step S221: The computing power bandwidth scheduling unit determines the first bandwidth information of the first channel according to the ratio of the total available bandwidth and the bandwidth value corresponding to the bandwidth demand information in the total bandwidth value corresponding to the total bandwidth demand information.

[0061] In this embodiment, taking Job A with a bandwidth demand of 40 Gbps, Job B with a bandwidth demand of 30 Gbps, and Job C with a bandwidth demand of 30 Gbps as an example, and the total bandwidth can be 100 Gbps, where 20 Gbps is the set fixed bandwidth value, so the total available bandwidth is 80 Gbps. First, the computing power bandwidth scheduling unit calculates the ratio of the total available bandwidth and the bandwidth value corresponding to the bandwidth demand information of each Job ID in the total bandwidth value corresponding to the total bandwidth demand information: The proportion of Job A is 40 Gbps / (40 + 30 + 30) Gbps = 40%; the proportion of Job B is 30 Gbps / 100 Gbps = 30%; the proportion of Job C is 30 Gbps / 100 Gbps = 30%. Based on this, it can be determined that the first bandwidth information of the first channel corresponding to Job A is 80 Gbps × 40% = 32 Gbps; the first bandwidth information of the first channel corresponding to Job B is 80 Gbps × 30% = 24 Gbps; the first bandwidth information of the first channel corresponding to Job C is 80 Gbps × 30% = 24 Gbps.

[0062] Among them, if Job C releases 10 Gbps of demand midway (i.e., the new demand is 20 Gbps) and the total demand becomes 90 Gbps, the ratio can be recalculated: The proportion of Job A is 40 / 90 ≈ 44.4%. At this time, the first bandwidth information of the first channel corresponding to Job A is 80 Gbps × 44.4% ≈ 35.5 Gbps; The proportion of Job B is 30 / 90 ≈ 33.3%. At this time, the first bandwidth information of the first channel corresponding to Job B is 80 Gbps × 33.3% ≈ 26.6 Gbps; The proportion of Job C is 20 / 90 ≈ 22.2%. At this time, the first bandwidth information of the first channel corresponding to Job C is 80 Gbps × 22.2% ≈ 17.7 Gbps.

[0063] In this way, when the total bandwidth is insufficient, the bandwidth of each job can be compressed proportionally, avoiding some jobs from over-preempting and causing starvation of other jobs, and improving the overall resource utilization rate; making the bandwidth allocation be in a direct proportional relationship with the actual job requirements, ensuring that jobs with high bandwidth requirements obtain more resources, and jobs with low requirements are allocated according to needs, reducing resource waste; and, through the proportional allocation mechanism, providing relatively fair bandwidth guarantees for jobs with different priorities, while ensuring the basic requirements of high-priority jobs, allowing low-priority jobs to reasonably use resources; when the total bandwidth or job requirements change, automatically recalculate the ratio and adjust the allocation without manual intervention, improving the system flexibility. It is applicable to the multi-tenant shared cluster scenario and can dynamically balance the bandwidth allocation according to the real-time job requirements. It realizes the dynamic adaptation of the computing power operation tasks and the bandwidth resources of the computing power, and effectively realizes the high-bandwidth network of the computing power.

[0064] In one embodiment, after the step S3, the method further includes: Step S7, when the GPU server receives the first scheduling data, it calls the multi-channel transceiver module to generate the first virtual channel, and each virtual channel in the first virtual channel has an independent cache and queue management mechanism; Step S8, the GPU server performs bandwidth scheduling based on the first virtual channel according to the first scheduling data.

[0065] In this embodiment, a first virtual channel with an independent cache and queue management mechanism is generated on the GPU server side through a multi-channel transceiver module, and the logical scheduling policy is converted into physically executable network resources; then bandwidth scheduling is performed based on this virtual channel. The independent caches of each virtual channel avoid cache contention between jobs with different priorities, and the traffic of high-priority jobs will not be blocked by low-priority traffic, improving the stability of network usage by important tenants of the intelligent computing center cloud platform and the stability of accessing some key computing power resources; and it allows configuring exclusive QoS policies (such as PFC / DCQCN) for different virtual channels to meet low-latency requirements such as gradient synchronization in AI training, and realizes the allocation of dedicated virtual channels for different job types (such as training / inference). For example, a low-latency queue is allocated for real-time inference traffic, and a high-bandwidth queue is allocated for batch training, improving the overall efficiency in a mixed-load scenario.

[0066] In addition, combined with hardware offloading technologies such as Smart NIC, the creation and management of virtual channels can be directly executed by the network card firmware, reducing CPU overhead and accelerating the data path; when congestion occurs in a certain virtual channel, the independent cache mechanism can prevent it from affecting other channels, improving the system fault tolerance. In this way, through the virtual channel visualization and refined management on the server side, the bandwidth scheduling policy is converted into physically executable traffic control behaviors, improving the end-to-end network performance and resource utilization rate.

[0067] In some embodiments, as Figure 3 shown, the method for dynamically scheduling the high-bandwidth network of the intelligent computing center cloud platform may include the following steps: The computing power job perception module subscribes to the computing power job metadata to achieve accurate perception of job priorities, bandwidth requirements, and network topologies; The computing power master control management unit analyzes the computing power job metadata. The computing power master control management unit has the ability to dock with the computing power scheduler and receive strategies such as the computing power operation task topology and bandwidth requirements, and can send the computing power job metadata to the computing power bandwidth scheduling unit; The computing power bandwidth scheduling unit adjusts the QoS policy according to the job priority, and injects different job data packets into different dscp values according to the computing power operation task priority, that is, enters the VNIC channels allocated by the multi-channel transceiver module for the GPU servers in the GPU cluster, and dynamically allocates the bandwidth of each VNIC channel according to the bandwidth required by each job, solving the problem of rough traditional scheduling granularity and realizing fine-grained resource isolation and priority guarantee; The link resource mapping engine coordinates with the TOR / Leaf switch for dynamic bandwidth allocation. By allocating bandwidth for each VSP channel to dynamically adjust the physical link resources, it can effectively solve problems such as multi-tenant bandwidth preemption conflicts and mismatches between communication topologies and computing power operation task behaviors; The multi-channel transceiver module supports concurrent communication of multiple independent virtual channels (VNIC and VSP). Each channel has an independent cache and queue management mechanism. The independent caches of each virtual channel avoid cache contention between jobs with different priorities; The RDMA enhanced engine optimizes the communication performance of large models and can support the RoCEv2 protocol.

[0068] In this way, on-demand scheduling of bandwidth is achieved, enabling end-to-end precise scheduling from logical resources to physical links, improving the network efficiency of heterogeneous computing power clusters, enhancing the network stability for important tenants of the intelligent computing center cloud platform, and the stability of accessing some key computing power resources.

[0069] Please refer to Figure 4 , Figure 4 FIG. 2 is a second structural diagram of a high-bandwidth network device for dynamically scheduling computing power in an intelligent computing center cloud platform according to an embodiment of the present invention. As Figure 4 shown, the high-bandwidth network device 400 for dynamically scheduling computing power in the intelligent computing center cloud platform includes: a computing power master control management unit 401, a computing power bandwidth scheduling unit 402, a link resource mapping engine 403, and a computing power job awareness module; The computing power master control management unit 401 is configured to send the computing power job metadata pushed by the computing power job awareness module to the computing power bandwidth scheduling unit. The computing power job metadata includes the computing power job identifier corresponding to the computing power job data packet, the bandwidth requirement information corresponding to the computing power job data packet, the priority information corresponding to the computing power job data packet, and network topology information. The computing power job awareness module is used to subscribe to the computing power job metadata of the GPU server; The computing power bandwidth scheduling unit 402 is configured to determine first scheduling data of the computing power job identifier according to the bandwidth requirement information and the priority information. The first scheduling data includes the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel and the bandwidth information of the virtual channel. Different virtual channels correspond to different priorities; The computing power bandwidth scheduling unit 402 is further configured to send the first scheduling data to the GPU server indicated by the network topology information. The GPU server is configured to perform bandwidth scheduling according to the first scheduling data; The link resource mapping engine 403 is configured to send second scheduling data to the switch indicated by the network topology information. The switch is configured to perform bandwidth scheduling according to the second scheduling data. The second scheduling data is mapping data of the first scheduling data generated by the link resource mapping engine according to the first scheduling data.

[0070] In one embodiment, the link resource mapping engine is further configured to obtain link layer discovery protocol information from the switch, where the link layer discovery protocol information is used for link resource mapping; The link resource mapping engine is further configured to generate the second scheduling data according to the link layer discovery protocol information and the first scheduling data sent by the computing power bandwidth scheduling unit.

[0071] In one embodiment, the computing power bandwidth scheduling unit 402 is specifically configured to: The computing power bandwidth scheduling unit sets a differentiated service code point for the job data indicated by the computing power job identifier according to the priority information, where the differentiated service code point is used to determine the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel; When the differentiated service code point indicates that the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel is the first channel, the computing power bandwidth scheduling unit determines the first bandwidth information of the first channel according to the bandwidth requirement information, and the first scheduling data includes the differentiated service code point and the first bandwidth information.

[0072] In one embodiment, the computing power bandwidth scheduling unit 402 is further specifically configured to: When the differentiated service code point indicates that the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel is the second channel, the computing power bandwidth scheduling unit generates a scheduling instruction for the second channel, where the scheduling instruction is used to indicate that the computing power job data packet is allocated to the second channel, the second bandwidth information of the second channel is a fixed bandwidth value, the first scheduling data includes the differentiated service code point and the scheduling instruction, and the priority of the second channel is higher than that of the first channel.

[0073] In one embodiment, the computing power bandwidth scheduling unit 402 is further specifically configured to: The computing power bandwidth scheduling unit determines the first bandwidth information of the first channel according to the ratio of the total available bandwidth to the bandwidth value corresponding to the bandwidth requirement information in the total bandwidth value corresponding to the total bandwidth requirement information.

[0074] In one embodiment, the device further includes a GPU server: When receiving the first scheduling data, the GPU server calls a multi-channel transceiver module to generate the first virtual channel, and each virtual channel in the first virtual channel has an independent cache and queue management mechanism; The GPU server performs bandwidth scheduling based on the first virtual channel according to the first scheduling data.

[0075] The device for the high - bandwidth network for dynamically scheduling computing power in the intelligent computing center cloud platform provided by the embodiments of the present invention can implement each process of the above - mentioned method for the high - bandwidth network for dynamically scheduling computing power in the intelligent computing center cloud platform. The technical features correspond one by one and can achieve the same technical effects. To avoid repetition, they will not be elaborated here.

[0076] It should be noted that the device for the high - bandwidth network for dynamically scheduling computing power in the intelligent computing center cloud platform in the embodiments of the present invention can be a device, or a component, an integrated circuit, or a chip in an electronic device.

[0077] The embodiments of the present invention also provide an electronic device. Refer to Figure 5 , Figure 5 which is a schematic structural diagram of an electronic device provided by the embodiments of the present invention. The electronic device includes a memory 501, a processor 502, and a program or instruction running on the memory 501. When the program or instruction is executed by the processor 502, it can implement Figure 1 any step in the corresponding method embodiment of the high - bandwidth network for dynamically scheduling computing power in the intelligent computing center cloud platform and achieve the same beneficial effects. They will not be elaborated here.

[0078] Among them, the processor 502 can be a CPU, an ASIC, an FPGA, or a GPU.

[0079] Those of ordinary skill in the art can understand that all or part of the steps of implementing the method embodiment of the high - bandwidth network for dynamically scheduling computing power in the intelligent computing center cloud platform can be completed by hardware related to program instructions. The program can be stored in a readable medium.

[0080] The embodiments of the present invention also provide a readable storage medium. A computer program is stored on the readable storage medium. When the computer program is executed by a processor, it can implement the above - mentioned Figure 1 any step in the corresponding method embodiment of the high - bandwidth network for dynamically scheduling computing power in the intelligent computing center cloud platform and can achieve the same technical effects. To avoid repetition, they will not be elaborated here. The storage medium can be, for example, a read - only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0081] The present invention also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement each process of the above - mentioned Figure 1 corresponding method embodiment of the high - bandwidth network for dynamically scheduling computing power in the intelligent computing center cloud platform and can achieve the same technical effects. To avoid repetition, they will not be elaborated here.

[0082] The terms "first", "second", etc. in the embodiments of the present invention are used to distinguish similar objects and do not necessarily describe a specific order or sequence. In addition, the terms "comprising", "having" and any of their variants are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. In addition, in this application, "and / or" is used to indicate at least one of the connected objects. For example, A and / or B and / or C means including A alone, B alone, C alone, as well as the cases where A and B exist together, B and C exist together, A and C exist together, and A, B and C exist together, a total of 7 cases.

[0083] It should be noted that in this text, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not clearly listed, or also includes elements inherent to such a process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising that element.

[0084] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of this application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or a second terminal device, etc.) to execute the methods of the various embodiments of this application.

[0085] The above describes the embodiments of this application in conjunction with the accompanying drawings, but this application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of this application, those of ordinary skill in the art can also make many forms without departing from the purpose of this application and the scope protected by the claims, and all of them belong to the protection scope of this application.

Claims

1. A method for a high-bandwidth network for dynamically scheduling computing power in an intelligent computing center cloud platform, characterized in that, The method includes: Step S1, the computing power master management unit sends the computing power job metadata pushed by the computing power job perception module to the computing power bandwidth scheduling unit. The computing power job metadata includes the computing power job identifier corresponding to the computing power job data packet, the bandwidth requirement information corresponding to the computing power job data packet, the priority information corresponding to the computing power job data packet, and network topology information. The computing power job perception module is used to subscribe to the computing power job metadata of the GPU server; Step S2, the computing power bandwidth scheduling unit determines the first scheduling data of the computing power job identifier according to the bandwidth requirement information and the priority information. The first scheduling data includes the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel and the bandwidth information of the virtual channel. Different virtual channels correspond to different priorities; Step S3, the computing power bandwidth scheduling unit sends the first scheduling data to the GPU server indicated by the network topology information. The GPU server is used to perform bandwidth scheduling according to the first scheduling data; Step S4, the link resource mapping engine sends the second scheduling data to the switch indicated by the network topology information. The switch is used to perform bandwidth scheduling according to the second scheduling data. The second scheduling data is the mapping data of the first scheduling data generated by the link resource mapping engine according to the first scheduling data.

2. The method according to claim 1, wherein After step S3 and before step S4, the method further includes: Step S5, the link resource mapping engine obtains link layer discovery protocol information from the switch. The link layer discovery protocol information is used for link resource mapping; Step S6, the link resource mapping engine generates the second scheduling data according to the link layer discovery protocol information and the first scheduling data sent by the computing power bandwidth scheduling unit.

3. The method according to claim 1, characterized in that Step S2 includes: Step S21, the computing power bandwidth scheduling unit sets a differentiated services code point for the job data indicated by the computing power job identifier according to the priority information. The differentiated services code point is used to determine the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel; Step S22, when the differentiated services code point indicates that the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel is the first channel, the computing power bandwidth scheduling unit determines the first bandwidth information of the first channel according to the bandwidth requirement information. The first scheduling data includes the differentiated services code point and the first bandwidth information.

4. The method according to claim 3, wherein Step S2 further includes: Step S23: When the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the differential service code point is the second channel in the first virtual channel, the computing power bandwidth scheduling unit generates a scheduling instruction for the second channel, where the scheduling instruction is used to indicate that the computing power job data packet is allocated to the second channel, the second bandwidth information of the second channel is a fixed bandwidth value, the first scheduling data includes the differential service code point and the scheduling instruction, and the priority of the second channel is higher than that of the first channel.

5. The method according to claim 3, wherein The step S22 includes: Step S221: The computing power bandwidth scheduling unit determines the first bandwidth information of the first channel according to the ratio of the total available bandwidth and the bandwidth value corresponding to the bandwidth requirement information in the total bandwidth value corresponding to the total bandwidth requirement information.

6. The method according to any one of claims 1 to 5, characterized in that After the step S3, the method further includes: Step S7: When the GPU server receives the first scheduling data, it calls the multi-channel transceiver module to generate the first virtual channel, and each virtual channel in the first virtual channel has an independent cache and queue management mechanism; Step S8: The GPU server performs bandwidth scheduling based on the first virtual channel according to the first scheduling data.

7. An apparatus for a high-bandwidth network for dynamically scheduling computing power in an intelligent computing center cloud platform, characterized in that, The device includes a computing power main control management unit, a computing power bandwidth scheduling unit, a link resource mapping engine, and a computing power job perception module; The computing power main control management unit is configured to send the computing power job metadata pushed by the computing power job perception module to the computing power bandwidth scheduling unit. The computing power job metadata includes the computing power job identifier corresponding to the computing power job data packet, the bandwidth requirement information corresponding to the computing power job data packet, the priority information corresponding to the computing power job data packet, and the network topology information. The computing power job perception module is used to subscribe to the computing power job metadata of the GPU server; The computing power bandwidth scheduling unit is configured to determine the first scheduling data of the computing power job identifier according to the bandwidth requirement information and the priority information. The first scheduling data includes the virtual channel corresponding to the computing power job data packet indicated by the computing power job identifier in the first virtual channel and the bandwidth information of the virtual channel, and different virtual channels have different priorities; The computing power bandwidth scheduling unit is further configured to send the first scheduling data to the GPU server indicated by the network topology information, and the GPU server is configured to perform bandwidth scheduling according to the first scheduling data; The link resource mapping engine is configured to send the second scheduling data to the switch indicated by the network topology information, and the switch is configured to perform bandwidth scheduling according to the second scheduling data. The second scheduling data is the mapping data of the first scheduling data generated by the link resource mapping engine according to the first scheduling data.

8. An electronic device, characterized in that, It includes: A processor, a memory, and a program stored on the memory and executable on the processor. When the program is executed by the processor, it implements the steps of the method for the high-bandwidth network for dynamically scheduling computing power of the intelligent computing center cloud platform as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the method for a high-bandwidth network for dynamically scheduling computing power of the intelligent computing center cloud platform according to any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that, It includes computer instructions, and when the computer instructions are executed by a processor, the steps of the method for a high-bandwidth network for dynamically scheduling computing power of the intelligent computing center cloud platform according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Data transmission method and data transmission device

    CN103138888A

  • Computing power scheduling system, method and device and storage medium

    CN114756340A

  • Service transmission method and device and storage medium

    CN115734291A

  • Multi-path forwarding method and system in computing power network

    CN116319522A

  • Calculation power scheduling method and system

    CN119127473A