Switching communication device and method, server

By designing a switching communication device, we have achieved proactive management of GPU resources and dynamic optimization of network topology, solving the problems of inflexible resource scheduling and low communication efficiency in traditional data center architectures, and improving the efficiency and topology adaptability of GPU communication.

CN121433923BActive Publication Date: 2026-03-27INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-31
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In traditional data center architectures, communication between graphics processors and servers relies on central processing unit scheduling, which makes it difficult to flexibly schedule and pool resources, resulting in large latency and low efficiency in cross-node communication. Furthermore, existing accelerator processors have limited functionality and cannot achieve intelligent adaptation and dynamic optimization of network topology.

Method used

Design a switching communication device that connects to the CPU in endpoint mode and to the GPU in root complex mode. It integrates a switching logic module, multiple MAC ports, a data processing module, and an on-chip management system to achieve active management of GPU resources and dynamic optimization of network topology, and supports data interaction in direct connection and switching modes.

Benefits of technology

It enables pooling and elastic scheduling of GPU resources, improving communication efficiency and network topology adaptability, supporting high-performance direct-connect networks and large-scale switching networks, breaking through the limitations of fixed topologies, and improving resource management capabilities and the flexibility of switching strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121433923B_ABST
    Figure CN121433923B_ABST
Patent Text Reader

Abstract

The application discloses a kind of switching communication device and method, server, it is related to computer technical field, by being configured as the second interface of root complex mode actively manages GPU, equipment level decoupling and resource poolization foundation are realized.The on-chip management system and switching logic module constitute control core, dynamically execute data exchange strategy, realize from passive forwarding to the change of autonomous management.Virtual port module carries out software definition and dynamic allocation to network resource, improve the flexibility of resource utilization.Meanwhile, support the MAC port of two modes of direct connection and switching, so that device can flexibly build high-performance direct connection network or large-scale switching network, break the limit of fixed topology.Data processing module then ensures the efficiency of data conversion between different interface protocols.The device improves the communication efficiency and schedulability between GPUs by integrating high-speed switching, intelligent management, resource virtualization and flexible networking capabilities.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, and particularly relates to a switching communication device and method and a server. BACKGROUND

[0002] In a traditional data center architecture, a graphics processing unit (GPU) is tightly coupled with a server as a peripheral component interconnect express (PCIe) device, and communication thereof completely depends on scheduling and forwarding of a central processing unit (CPU) and a motherboard chipset of the server, which leads to difficulty in flexible scheduling and pooling of GPU resources, large communication delay and low efficiency across nodes.

[0003] In the related art, although a scheme uses an acceleration processor as a switching card between the GPU and the server to assist communication, the acceleration processor in the related art has a single function and can only passively forward data according to external configuration information and instructions. The acceleration processor lacks active management capability of the GPU, has a fixed network topology and cannot be intelligently adapted, cannot realize dynamic optimization of resources and large-scale flexible networking, and still has obvious limitations in dealing with modern heterogeneous computing scenarios. SUMMARY

[0004] The present application provides a switching communication device and method and a server to at least solve the problems of weak resource management capability, poor network topology adaptability and inflexible switching strategy of a data switching device in the related art.

[0005] The present application provides a switching communication device, wherein the device is in communication connection with at least one GPU and at least one CPU, and comprises:

[0006] a first interface configured in an endpoint mode and used for connecting with the CPU;

[0007] at least one second interface configured in a root complex mode and used for connecting with the GPU;

[0008] a switching logic module connected to the first interface and the at least one second interface and used for realizing data interaction between the GPU and the CPU and data interaction between the GPUs connected by the at least one second interface respectively;

[0009] a plurality of MAC ports used for directly connecting to at least one other switching communication device in a direct connection mode or connecting to an Ethernet switching network in a switching mode to realize data interaction between other GPUs and the GPU under connection of the other switching communication device;

[0010] a data processing module, configured to perform data conversion and processing on data between the first interface, the second interface and the plurality of MAC ports;

[0011] a virtual port module, configured to virtualize and allocate the plurality of MAC port resources to the GPU;

[0012] an on-chip management system, configured to manage and configure configuration information for data interaction.

[0013] The application further provides a switching communication method, wherein the method is applied to the on-chip management system of the switching communication device, and comprises the following steps of:

[0014] performing resource discovery and virtualization management on the GPU connected to each of the at least one second interface, determining a target GPU available, and reporting GPU resource information of the target GPU to the CPU;

[0015] performing network topology detection through the plurality of MAC ports, determining a target network topology available, and the target network topology comprises a direct connection mode and / or a switching mode;

[0016] generating, according to the received task request, the GPU resource information and the target network topology, a routing strategy applied to the switching logic module for data stream transmission, an allocation strategy applied to the virtual port module for virtual resource allocation, and a processing rule applied to the data processing module for data processing;

[0017] in response to a data communication request of the target GPU, determining a data interaction path of the data stream between the target GPU, the CPU and the plurality of MAC ports according to the routing strategy, the allocation strategy and the processing rule.

[0018] The application further provides an electronic device, comprising a memory configured to store a computer program, and a processor configured to execute the computer program to implement the steps of any of the switching communication methods.

[0019] The application further provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of any of the switching communication methods.

[0020] The application further provides a computer program product, comprising a computer program, and the computer program is executed by a processor to implement the steps of any of the switching communication methods.

[0021] The application provides a switching communication device and method and a server, wherein the second interface configured as a root complex mode actively manages a GPU, device-level decoupling and resource pooling are realized, an on-chip management system and a switching logic module constitute a control core, a dynamic data switching strategy is executed, and a change from passive forwarding to autonomous management is realized, a virtual port module performs software definition and dynamic allocation on network resources, and flexibility of resource utilization is improved. Meanwhile, the MAC port supports direct connection and switching modes, so that the device can flexibly construct a high-performance direct connection network or a large-scale switching network, and the limitation of a fixed topology is broken. The data processing module ensures the efficiency of data conversion between different interface protocols. The device integrates high-speed switching, intelligent management, resource virtualization and flexible networking capabilities, constructs an active and programmable interconnection platform, improves the communication efficiency and schedulability between GPUs, and improves the resource management capability, network topology adaptability and switching strategy flexibility in the communication process between GPUs. BRIEF DESCRIPTION OF DRAWINGS

[0022] In order to more clearly illustrate the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0023] Figure 1 A structural schematic diagram of a switching communication device provided by an embodiment of the present application is shown in the figure.

[0024] Figure 2 A structural schematic diagram of a switching logic module provided by an embodiment of the present application is shown in the figure.

[0025] Figure 3 A direct connection mode schematic diagram of a switching communication device provided by an embodiment of the present application is shown in the figure.

[0026] Figure 4 A switching mode schematic diagram of a switching communication device provided by an embodiment of the present application is shown in the figure.

[0027] Figure 5 A flowchart of a switching communication method provided by an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION

[0028] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.

[0029] It should be noted that in the description of the present application, the term "comprising", "containing" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or further includes elements inherent to such a process, method, article or device. The terms "first", "second" and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.

[0030] The present application relates to the technical field of computing, in particular to a switching communication device for connecting GPU and CPU. The device aims to build an efficient, flexible and intelligent interconnection intermediate layer to overcome the various limitations brought by the tight coupling of computing acceleration devices and host systems in traditional architecture.

[0031] The core idea of the switching communication device is that it is an independent, functionally integrated hardware entity, physically located between the GPU and the server motherboard (such as CPU). Here, not only refers to the physical location intervention, but more importantly, on the logic and data path, the device becomes a necessary hub and management node for GPU and CPU communication. Through this intervening design, the device can uniformly abstract and manage the GPU resources mounted on it, thereby providing a hardware foundation for the pooling, elastic scheduling and flexible networking of GPU resources across nodes.

[0032] In order to enable those skilled in the art to better understand the present application, the present application will be further described in detail below in conjunction with the drawings and specific embodiments.

[0033] Figure 1 A structural schematic diagram of a switching communication device provided for the embodiments of the present application is shown, which is described in detail in combination with the structure of the switching communication device.

[0034] As shown in Figure 1 The switching communication device, which is in communication connection with at least one GPU and at least one CPU, comprises:

[0035] A first interface configured in endpoint mode for connecting with the CPU;

[0036] At least one second interface configured in root complex mode for connecting with the GPU respectively;

[0037] A switching logic module connected to the first interface and the at least one second interface, for realizing data interaction between the GPU and the CPU, and data interaction between the GPUs connected by the at least one second interface respectively;

[0038] a plurality of MAC ports for connecting to at least one other switching communication device directly through a direct connection mode or to an Ethernet switching network through a switching mode to realize data interaction between other GPUs and GPUs connected through the other switching communication device;

[0039] a data processing module for data conversion and processing of data between the first interface, the second interface and the plurality of MAC ports;

[0040] a virtual port module for virtualizing and allocating the plurality of MAC port resources to the GPUs;

[0041] an on-chip management system for managing and configuring configuration information for data interaction.

[0042] In embodiments of the present application, the first interface is physically and logically set to an endpoint mode, indicating that the responder responds to requests from an upper bus manager. The first interface is specifically used to establish a communication connection with a CPU. This means that, from the perspective of a server operating system, the switching communication device of the present application is recognized and driven as a standard PCIe endpoint device (for example: a functional complex network card or acceleration card). Access and control of all CPUs in the server to the GPUs connected behind the switching communication device are performed in the form of standard PCIe transactions through the first interface, laying a foundation for software-level compatibility of the switching communication device with existing server systems, for example: the first interface is a PCIe endpoint (PCIe EP) interface, and the first interface is used as an Endpoint mode to connect the interactive network device and the server motherboard using the PCIe Gen5x16 standard.

[0043] Corresponding to the CPU-oriented first interface, at least one second interface is also provided, and the second interface is configured as a root complex mode. The root complex mode second interface is responsible for initiating bus enumeration, configuration and management of all devices downstream thereof. By setting the second interface as a root complex mode, the switching communication device obtains complete management rights over the GPUs connected downstream thereof. Each second interface can connect one GPU respectively and enumerate, initialize and allocate resources to it as a host. This makes the GPU no longer directly connected to the server motherboard in a physical sense, but becomes a managed sub-device of the switching communication device in a logical sense, thereby realizing decoupling of the GPU and the native server motherboard at the hardware connection level, for example: the second interface is a PCIe root complex (PCIe RC) interface, and the second interface is used as a Root Complex mode to connect the GPU card using the PCIe Gen5 standard.

[0044] To realize efficient data flow, a switching logic module is integrated inside the switching communication device. The switching logic module is connected to the first interface and all the second interfaces in hardware logic, and constitutes the data switching core inside the switching communication device. Its function is to be responsible for path selection and data forwarding, specifically, to realize two types of data interaction: one is to pass through the switching communication device for uplink and downlink data interaction between the GPU and the CPU; the other is to realize horizontal data interaction between GPUs connected by different second interfaces inside the switching communication device. For example, when a GPU needs to transmit data to the CPU, the switching logic module will guide the data stream from the second interface to which the GPU belongs to the first interface; when two GPUs managed by the switching communication device need to directly exchange data, the switching logic module can complete the data routing and switching between the two inside the switching communication device without the need to transmit data to the CPU, which greatly reduces the communication delay and server bandwidth pressure.

[0045] To expand the interconnection range, the device is also equipped with multiple MAC ports, i.e. Media Access Control (MAC) ports. The MAC port is a physical channel for the device to communicate with the external network environment. It supports at least two working modes: direct connection mode and switching mode. In the direct connection mode, the MAC ports can be used to directly connect to at least one other switching communication device of the same type to establish a point-to-point or chain-type dedicated high-speed link through a cable or an optical cable. In the switching mode, the MAC ports are used to connect to a standard Ethernet switch in the data center, i.e. an Ethernet switch commonly used in the data center. The two modes together enable the device to flexibly build interconnection networks of different scales and topologies: through the direct connection mode, the device can realize ultra-low delay and high-bandwidth direct communication of GPUs within a rack or between adjacent nodes; through the switching mode, the device can be integrated into a larger cluster network to realize GPU resource interconnection across racks and nodes. No matter which mode is adopted, the fundamental purpose is to realize data interaction between the GPU managed by the device and other GPUs managed by other devices of the same type.

[0046] The data processing module in the device is responsible for protocol conversion and data processing. Since the data can come from the first interface (following the PCIe protocol), the second interface (also following the PCIe protocol), or multiple MAC ports (following network protocols such as Ethernet and RDMA over Converged Ethernet (RoCE)), the data processing module needs to perform necessary data conversion and processing on these different sources of data following different protocols. This includes but is not limited to: parsing and packaging PCIe read-write transactions, packetizing and unpackaging network packets, converting between different data formats, and possibly data checking, encryption and decryption operations. This module is a functional unit that ensures seamless, accurate and efficient flow of cross-protocol data.

[0047] To manage network resources more finely and flexibly, the device introduces a virtual port module. Virtualization is a logical technology that abstracts, divides and combines physical resources. The specific function of the virtual port module is to virtualize the bandwidth and queue resources provided by multiple MAC ports in the physical world, forming multiple logically independent and flexibly configurable virtual communication channels. Then, the virtual port module dynamically allocates these virtualized channel resources to each GPU managed by the device. This means that a physical MAC port can be shared by multiple GPUs, or a GPU can exclusively use the aggregated bandwidth of multiple physical ports. This breaks the fixed binding relationship between physical ports and GPUs, realizes software-defined and on-demand allocation of network resources, and greatly improves the flexibility and efficiency of resource utilization.

[0048] The on-chip management system coordinates and controls the work of all the above modules. It is a control core integrated in the hardware of the device, and its function is to manage and configure configuration information for data interaction. Configuration information is a broad concept that covers various strategies, parameters and state tables required for the normal operation of the device. For example: configuring address routing tables and port forwarding rules for the switching logic module; configuring protocol conversion parameters for the data processing module; configuring mapping strategies from physical ports to virtual channels for the virtual port module; configuring network addresses, rates and working modes (direct connection or switching) for multiple MAC ports. The on-chip management system can respond to control instructions from the CPU, or adjust these configuration information autonomously or dynamically according to the network topology and load conditions it perceives, so that the entire device can intelligently and adaptively optimize data flow paths and resource allocation, realizing the transition from passive connection to active management, i.e. performing resource management of the processor, network topology discovery and data interaction strategy generation.

[0049] The application integrates an endpoint mode interface, a root complex mode interface, internal switching logic, multi-mode MAC ports, protocol data processing, port resource virtualization, and on-chip intelligent management, and constructs a heterogeneous computing interconnection switching platform with complete functions and intelligent management. The device not only effectively decouples the GPU and the server hardware, creating conditions for GPU resource pooling, but also supports the construction of various efficient interconnection topologies from direct connection to switching and from local to large-scale through its flexible port mode and virtualization capability. At the same time, the on-chip management system gives the device dynamic configuration and optimization capability, significantly improving the flexibility, scalability, and data processing efficiency of the overall interconnection system.

[0050] In an implementation manner of the embodiment of the application, the virtual port module is dynamically configured by the on-chip management system based on the number of MAC ports and the number of GPUs, and is specifically used for:

[0051] performing physical port resource virtualization on the plurality of MAC ports, and dynamically allocating the virtualized physical port resources to the GPUs.

[0052] In the embodiment of the application, the virtual port module is not operated in a fixed or predefined manner, but is managed and driven in real time and adaptively by the on-chip management system, and the core is dynamic configuration. The dynamic configuration here means that the operating parameters and resource mapping strategies of the virtual port module are not fixed, but can be adjusted and optimized in real time according to the actual hardware environment and load status of the device. Specifically, the key basis for the configuration decision of the on-chip management system is two types of real-time information: one is the number of MAC ports provided by the device itself, that is, the total of the currently available physical network connection resources; the other is the number of GPUs connected and managed through the second interface, that is, the total amount of data processing units currently needing to be served. By continuously monitoring and evaluating these two key quantities, the on-chip management system can accurately perceive the corresponding relationship between physical resource supply and logical service demand, and thus calculate the most reasonable resource configuration strategy for the virtual port module, that is, the virtual port module can be flexibly configured according to the number of GPUs and the number of MAC ports.

[0053] Based on the above dynamic evaluation, the specific function execution of the virtual port module is divided into two logical stages. The first stage is physical port resource virtualization. The physical port resource here refers to the communication capability possessed by each MAC port, including its bandwidth, transceiver queue, cache space, and related processing unit. Virtualization is a resource abstraction technology, and its purpose is to break the rigid boundaries of physical resources. The virtual port module aggregates all the communication resources contributed by these physical ports into a unified and flexible resource pool through software-defined manner. This resource pool is no longer strictly bound to a specific physical port, but is abstracted and divided into multiple logically independent and feature-configurable virtual channels. Each virtual channel represents a certain share of communication capability and can independently carry data streams.

[0054] The second stage is to dynamically allocate these virtualized physical port resources to the GPU. Dynamic allocation means that the allocation behavior is not pre-set statically, but according to real-time strategies, the created virtual channels are logically bound to specific GPUs. For example: when it is monitored that a certain GPU is undertaking a high-intensity cross-node synchronization task, the on-chip management system can temporarily aggregate multiple virtual channels (which may correspond to the bandwidth of multiple physical ports) and allocate them to the GPU through the virtual port module, providing a super-high-speed logical link for it. Conversely, for GPUs in an idle or light load state, the most basic bandwidth channel can be allocated to them to save resources. This allocation can be dynamically established and released according to the life cycle of the task, realizing on-demand supply of network resources.

[0055] The application realizes fine and intelligent management of network connection resources through a dynamic configuration mechanism based on real-time resource and demand quantity perception. It changes the fixed connection mode of MAC ports and computing devices in the traditional architecture, and the rigid mode of resource utilization. It enables limited physical MAC port resources to serve a variable number of GPU clusters with different tasks with high flexibility, improving the efficiency and flexibility of resource utilization. At the same time, it also provides convenience for upper-layer scheduling, so that the scheduling of GPU resources is no longer limited by the topology constraints of the underlying physical network, realizing decoupling and collaborative optimization of computing resources and network resources.

[0056] In an implementable manner of an embodiment of the application, the apparatus further comprises:

[0057] The bus control module is configured to aggregate and schedule the requests of the switching logic module and the data processing module to access the local memory based on the on-chip bus.

[0058] In the embodiments of the present application, the bus control module is a coordinator of data paths and storage accesses inside the device, which is introduced to efficiently and orderly manage the concurrent accesses of the shared storage resources by the functional modules inside the device, and to ensure the overall performance and stability of the system.

[0059] The bus control module collects and schedules the requests of the switching logic module and the data processing module to access the local memory. Here, the local memory refers to the physical storage unit integrated inside the switching communication device, such as a Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM) chip or an on-chip cache. It provides the necessary high-speed data storage space for the switching logic module (used to cache the data packets to be forwarded or store the routing table) and the data processing module (used to temporarily store the intermediate data being processed or bear the direct memory access buffer).

[0060] When the switching logic module and the data processing module simultaneously or frequently need to read and write the shared local memory, access competition will occur. The role of the bus control module is to act as an arbitrator and commander, receive (i.e., collect) all access requests from the two modules, then sort and coordinate (i.e., schedule management) these requests according to the preset priority, fairness policy or real-time load situation, and finally execute these memory read / write operations in an orderly manner, preventing data conflicts and maximizing the utilization of memory bandwidth.

[0061] The bus control module works based on an on-chip bus, such as the Advanced eXtensible Interface 4 (AXI4) bus, which is a high-performance, high-frequency on-chip interconnection protocol standard. It is suitable for connecting master devices (such as processors, Direct Memory Access (DMA) controllers, hardware accelerators) and slave devices (such as memory controllers, peripherals) that require high bandwidth data processing. Selecting the AXI4 bus means that the bus control module follows the standard of the protocol to design and implement its internal interfaces, command / data channels, and handshake mechanisms. This allows the switching logic module and the data processing module to act as master devices, initiating requests to the bus control module through standard AXI4 read / write transactions; while the bus control module acts as the interconnection center, also communicating with the memory controller of the local memory using the standard AXI4 protocol. Using this standardized interface greatly simplifies the integration design between modules, improves interoperability and design reusability, and at the same time can fully exploit the performance advantages of the AXI4 protocol in supporting out-of-order completion, burst transmission, etc.

[0062] The application introduces a bus control module based on an on-chip bus, and a bus access control module based on the on-chip bus is responsible for processing the aggregated local memory access management, creating an efficient and standardized internal data path, which not only solves the problem of competition for shared memory resources by multiple high-performance modules, but also avoids access congestion through professional scheduling management, ensuring that the data forwarding of the exchange logic module and the protocol conversion of the data processing module can be parallel without contradiction smoothly and with low delay. From the internal interconnection level, the overall data processing capability of the device is fully utilized, and the internal coordination and operation efficiency of the device are improved.

[0063] In an implementation manner of the embodiment of the application, the exchange logic module comprises a configurable exchange switch matrix.

[0064] The exchange switch matrix arbitrates and directionally transmits the received data stream based on a routing table or an address mapping table configured by the on-chip management system, to realize data interaction.

[0065] In the embodiment of the application, the exchange logic module is a specific execution unit for realizing the core exchange function thereof. It internally comprises a configurable exchange switch matrix. The exchange switch matrix is the physical basis for realizing high-speed and parallel data forwarding. The exchange switch matrix is essentially a high-speed interconnection network composed of a large number of programmable electronic switch nodes. The connection state of these switch nodes is not fixed but configurable, which means that the internal data path connection relationship can be dynamically set, modified or reconfigured through software or firmware instructions. This configurability gives the exchange switch matrix high flexibility, enabling it to adapt to different data flow modes and topology requirements, rather than being limited to a fixed exchange mode.

[0066] The operation decision of the configurable switch matrix is not generated independently, but is strictly based on the routing table or address mapping table configured by the on-chip management system. The routing table is a data structure that contains rule entries about the forwarding path of data packets. Each rule usually indicates that the data packet matching a certain feature (such as a target address or a virtual channel identifier) should be output from which physical or logical port. The address mapping table is another key data structure that is responsible for resolving or translating an abstract logical address (for example, a target GPU identifier specified by the upper software) to specific physical location information (for example, a virtual channel corresponding to a second interface or a network port inside the device). The routing table or address mapping table together constitutes the address and rules that guide data exchange. The generation, maintenance and update of the routing table or address mapping table are the responsibility of the on-chip management system as the control core. The on-chip management system dynamically calculates and refreshes the entries in the routing table or address mapping table according to the global network topology information, load balancing strategy and real-time task demand, and then issues the complete and latest table configuration information to the switch logic module, so as to ensure that the switching decision is consistent with the global system state.

[0067] At runtime, the configurable switch matrix performs two core operations on the received data stream: arbitration and directional transmission. When data packets or data units from different input ends (such as different second interfaces, first interfaces or data processing modules) arrive almost simultaneously and may compete for the same output resource, the arbitration mechanism is activated. According to the preset priority strategy, fairness algorithm or quality of service rule, the arbitration mechanism determines which data stream has the right to be processed preferentially in the current period, thereby avoiding conflicts and ensuring low latency of critical data. After the arbitration determines the processing order, directional transmission is performed. At this time, the switch matrix analyzes the key identification information (such as the target address) in the current data stream, and refers to the routing table or address mapping table loaded in it to quickly find the corresponding output target. Then, the matrix dynamically configures its internal switch nodes to establish a temporary and exclusive physical or logical path from the input port of the data stream to the specified output port, efficiently and accurately guiding the data stream to the destination, and finally realizing various data interactions.

[0068] Specifically, regarding the structure of the switch logic module, the embodiment of the present application also provides a structural diagram of a switch logic module, as shown in Figure 2 The switch matrix inside the module can output data streams according to port routing and address routing caching and arbitration.

[0069] The application realizes flexible, intelligent and efficient switching mechanism by adopting a switching logic module composed of a configurable switching switch matrix and making it work under the accurate guidance of a routing table or an address mapping table dynamically provided by an on-chip management system. The data can be quickly and accurately forwarded between the GPU and the CPU, between the GPUs and between the cross-device GPUs, so that the path and strategy of data exchange are no longer static and predetermined, but can be dynamically optimized with the changes of system state, network topology and computing tasks.

[0070] In an implementation manner of the embodiment of the application, the on-chip management system is specifically used for:

[0071] configuring a network protocol and a port routing table of the plurality of MAC ports, configuring an interface routing table representing the interaction logic between the first interface and the at least one second interface, configuring an address mapping table and configuring a data interaction strategy;

[0072] configuring a load balancing algorithm to dynamically adjust a data interaction path in the data interaction strategy according to the traffic state in the plurality of MAC ports;

[0073] configuring a virtual port module to map physical port resources in the plurality of MAC ports into a plurality of virtual channels and to respectively assign the plurality of virtual channels to different GPUs.

[0074] In the embodiment of the application, the on-chip management system has multi-dimensional and fine configuration management functions. The on-chip management system as the control core of the device involves comprehensive and coordinated programming and control of the data plane, the control plane and the resource management plane to ensure that the device can operate in the optimal state. The on-chip management system is a high-performance real-time task processing core based on a Reduced Instruction Set Computer - V (RISCV), which is responsible for updating and configuring the routing table and the address mapping table of the PCIe switching logic, configuring the network port routing table and the load balancing algorithm, configuring the network protocol and other control information.

[0075] The specific tasks of the on-chip management system are first embodied in the basic configuration of the network communication plane and the internal switching plane. This includes the configuration of the network protocol and port routing table of multiple MAC ports. The MAC port is the specific implementation entity of the network port at the data link layer, and each physical network port corresponds to a MAC port, which is responsible for processing frame encapsulation, addressing and error checking. The on-chip management system needs to assign appropriate network protocol stack parameters to the MAC port, such as: Ethernet frame format, rate negotiation mode, and whether to enable remote direct memory access enhanced protocols such as RoCE. At the same time, it also needs to generate and maintain a port routing table that determines the rules based on which data sent from the internal device should be directed to which specific physical MAC port.

[0076] In addition to managing external network connections, the on-chip management system also regulates the internal interconnection logic of the device, that is, the interface routing table representing the interaction logic between the first interface and at least one second interface. This interface routing table is different from the network port routing table, and it is specifically used to manage the data flow of the device kernel, defining the rules for how data should flow between the first interface (endpoint mode) as an upstream channel and at least one second interface (root complex mode) as a downstream connection to various GPUs. For example: define which specific address range access should be routed to the CPU, and which should be intercepted and exchanged between local GPUs, thereby efficiently implementing data isolation and intercommunication between GPUs and CPUs, and between GPUs.

[0077] In order to further realize flexible resource addressing, an address mapping table also needs to be configured. The address mapping table is a conversion hub of logical addresses to physical or virtual resources. It maps the abstract GPU identifiers or memory addresses used by the upper software to specific entities identifiable within the device, such as: corresponding second interface number, virtual channel identifier or physical address in local memory. This layer of mapping is the basis for hardware resource virtualization and transparency to the upper software.

[0078] After all the above configurations, the higher-level task of the on-chip management system is to configure the data interaction strategy. It is a comprehensive strategy set that integrates routing tables, mapping tables and various protocol parameters to form a complete set of guidelines for determining when, with what priority, and via what path data should be transmitted.

[0079] Moreover, the on-chip management system is also capable of traffic optimization, i.e. configuring a load balancing algorithm. The load balancing algorithm is a dynamic decision-making program, which aims to dynamically adjust the data interaction path in the data interaction strategy according to the traffic status in multiple MAC ports. The traffic status includes but is not limited to the real-time bandwidth utilization, queue depth, packet delay and packet loss rate of each port. By monitoring these statuses, the configured load balancing algorithm can automatically redirect part of the subsequent data traffic to other relatively idle MAC ports or alternative paths when congestion is detected in a certain port or path, thereby avoiding network hotspots, smoothing the overall traffic, maximizing the utilization of all available network bandwidth resources, and ensuring efficient and stable data transmission.

[0080] Finally, the on-chip management system is directly responsible for driving the resource virtualization process, i.e. configuring the virtual port module. The physical port resources in multiple MAC ports are mapped to multiple virtual channels. The physical port resources refer to the real bandwidth, cache and queue capabilities of each MAC port. Through software-defined manner, these physical resources are cut, aggregated or recombined to form multiple logically independent virtual channels with flexible bandwidth and customizable characteristics. Then, resource allocation is performed to allocate multiple virtual channels to different GPUs. This allocation is dynamically adjustable, and the on-chip management system can adjust the allocation scheme in real time according to the computing task requirements of the GPU and the overall network status.

[0081] The on-chip management system of the present application, through the execution of collaborative configuration tasks, transforms the device from a passive connection component to an active, programmable intelligent interconnection hub. Not only does it establish the basic rules of data flow, but it also real-time perceives the system status and continuously optimizes the performance and efficiency of the entire interconnection system through dynamic adjustment of routing strategy, load distribution and resource mapping. Through deep and full-stack configuration management capabilities, the device achieves high performance, high flexibility and high resource utilization, enabling it to intelligently adapt to complex and variable heterogeneous computing environments.

[0082] In an implementable manner of an embodiment of the present application, the data processing module is connected with the switching logic module and the multiple MAC ports respectively, and is specifically configured to convert transaction requests from the first interface, the second interface and the multiple MAC ports into data movement transactions for processing;

[0083] The data processing module is further configured to, after converting the transaction requests into data movement transactions, perform direct memory access operations to move data between the GPU, the local memory and the multiple MAC ports.

[0084] In the embodiments of the present application, the connection relationship of the data processing module is specifically that it is connected with the switching logic module and the plurality of MAC ports respectively. This connection relationship establishes the central position of the data processing module in the data flow inside the device. On the one hand, it is connected with the switching logic module responsible for path scheduling, and receives data from or sent to the switching logic module; on the other hand, it is directly connected with the plurality of MAC ports responsible for physical network transmission and reception, and processes data flow in and out of the external network. This makes the data processing module able to directly process data from the network side and the internal switching side, reduces intermediate links, and provides a channel for high-performance data processing.

[0085] One of the functions of the data processing module is to convert transaction requests from the first interface, the second interface and the plurality of MAC ports into data movement transactions for processing. Transaction request refers to an operation instruction or data packet with specific format and semantics generated under different interface protocol specifications. For example: the read-write request transaction from the first interface or the second interface may conform to the PCIe protocol; the data frame conforming to the Ethernet protocol or the message carrying the RoCE protocol from the plurality of MAC ports. These transaction requests have different purposes and different formats, and cannot be directly and uniformly executed. Therefore, the data processing module needs to parse these heterogeneous transaction requests, understand their operation intentions (such as reading a block of memory or sending data to a specific destination), and convert them into an internal unified and efficiently executable intermediate representation, i.e. data movement transaction. Data movement transaction is a standardized description closer to the underlying hardware operation, which clearly defines the source address, target address, data length and movement control parameters (such as whether data format conversion is needed) of data. Through this conversion, complex operations from different interface protocols are unified and simplified to the core data movement problem, greatly simplifying the subsequent processing logic.

[0086] After the conversion is completed, the data processing module performs further direct memory access operations. Direct memory access is a technology that allows a specific hardware subsystem to independently transfer data between memory and input / output devices without direct intervention of the central processor. The data processing module uses direct memory access operations to efficiently move data directly between the GPU's video memory, local memory, and the buffer of the MAC port. Specifically, when data in a piece of GPU video memory needs to be sent to the network, the data processing module initiates a DMA operation to read the data from the GPU video memory directly into its internal buffer or local memory for necessary protocol encapsulation, and then initiates another DMA operation to write the encapsulated data from the buffer directly to the sending queue of the target MAC port. The entire process minimizes the number of unnecessary intermediate copies of data in memory and avoids frequent intervention by the CPU, thereby achieving low-latency, high-bandwidth data transmission.

[0087] The data processing module can convert transaction requests from the network (MAC port) and PCIe channel (first interface, second interface) into data movement transactions for processing, and internally includes a PCIe transaction layer read / write processing module, a network frame buffer, a framing and parsing module, and the like.

[0088] The data processing module of the present application, through the connection architecture and the two-stage processing flow, constitutes an engine for efficient data processing within the device. It uniformly converts heterogeneous interface protocol transactions into internal data movement tasks and relies on direct memory access operations to perform these tasks, thereby establishing an efficient and direct data path between the GPU, local memory, and network MAC port. This improves the throughput and efficiency of the device in processing cross-domain data exchange, reduces the occupation of other processing resources, and achieves high performance and low overhead of the device.

[0089] In one implementation manner of the embodiment of the present application, in the direct connection mode, at least one MAC port in the plurality of MAC ports is directly connected with at least one other same device to form a high-dimensional GPU interconnection network.

[0090] In the embodiment of the present application, the direct connection mode supported by the plurality of MAC ports provides a physical basis for an efficient and dedicated interconnection network. Specifically, in the direct connection mode, the connection behavior of the device is that at least two MAC ports in the plurality of MAC ports are directly connected with at least one other same device. Specifically, regarding the direct connection mode in the present application, the present application provides a direct connection mode schematic diagram of a switching communication device, as shown in Figure 3As shown in the figure, there are 4 switching communication devices (i.e., switching devices 1-4), each having 6 MAC ports. The number of direct connection channels between any two devices can be configured according to the actual bandwidth requirement. Any GPU can communicate directly through two interconnected devices. The on-chip management core can configure the bandwidth of the interconnection channel and balance the task load according to the actual communication requirement. The GPUs in the node can communicate through the original PCIe switch and the CPU interconnection (Quick Path Interconnect, QPI) channel, or through the direct connection channel realized by the device.

[0091] The device is not only used for communication with the outside world through a single network link, but also can simultaneously enable multiple physical MAC ports, each of which is independently and point-to-point connected to another switching communication device of the same type and function. The multi-link parallel connection mode discards the traditional single-line connection or cascade mode and provides physical possibility for constructing a more complex topology.

[0092] The multi-link direct connection of the present application forms a high-dimensional GPU interconnection network. High dimension refers to a complex network topology formed by multiple switching communication devices as nodes, which are interwoven through multiple parallel direct connection links between them, which exceeds the simple linear or ring structure. For example, when each device uses two MAC ports to connect two different neighbors, all devices can form a ring topology. If more than two ports are used for regular connection, a two-dimensional mesh (Mesh), a three-dimensional torus (Torus), or a partially connected hypercube (Hypercube) and other classic high-performance computing interconnection topologies can be formed. In the high-dimensional interconnection network, there are usually multiple parallel physical communication paths between any two GPUs managed by different devices.

[0093] The high-dimensional GPU interconnection network formed by direct connection in the present application improves the aggregated communication bandwidth within the GPU cluster. Since data can be transmitted simultaneously on multiple physical links, the total throughput of the network is multiplied, which can meet the huge parameter synchronization and gradient exchange demand generated in large-scale model training. Secondly, it effectively reduces the communication delay and avoids network hotspots. The existence of multiple paths allows communication tasks to be dispersed to different links through load balancing strategies, reducing the congestion of a single path, while providing shorter, possibly bypassing intermediate nodes, direct or nearly direct communication paths. In addition, this topology also enhances the fault tolerance and scalability of the network. The failure of part of the link usually does not cause the network to be split, and the data flow can be detoured through other redundant paths.

[0094] Therefore, the direct connection mode of the present application promotes the direct connection mode capability of the device from establishing a simple point-to-point link to a high level of constructing a complex and efficient system-level interconnection network. The GPU cluster constructed based on the device can flexibly deploy the most suitable high-dimensional network topology according to the communication mode requirement of the computing task, thereby providing a communication infrastructure with low latency, high bandwidth and high reliability at the hardware interconnection level for large-scale parallel computing applications.

[0095] In an implementation manner of the embodiment of the present application, in the switching mode, the device is connected to an Ethernet switching network through at least one MAC port in the plurality of MAC ports, so as to construct an expandable GPU interconnection network with other switching communication devices connected to the Ethernet switching network.

[0096] The device and the other switching communication devices are in the same or different server nodes.

[0097] In the embodiment of the present application, the switching mode supported by the plurality of MAC ports provides another network integration scheme which is standardized and easy to deploy on a large scale. Specifically, in the switching mode, the connection behavior of the device is to connect to an Ethernet switching network through at least one MAC port in the plurality of MAC ports. The Ethernet switching network refers to a standard data center network infrastructure based on Ethernet technology and taking a switch as a core node. The Ethernet switching network provides a general service of exchanging and forwarding data packets based on MAC addresses or Internet Protocol (IP) addresses. By connecting the MAC port of the device to this network, the device integrates itself and the GPU resources managed thereby into the general network plane of the data center, and can utilize the existing mature network management, routing and security policies.

[0098] Specifically, regarding the direct connection mode in the present application, the present application provides a switching mode diagram of a switching communication device, as shown in Figure 4 There are four switching communication devices (i.e., switching devices 1-4), each of which has 32 MAC ports, directly connected to an Ethernet switching network composed of multiple multi-port Ethernet switches. The GPUs controlled by the switching device under different CPUs can communicate through the Ethernet network interconnection. The internal topology of the Ethernet network can be freely combined and flexibly configured, facilitating large-scale interconnection of GPU devices.

[0099] The exchange connection mode facilitates communication with other exchange communication devices connected to the Ethernet exchange network. Here, the other exchange communication devices refer to the same type of devices as the device, which have the same structure and function. When multiple such devices connect their MAC ports to the same shared Ethernet exchange network, they no longer need dedicated direct lines between them, but communicate through this common network infrastructure. The core exchange device (switch) of the Ethernet exchange network is responsible for routing and forwarding data between the ports of the devices. This networking method based on a standard exchange network makes it extremely easy and flexible to build a large-scale GPU interconnection network. When the system is expanded, new devices only need to be connected to the existing Ethernet exchange network, and the corresponding network address and route are configured, and the new GPU resources can be integrated into the cluster, thereby realizing linear or even super-linear scalability of the network size. This scalability is not strictly limited by the complexity of physical wiring, and is suitable for building large computing clusters with a large number of GPU nodes.

[0100] The exchange communication device and the other exchange communication devices are in the same or different server nodes. A server node usually refers to a physical server or an independent computing unit including a mainboard, a CPU, a memory, and the like. This means that multiple devices (and the GPUs they manage) interconnected by exchange mode and direct connection mode can be deployed within the same server node (for example, different expansion slots in the same multi-GPU server) or across the chassis boundary (for example, distributed in multiple servers in different racks). When the devices are in the same server node, they can use the exchange network to achieve more flexible logical reorganization and communication of GPU resources within the node; when they are in different server nodes, they can achieve true GPU resource pooling and clustering across physical servers. The wide-area coverage capability of the Ethernet exchange network enables geographically dispersed GPU resources to be logically integrated into a unified interconnection plane.

[0101] The present application extends the networking capability of the device from a dedicated direct connection scenario to a general data center network environment through exchange mode. Using the Ethernet exchange network as a carrier, the GPU interconnection scheme based on the present application has deployment flexibility, scale expansion capability, and seamless integration with existing infrastructure. It is suitable for large-scale cloud deployment and artificial intelligence training scenarios that require dynamic integration of heterogeneous and distributed GPU computing resources, and provides interconnection networking support for realizing flexible and efficient computing resource pools.

[0102] Figure 5 A flowchart of a switching communication method provided for an embodiment of the present application is provided, and the switching communication method is described in detail in combination with the flowchart.

[0103] As Figure 5As shown, the exchange communication method is applied to a server as shown Figure 1 As shown, the on-chip management system in the exchange communication device comprises the following steps:

[0104] Step 501, resource discovery and virtualization management are performed on the GPUs connected through the at least one second interface, available target GPUs are determined, and GPU resource information of the target GPUs is reported to the CPU.

[0105] In embodiments of the present application, resource discovery refers to the on-chip management system actively initiating an enumeration and query process to the physically connected downstream bus through the second interface managed by it and configured in root complex mode, identifying each specific connected GPU device, and obtaining its model, video memory capacity, computing capability and other key attributes.

[0106] Virtualization management refers to the on-chip management system not directly exposing the original information of the physical GPU to the upper layer, but performing a layer of logical abstraction, for example: creating one or more virtual device identifiers for each physical GPU, or dividing its computing resources into finer-grained virtual units. Through this process, the system can determine the available target GPUs, i.e., the currently available GPU resources for scheduling. Finally, the system reports these abstracted and sorted GPU resource information to the connected CPU's driver and operating system through its first interface (endpoint mode) in a standardized device reporting protocol. This allows the server to manage these GPUs, but perceives the logical resources after virtualization processing by the device, rather than the original physical connection, thereby completing the decoupling of the GPU and the server hardware at the software level.

[0107] Meanwhile, one of the important outputs of virtualization management is to dynamically associate and map the created logical GPU units with the virtual channels (i.e., virtualized network port resources) created by the virtual port module. This means that a logical GPU unit is not fixedly bound to a certain physical MAC port when it needs to access the network, but is assigned one or more virtual channels with specific bandwidth and quality of service attributes. This mapping relationship is dynamically maintained by the on-chip management system.

[0108] Step 502, network topology detection is performed through multiple MAC ports to determine available target network topologies, which include direct mode and / or exchange mode.

[0109] In embodiments of the present application, network topology detection is a proactive network discovery process in which the on-chip management system instructs the multiple MAC ports managed by it to send specific discovery packets (such as Link Layer Discovery Protocol (LLDP) protocol extension frames or custom beacon frames) and listen to responses from network neighbors.

[0110] By analyzing the response message, the system can determine the current connection state of each MAC port: is it directly connected to another device of the same kind (forming a direct connection mode), or is it connected to a standard Ethernet switch (entering a switching mode). By integrating the detection results of all ports, a real-time and accurate network connection map is constructed and maintained locally, i.e., the available target network topology is determined. This target network topology is dynamic and may contain both point-to-point links in direct connection mode and switch access links in switching mode, so the on-chip management system needs to understand and manage this mixed topology environment.

[0111] Step 503, according to the received task request, GPU resource information and target network topology, generate routing strategy applied to switching logic module for data flow transmission, allocation strategy applied to virtual port module for virtual resource allocation and processing rule applied to data processing module for data processing.

[0112] In the embodiments of the present application, the task request originates from the upper application or cluster scheduler and is issued by CPU, and its essence is data communication demand, for example: sending data of GPU A to GPU B of node X. The on-chip management system comprehensively analyzes the task request, combines the known GPU resource information (such as the logical position of GPU) and the target network topology (such as the available path to node X), and uses the built-in algorithm for intelligent calculation. The calculation result is embodied in three sets of interrelated and executable configuration strategies: one is the routing strategy guiding the data packet forwarding path, which will be issued to the switching logic module, filled into its switching table entry, and determine the direction of data flow within the device; the second is the allocation strategy guiding the network bandwidth resource division, which will be issued to the virtual port module, and specify how to map and allocate the physical MAC port bandwidth to specific virtual GPU or data flow; the third is the processing rule specifying the data format conversion, encapsulation, etc. details, which will be issued to the data processing module to ensure correct conversion of data between different protocols. The generation of these strategies is a dynamic and optimized process, aiming to find the most efficient resource utilization and communication path scheme for the current task.

[0113] Step 504, in response to the data communication request of the target GPU, according to the routing strategy, the allocation strategy and the processing rule, determine the data interaction path between the target GPU, the CPU and the plurality of MAC ports.

[0114] In the embodiments of the present application, when a GPU mounted on the device initiates an actual data communication request (for example: writing a specific address through DMA) due to the need of a computing task, the request is captured by the device. The on-chip management system acts as a coordinator and does not directly process data, but performs real-time query and matching according to the generated and effective routing strategy, allocation strategy and processing rule.

[0115] Through this process, an explicit end-to-end data interaction path is determined for the current communication request. This data interaction path clearly specifies whether the data read out from the source GPU memory is directed to the CPU through the switching logic module or to a specific virtual channel allocated by the virtual port module and sent to the network through the corresponding MAC port, and also specifies the specific operation rules to be followed in the data processing module. The determination of this path enables subsequent data transmission actions to be automatically and pipelined efficiently executed by various hardware modules within the device.

[0116] The present application describes the global management process of the on-chip management system from resource perception, topology discovery, to intelligent strategy generation, and to real-time path determination. The isolated hardware connection is transformed into a dynamic service defined by intelligent software, enabling the device to actively adapt to changes in computing tasks and network environment, thereby achieving efficient pooling of GPU resources, flexible use of network topology, and dynamic optimization of data transmission paths at the system level, and ultimately providing high-performance and flexible communication support for large-scale heterogeneous computing.

[0117] In one implementation manner of the embodiments of the present application, in order to enable the on-chip management system not only to generate strategies based on initial information, but also to continuously perceive changes in network state during operation and dynamically and intelligently adjust its control behavior accordingly, thereby realizing real-time optimization of performance, the following methods can also be used but are not limited to: monitoring the data traffic of multiple MAC ports; executing a load balancing algorithm based on the data traffic to update the routing strategy, the allocation strategy and the processing rule, and redistributing the data traffic to different MAC ports or data interaction paths.

[0118] In the embodiments of the present application, the data traffic of multiple MAC ports is monitored, and the monitoring refers to continuous active measurement and data collection. The on-chip management system acquires data traffic information on each physical MAC port in real time and periodically through its hardware counter or dedicated management interface. These data traffic information is quantitative and multi-dimensional, and is not limited to total bandwidth utilization, but usually includes but is not limited to: real-time inbound and outbound throughput rate of the port, data packet forwarding delay, depth of sending and receiving queues, data packet loss rate and error frame count, etc. Monitoring is used to understand the real-time load and health status of the external network environment, and provides an objective and quantitative data basis for subsequent intelligent decision-making.

[0119] After real-time traffic data is acquired, a load balancing algorithm is executed based on the data traffic, which is a set of pre-defined or dynamically generated mathematical rules and decision logic, designed to optimize resource usage, maximize throughput, minimize response time, and avoid overloading any single resource. The on-chip management system takes the collected data traffic data as input, feeds it into the load balancing algorithm for calculation and analysis. It assesses whether the current traffic distribution is balanced, identifies MAC ports or transmission paths that may be congested, and discovers ports and paths that have insufficient utilization and still have redundant capacity.

[0120] The results of the algorithm analysis will directly drive the dynamic update of the control strategy, i.e., updating the routing strategy, allocation strategy, and processing rules. Here, the update is a dynamic reconfiguration process, rather than a one-time setting. Specifically:

[0121] For the update of the routing strategy, it may mean modifying the routing table entries in the switch logic module, redirecting part of the data flow originally planned to a congested port to another path corresponding to a less loaded port.

[0122] For the update of the allocation strategy, it may involve adjusting the configuration of the virtual port module, changing the mapping ratio of physical port bandwidth to each virtual channel, or switching a virtual channel being used by a virtual GPU from a high-load physical port to a low-load physical port.

[0123] For the update of the processing rules, it may include adjusting data encapsulation priorities, modifying queue scheduling weights, etc., to cooperate with new paths and resource allocation.

[0124] The ultimate operational goal of all strategy updates is to redistribute data traffic to different MAC ports or data interaction paths. Redistribution is a real-time traffic scheduling action. This means that for newly arrived data communication requests, a new and more optimal exit MAC port will be selected for them based on the updated strategy; for data streams that are already in transit and can be split, even the subsequent part of the data may be directed to a new path through techniques similar to multi-path transmission. Data interaction path is an end-to-end logical concept, which may span the internal switch logic, specific virtual channels, and external physical network links.

[0125] The control loop of monitoring, analyzing, deciding and executing of the application enables the on-chip management system to go beyond the limitation of static configuration and become an intelligent agent that can respond to network traffic changes in real time, actively prevent congestion and continuously optimize global resource utilization. By dynamically redistributing data traffic from busy ports and paths to relatively idle resources, the overall utilization efficiency of all MAC ports is improved, the transmission delay and jitter of critical data streams are effectively reduced, and the stability and performance predictability of the entire GPU interconnection network under variable load are enhanced.

[0126] Further, it needs to be explained that the switching communication device of the application has at least the following functions and configurations:

[0127] 1) It can be used as a switching network device to directly mount multiple GPU card devices to provide network ports for GPU external interconnection, 2) It has the attribute functions of PCIe Switch (switch) and network card Switch (switch), and the host manages GPU devices through the device, 3) The device has an on-chip high-performance real-time task processing core, which loads the management of the routing table, address conversion table, load balancing algorithm and network protocol configuration of the PCIe domain and MAC network domain; 4) The virtual port module can distribute network MAC ports to different GPUs through configuration to realize communication bandwidth reconfiguration, 5) The transaction conversion processing module receives PCIe transactions and network message frames and parses and forwards them to the internal data moving engine to realize access control local memory, 6) According to the multiple MAC ports and Ethernet switches of the device, different GPU communication network topologies can be formed, including cross-node direct connection network topology and cross-node switching network topology, wherein the direct connection network topology can be used when there are not many GPU devices, and the switching network topology is suitable for communication topology when more GPU devices are expanded.

[0128] In an implementable manner of an embodiment of the application, a server is also provided, wherein the server comprises: at least one GPU, a CPU and a switching communication device as shown in Figure 1

[0129] The CPU is connected to the first interface of the switching communication device, and the at least one GPU is connected to at least one second interface of the switching communication device respectively to realize data interaction between the at least one GPU and the CPU;

[0130] The switching communication device is directly connected to at least one other switching communication device in a direct connection mode through multiple MAC ports, and the at least one other switching communication device is located in the server and / or in other servers;

[0131] ​Or, the switching communication device is connected to an Ethernet switching network in a switching mode through multiple MAC ports to establish a communication connection with at least one other switching communication device connected to the Ethernet switching network.

[0132] The switching communication device realizes data interaction between the at least one GPU and the other GPU under the connection of the at least one other switching communication device through the direct connection mode and the switching mode.

[0133] In the embodiments of the present application, the server components are connected through physical and electrical connections to form a cooperative whole, and specifically, the CPU is connected to the first interface of the corresponding switching communication device through a system bus or a dedicated interconnection interface. The first interface is usually a high-speed interface facing the control plane or the management plane, and is used to transmit control instructions, management information and non-pipelined data requiring CPU processing. Meanwhile, the at least one GPU is connected to at least one second interface of the corresponding switching communication device. The second interface is a super-high-speed interface facing the data plane, usually has extremely high bandwidth, and is optimized for carrying large-scale data flow between GPUs or between GPUs and the outside. Through this connection topology, data interaction is realized between the at least one GPU and the corresponding CPU. The CPU can assign computing tasks or exchange control parameters to the GPU, and the computing results generated by the GPU can also be efficiently returned to the CPU or sent out through the switching communication device, forming a tightly coupled heterogeneous computing unit.

[0134] The network interconnection capability of the server is flexibly realized through multiple MAC ports of the switching communication device, and supports two main networking modes. The first is the direct connection mode. In this mode, the switching communication device directly connects to at least one other switching communication device in the server and / or at least one other switching communication device in other servers through its multiple MAC ports in a direct connection mode. This means that in a cluster composed of multiple such servers, their switching communication devices can be directly connected point-to-point through backplane wiring, cables or optical modules, forming a low-latency, high-bandwidth dedicated interconnection network. This mode bypasses the hop of external switches, greatly reducing the end-to-end communication delay.

[0135] The second is the switching mode. In this mode, the switching communication device is connected to an Ethernet switching network in a switching mode through multiple MAC ports. The Ethernet switching network refers to a standard data center network architecture built based on the Ethernet protocol, which may include access switches, aggregation switches and other network devices. When the switching communication device accesses this network, it can realize communication connection between the switching communication device and other switching communication devices connected to the Ethernet switching network.

[0136] Through the direct connection mode or the switching mode, data interaction between other GPUs connected to the switching communication device and at least one GPU connected to the switching communication device can be realized. Even if the GPUs are located in different physical servers, their data packets can be routed and forwarded through a standard Ethernet network to realize large-scale and flexible cluster communication. This mode provides excellent scalability and topology flexibility.

[0137] Through the above architecture, the server provided by the present application realizes high integration of computing and network. By embedding the intelligent switching communication device in the server and directly interconnecting the GPU and the CPU with the intelligent switching communication device, the performance bottleneck between the network card and the computing unit in the traditional architecture is eliminated, the boundary between the computing node and the network is minimized, and the communication performance is provided for the distributed computing application. The two modes of direct connection and switching are supported, so that the server can be used to build a low-latency tightly coupled computing group and can be seamlessly integrated into an existing standard data center network, having excellent deployment flexibility and topology adaptability. Finally, since the switching communication device itself has high-precision intelligent monitoring capability, the internal data interaction state of the entire server and the external network traffic can be perceived and intelligently analyzed at a millisecond level, realizing panoramic and high-precision operation and maintenance visibility from the chip level to the network level, and significantly improving the reliability and maintainability of the large-scale computing cluster.

[0138] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and the necessary general hardware platform, of course, it can also be realized by hardware, but in many cases, the former is a better embodiment.

[0139] The embodiment of the present application also provides an electronic device, comprising a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above switching communication method embodiments.

[0140] The embodiment of the present application also provides a computer readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps in any of the above switching communication method embodiments when running.

[0141] In an exemplary embodiment, the above computer readable storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.

[0142] The embodiment of the present application further provides a computer program product, which comprises a computer program, and the computer program realizes the steps in any of the exchange communication method embodiments when executed by a processor.

[0143] The embodiment of the present application further provides another computer program product, which comprises a nonvolatile computer readable storage medium, and the nonvolatile computer readable storage medium stores a computer program, and the computer program realizes the steps in any of the exchange communication method embodiments when executed by a processor.

[0144] Those skilled in the art can further understand that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized in electronic hardware, computer software or combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in general in the above description. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0145] The above describes in detail the exchange communication device and method and the server provided by the present application. The principles and implementation manners of the present application are described by applying specific examples in the present application. The above description of the examples is only for helping to understand the method of the present application and its core idea. It should be pointed out that, for those skilled in the art, some improvements and modifications can be made to the present application without departing from the principles of the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A switching communication device, characterized in that, The device is communicatively connected to at least one GPU and at least one CPU, and the device includes: The first interface is configured in endpoint mode for connection to the CPU; At least one second interface is configured in root complex mode for connecting to the GPU, respectively; A switching logic module, connected to the first interface and the at least one second interface, is used to realize data interaction between the GPU and the CPU, and data interaction between the GPUs connected to the at least one second interface respectively; Multiple MAC ports are used to directly connect to at least one other switching communication device in direct connection mode, or to connect to an Ethernet switching network in switching mode, so as to realize data interaction between other GPUs connected to the other switching communication device and the GPU. The data processing module is used to perform data conversion and processing on the data between the first interface, the second interface, and the plurality of MAC ports; The virtual port module is used to virtualize the multiple MAC port resources and allocate them to the GPU; The on-chip management system is used to manage and configure configuration information that affects data interaction; The on-chip management system is specifically used for: Configure the network protocol and port routing table of the multiple MAC ports, configure the interface routing table representing the interaction logic between the first interface and the at least one second interface, configure the address mapping table, and configure the data interaction strategy; Configure a load balancing algorithm to dynamically adjust the data interaction path in the data interaction strategy based on the traffic status of the multiple MAC ports; Configure the virtual port module to map the physical port resources of the multiple MAC ports into multiple virtual channels, and allocate the multiple virtual channels to different GPUs respectively.

2. The switching communication device according to claim 1, characterized in that, The virtual port module is dynamically configured by the on-chip management system based on the number of the multiple MAC ports and the number of GPUs, and is specifically used for: The physical port resources of the multiple MAC ports are virtualized, and the virtualized physical port resources are dynamically allocated to the GPU.

3. The switching communication device according to claim 1, characterized in that, The device further includes: The bus control module is used to aggregate and schedule requests from the switching logic module and the data processing module to access local memory based on the on-chip bus.

4. The switching communication device according to claim 1, characterized in that, The switching logic module includes: a configurable switching switch matrix; The switching matrix arbitrates and directs the received data stream based on the routing table or address mapping table configured by the on-chip management system to achieve data interaction.

5. The switching communication device according to claim 3, characterized in that, The data processing module is connected to the switching logic module and the plurality of MAC ports respectively, and is specifically used to convert transaction requests from the first interface, the second interface and the plurality of MAC ports into data migration transactions for processing; The data processing module is further configured to, after converting the transaction request into the data transfer transaction, perform a direct memory access operation to transfer data between the GPU, the local memory, and the multiple MAC ports.

6. The switching communication device according to claim 1, characterized in that, In the direct connection mode, at least one of the plurality of MAC ports is directly connected to at least one other switching communication device to form a high-dimensional GPU interconnect network.

7. The switching communication device according to claim 1, characterized in that, In the switching mode, at least one of the plurality of MAC ports is connected to the Ethernet switching network to facilitate the construction of a scalable GPU interconnect network with the other switching communication devices connected to the Ethernet switching network; The device is located on the same or different server nodes as the other exchange and communication devices.

8. A switching communication method, characterized in that, The method is applied to the on-chip management system of the switching communication device as described in any one of claims 1 to 7, comprising: Resource discovery and virtualization management are performed on GPUs connected through at least one second interface to identify available target GPUs and report the GPU resource information of the target GPUs to the CPU. Network topology detection is performed using multiple MAC ports to determine available target network topologies, including direct-connect mode and / or switched mode. Based on the received task request, the GPU resource information, and the target network topology, a routing strategy for data flow transmission in the switching logic module, an allocation strategy for virtual resource allocation in the virtual port module, and a processing rule for data processing in the data processing module are generated. In response to the data communication request of the target GPU, the data interaction path of the data flow between the target GPU, the CPU and the multiple MAC ports is determined according to the routing policy, the allocation policy and the processing rules.

9. The switching communication method according to claim 8, characterized in that, The method further includes: Monitor the data traffic of the multiple MAC ports; A load balancing algorithm is executed based on the data traffic to update the routing policy, the allocation policy, and the processing rules, and the data traffic is redistributed to different MAC ports or data interaction paths.

10. A server, characterized in that, The server includes: at least one GPU, a CPU, and a switching communication device as described in any one of claims 1 to 7; The CPU is connected to a first interface of the switching communication device, and the at least one GPU is connected to at least one second interface of the switching communication device to realize data interaction between the at least one GPU and the CPU; The switching communication device is directly connected to at least one other switching communication device in direct connection mode via multiple MAC ports, and the at least one other switching communication device is located in the server and / or other servers; Alternatively, the switching communication device may be connected to the Ethernet switching network in a switching mode through the plurality of MAC ports to establish a communication connection with the at least one other switching communication device connected to the Ethernet switching network. The switching communication device enables data interaction between other GPUs connected to the at least one other switching communication device and the at least one GPU through the direct connection mode and the switching mode.

Citation Information

Patent Citations

  • Virtual machine data communication method and system, and virtual machine configuration method and apparatus

    CN109445905A

  • Heterogeneous acceleration pooling server system, pooling resource allocation method and medium

    CN119883590A