Device interconnection system and method, device, non-volatile readable storage medium and program product
By introducing resource pools and switching chips into the device interconnection system in the bus topology architecture, the problem of the limited number of PCIe switch ports is solved, low-latency device interconnection and communication are achieved, and cabling layout is simplified.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2026-03-12
AI Technical Summary
In the existing bus topology, the limited number of PCIe switch ports leads to complex wiring and configuration, increased communication latency, and an inability to meet the business requirements with strict low latency requirements.
The device interconnection system includes at least two resource pools, each containing a switching chip. The upstream port of the switching chip connects to the switching chips of other resource pools, and the downstream port connects to hardware resource devices. The switching chip forwards data according to the destination address and realizes data conversion and communication through modules such as routing forwarding engine module, storage module, and direct memory access engine module.
It simplifies the complexity of device interconnection and wiring layout, reduces the number of connection cables, lowers communication latency, and meets the business communication requirements for low latency.
Smart Images

Figure CN2025082496_12032026_PF_FP_ABST
Abstract
Description
Device interconnection system, method, device, nonvolatile readable storage medium and program product
[0001] Cross-reference to Related Applications
[0002] The present application claims priority to the Chinese patent application No. 202411247080.0, filed on September 6, 2024, and entitled "A device interconnection system, method, device, medium and program product", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0003] The present application relates to the technical field of computer, and in particular, to a device interconnection system, method, device, nonvolatile readable storage medium and program product. BACKGROUND
[0004] The current bus topology architecture is mostly a tree structure based on PCIe bus and PCIe switch chip (PCIe supporting exchange chip). In this structure, the number of PCIe switch ports is limited. With the expansion of system size and the increase of level number, the number of PCIe switches is large, resulting in complex wiring and configuration. In addition, the increase of hop number in multi-level topology also leads to the increase of communication delay, which cannot meet the strict business requirements for low delay.
[0005] There is a technical problem of high complexity of device interconnection and wiring layout in the related art. SUMMARY
[0006] Therefore, the purpose of the present application is to provide a device interconnection system, method, device, nonvolatile readable storage medium and program product to simplify the complexity of device interconnection and wiring layout. The optional solutions are as follows:
[0007] In a first aspect, the present application provides a device interconnection system, comprising:
[0008] at least two resource pools, each resource pool comprising at least one exchange chip;
[0009] At least one downstream port of at least one exchange chip is set to connect a hardware resource device; the hardware resource devices in the same resource pool are memory devices, hardware acceleration devices or computing devices, and the hardware resource devices in different resource pools are different;
[0010] At least one upstream port of at least one exchange chip is set to connect the upstream port of at least one exchange chip in other resource pool; at least one upstream port of at least one exchange chip is set to connect the upstream port of other exchange chip in the resource pool where the current exchange chip is located;
[0011] The at least one switch chip is configured to forward upstream data sent by any one of the upstream ports to a corresponding downstream port according to a destination address of the upstream data, and / or forward downstream data sent by any one of the downstream ports to a corresponding upstream port according to a destination address of the downstream data.
[0012] In some embodiments, the at least one switch chip comprises a routing and forwarding engine module.
[0013] The routing and forwarding engine module is configured to determine, according to a routing table, a destination address of upstream data sent by any one of the upstream ports in the current switch chip, and / or a destination address of downstream data sent by any one of the downstream ports in the current switch chip.
[0014] In some embodiments, the at least one switch chip comprises a storage module.
[0015] The storage module is configured to store the routing table, upstream data sent by any one of the upstream ports in the current switch chip, and downstream data sent by any one of the downstream ports in the current switch chip.
[0016] In some embodiments, the at least one switch chip comprises a direct memory access engine module.
[0017] The direct memory access engine module is configured to realize communication between other resource pools connected by any one of the upstream ports in the current switch chip and hardware resource devices connected by any one of the downstream ports in the current switch chip by using a direct memory access technology.
[0018] In some embodiments, the at least one switch chip comprises a control logic module.
[0019] The control logic module comprises a data packet conversion unit, a register unit and a resource awareness module.
[0020] The data packet conversion unit is configured to perform data format conversion on upstream data sent by any one of the upstream ports in the current switch chip and / or downstream data sent by any one of the downstream ports in the current switch chip.
[0021] The register unit is configured to configure port information of any one of the upstream ports and / or any one of the downstream ports in the current switch chip; the port information comprises port parameters, port attributes and port states.
[0022] The resource awareness unit is configured to monitor whether each of the downstream ports in the current switch chip is connected to a hardware resource device, and / or monitor resource usage information of the hardware resource devices connected by the downstream ports in the current switch chip, and generate a corresponding resource awareness table.
[0023] In some embodiments, each downstream port of the at least one switch chip supports the first protocol, and each upstream port of the at least one switch chip supports the second protocol.
[0024] Correspondingly, the data packet conversion unit is configured to convert upstream data sent by any one of the upstream ports in the current switch chip from the second protocol format to the first protocol format, and / or convert upstream data sent by any one of the downstream ports in the current switch chip from the first protocol format to the second protocol format.
[0025] In some embodiments, the register unit is configured to configure the routing table and the resource awareness table.
[0026] In some embodiments, the resource awareness unit is configured to update the resource awareness table in real time according to whether each downstream port is connected to a hardware resource device and / or resource usage information of the hardware resource device obtained in real time.
[0027] In some embodiments, any hardware resource device is configured to send idle resource information in itself to the resource awareness unit at a preset period;
[0028] Correspondingly, the resource awareness unit updates the resource awareness table according to the information sent by the hardware resource device.
[0029] In some embodiments, the resource awareness unit is configured to periodically detect whether a communication link between each upstream port of the current switch chip and a connected resource pool is connected, and periodically detect whether a communication link between each downstream port of the current switch chip and a connected hardware resource device is connected.
[0030] In some embodiments, each upstream port of the at least one switch chip is provided with a decoder.
[0031] The decoder is configured to determine an address mapping relationship between a memory of a computing device connected to the upstream port where the current decoder is located and a hardware resource device connected to a downstream port in the current switch chip.
[0032] In some embodiments, the decoder is configured to store memory capacity information of a computing device connected to the upstream port where the current decoder is located.
[0033] In some embodiments, the at least one switch chip comprises a cache coherency interface module.
[0034] The cache coherency interface module is configured to provide a cache coherency transmission channel between hardware resource devices connected to each downstream port in the current switch chip.
[0035] In some embodiments, the at least one switch chip comprises a clock and reset module.
[0036] The clock and reset module is configured to maintain clock synchronization among devices in the system.
[0037] In some embodiments, the at least one switch chip further comprises a switch function module.
[0038] In some embodiments, the at least one switch chip comprises a power management module.
[0039] The power management module is configured to provide power for the switch function module and the switch chips in the resource pools, and manage power consumption of the switch function module and the switch chips in the resource pools.
[0040] In some embodiments, each hardware resource device in each resource pool supports a CLOS topology and / or a full interconnection topology.
[0041] In a second aspect, the present application provides a device interconnection method, applied to any of the device interconnection systems described above, comprising:
[0042] For any resource pool in the device interconnection system, connecting at least one upstream port of at least one switch chip in the current resource pool to an upstream port of at least one switch chip in another resource pool; connecting at least one downstream port of at least one switch chip in the current resource pool to a hardware resource device; and connecting at least one upstream port of at least one switch chip in the current resource pool to an upstream port of another switch chip in the current resource pool.
[0043] In the device interconnection system, each resource pool comprises at least one switch chip; the hardware resource devices in the same resource pool are memory devices, hardware acceleration devices or computing devices, and the hardware resource devices in different resource pools are different.
[0044] In the device interconnection system, each resource pool comprises at least one switch chip; the hardware resource devices in the same resource pool are memory devices, hardware acceleration devices or computing devices, and the hardware resource devices in different resource pools are different.
[0045] In the device interconnection system, each resource pool comprises at least one switch chip; the hardware resource devices in the same resource pool are memory devices, hardware acceleration devices or computing devices, and the hardware resource devices in different resource pools are different.
[0046] The memory is configured to store the computer program.
[0047] The processor is configured to execute the computer program to implement the device interconnection method disclosed above.
[0048] In a fourth aspect, the present application provides a non-volatile readable storage medium for storing a computer program, wherein the computer program is executed by a processor to implement the device interconnection method disclosed above.
[0049] In a fifth aspect, the present application provides a computer program product comprising computer programs / instructions which, when executed by a processor, implement the steps of the above disclosed device interconnection method.
[0050] According to the above scheme, the present application provides a device interconnection system, which comprises: at least two resource pools, each resource pool comprising at least one switching chip; at least one downstream port of at least one switching chip being configured to connect a hardware resource device; the hardware resource devices in the same resource pool being: a memory device, a hardware acceleration device or a computing device, the hardware resource devices in different resource pools being different; at least one upstream port of at least one switching chip being configured to connect an upstream port of at least one switching chip in another resource pool; at least one switching chip being configured to: forward upstream data to a corresponding downstream port according to a destination address of the upstream data sent by any one upstream port; and / or forward downstream data to a corresponding upstream port according to a destination address of the downstream data sent by any one downstream port.
[0051] It can be seen that the present application has the following beneficial effects: different hardware resource devices, such as memory devices, hardware acceleration devices or computing devices, are included in different resource pools; at least one upstream port of at least one switching chip in each resource pool is connected to an upstream port of at least one switching chip in another resource pool, thereby realizing interconnection between different resource pools; at least one downstream port of at least one switching chip in each resource pool is connected to a hardware resource device, so that the switching chip forwards upstream data to a corresponding downstream port according to a destination address of the upstream data sent by any one upstream port; and / or forwards downstream data to a corresponding upstream port according to a destination address of the downstream data sent by any one downstream port, thereby realizing communication within the system. The scheme realizes interconnection of devices within the system with as few ports as possible, the downstream devices do not need to be connected two by two, the number of connection lines and the complexity of wiring layout are reduced, the communication delay can be correspondingly reduced, and the communication needs of low-delay requirement services are met.
[0052] Correspondingly, the device interconnection method, device, non-volatile readable storage medium and program product provided by the present application also have the above technical effects. BRIEF DESCRIPTION OF DRAWINGS
[0053] In order to more clearly illustrate the technical solutions in the embodiments or the related art, the following will briefly introduce the drawings needed to be used in the embodiments or the related art description. Obviously, the drawings in the following description only belong to the embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor based on the provided drawings.
[0054] Fig. 1 is a schematic diagram of a device interconnection system disclosed by the present application;
[0055] Fig. 2 is a schematic diagram of a second device interconnection system disclosed by the present application;
[0056] Fig. 3 is a schematic diagram of a third device interconnection system disclosed by the present application;
[0057] Fig. 4 is a schematic diagram of a switch chip disclosed by the present application;
[0058] Fig. 5 is a schematic diagram of an electronic device disclosed by the present application;
[0059] Fig. 6 is a schematic diagram of a server provided by the present application;
[0060] Fig. 7 is a schematic diagram of a terminal provided by the present application. DETAILED DESCRIPTION
[0061] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other examples obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0062] At present, the bus topology architecture is mostly a tree structure based on a PCIe (Peripheral Component Interconnect express, high-speed serial computer expansion bus standard) bus and a PCIe switch chip. In this structure, the number of PCIe switch ports is limited. With the expansion of the system size and the increase in the number of levels, the number of PCIe switches is large, resulting in complex wiring and configuration. In addition, the increase in the number of hops in the multi-level topology also leads to an increase in communication delay, which cannot meet the strict requirements of low-delay business demands. Therefore, the present application provides a device interconnection scheme, which can simplify the complexity of device interconnection and wiring layout.
[0063] Referring to Fig. 1, the embodiment of the present application discloses a device interconnection system, comprising: at least two resource pools, each resource pool comprising at least one switching chip; at least one downstream port of at least one switching chip is configured to connect a hardware resource device; the hardware resource devices in the same resource pool are: memory devices, hardware acceleration devices or computing devices, and the hardware resource devices in different resource pools are different; at least one upstream port of at least one switching chip is configured to connect the upstream port of at least one switching chip in another resource pool. At least one upstream port of at least one switching chip is configured to connect the upstream port of another switching chip in the resource pool where the current switching chip is located.
[0064] In FIG. 1, each switch chip has ports P1-PN, where P1-P5 are upstream ports and P6-PN are downstream ports. Of course, the number of upstream ports can be expanded according to actual needs.
[0065] It should be noted that the memory device can be a DRAM (Dynamic Random Access Memory), and the hardware acceleration device can be a GPU (Graphics Processing Unit) or an FPGA (Field-Programmable Gate Array), and the computing device can include a CPU (Central Processing Unit), a NIC (Network Interface Controller), and a DRAM. In an embodiment, each hardware resource device in each resource pool supports a CLOS topology (Clos Network, a topology structure that uses a multi-level interconnection network to realize non-blocking switching) and / or a full interconnection topology.
[0066] In the method, at least one switch chip is configured to: forward upstream data sent by any one of the upstream ports to a corresponding downstream port according to a destination address of the upstream data; and / or forward downstream data sent by any one of the downstream ports to a corresponding upstream port according to a destination address of the downstream data.
[0067] In an embodiment, the at least one switch chip includes a routing and forwarding engine module, and the routing and forwarding engine module is configured to: determine, according to a routing table, a destination address of upstream data sent by any one of the upstream ports in the current switch chip; and / or a destination address of downstream data sent by any one of the downstream ports in the current switch chip.
[0068] In an embodiment, the at least one switch chip includes a storage module, and the storage module is configured to: store the routing table, upstream data sent by any one of the upstream ports in the current switch chip, and downstream data sent by any one of the downstream ports in the current switch chip.
[0069] In an embodiment, the at least one switch chip includes a direct memory access engine module, and the direct memory access engine module is configured to: implement communication between other resource pools connected to any one of the upstream ports in the current switch chip and hardware resource devices connected to any one of the downstream ports in the current switch chip by using a direct memory access (DMA) technology.
[0070] In an example, the at least one switch chip comprises: a control logic module; the control logic module comprises: a data packet conversion unit, a register unit and a resource awareness module.
[0071] The data packet conversion unit is configured to perform data format conversion on upstream data sent by any one of the upstream ports in the current switch chip and / or downstream data sent by any one of the downstream ports in the current switch chip.
[0072] The register unit is configured to configure port information of any upstream port and / or any downstream port in the current switch chip; the port information comprises: port parameters, port attributes and port states.
[0073] The resource awareness module is configured to monitor whether each downstream port in the current switch chip is connected with a hardware resource device and / or monitor resource usage information of the hardware resource device connected with the downstream port in the current switch chip, and generate a corresponding resource awareness table.
[0074] Correspondingly, each downstream port of the at least one switch chip supports a first protocol, and each upstream port of the at least one switch chip supports a second protocol; correspondingly, the data packet conversion unit is configured to convert upstream data sent by any one of the upstream ports in the current switch chip from the second protocol format to the first protocol format; and / or convert upstream data sent by any one of the downstream ports in the current switch chip from the first protocol format to the second protocol format. For example, each downstream port of the at least one switch chip supports a CCIX protocol (Cache Coherent Interconnect for Accelerators), and each upstream port of the at least one switch chip supports a CXL protocol (Compute Express Link); correspondingly, the data packet conversion unit is configured to convert upstream data sent by any one of the upstream ports in the current switch chip from the CXL protocol format to the CCIX protocol format; and / or convert upstream data sent by any one of the downstream ports in the current switch chip from the CCIX protocol format to the CXL protocol format.
[0075] Correspondingly, the register unit is configured to configure a routing table and a resource awareness table; the resource awareness table records parameters such as remaining memory, maximum bandwidth and remaining computing power of the hardware resource device.
[0076] Correspondingly, the resource awareness module is configured to update the resource awareness table in real time according to whether each downstream port is connected with a hardware resource device and / or resource usage information of the hardware resource device obtained by real-time monitoring.
[0077] In an embodiment, any hardware resource device is configured to send information of free resources in itself to the resource awareness unit at a preset period; accordingly, the resource awareness unit updates the resource awareness table according to the information sent by the hardware resource device.
[0078] Moreover, the resource awareness unit is configured to periodically detect whether the communication link between each upstream port of the current switch chip and the connected resource pool is connected; and periodically detect whether the communication link between each downstream port of the current switch chip and the connected hardware resource device is connected.
[0079] In an embodiment, each upstream port of at least one switch chip is provided with a decoder; the decoder is configured to determine the address mapping relationship between the memory of the computing device connected to the upstream port where the decoder is located and the hardware resource device connected to the downstream port of the current switch chip. In an example, the decoder is configured to store the memory capacity information of the computing device connected to the upstream port where the decoder is located.
[0080] In an embodiment, at least one switch chip comprises a cache coherence interface module; the cache coherence interface module is configured to provide a cache coherence transmission channel for the hardware resource devices connected to each downstream port of the current switch chip.
[0081] In an embodiment, at least one switch chip comprises a clock and reset module; the clock and reset module is configured to keep the clock synchronization between each device in the system.
[0082] In an embodiment, at least one switch chip further comprises a switching function module. The at least one switch chip comprises a power management module; the power management module is configured to provide power for the switching function module and the switch chips in each resource pool, and manage the power consumption of the switching function module and the switch chips in each resource pool.
[0083] It can be seen that the embodiment integrates different hardware resource devices such as memory devices, hardware acceleration devices or computing devices into different resource pools, at least one upstream port of at least one switching chip in each resource pool is connected to an upstream port of at least one switching chip in another resource pool, interconnection between different resource pools is realized, at least one downstream port of at least one switching chip in each resource pool is connected to a hardware resource device, so that the switching chip forwards upstream data to a corresponding downstream port according to the destination address of the upstream data sent by any one upstream port, and / or forwards downstream data to a corresponding upstream port according to the destination address of the downstream data sent by any one downstream port, realizing intrasystem communication. The scheme realizes interconnection of intrasystem devices with as few port numbers as possible, and downstream devices do not need to be connected two by two, reducing the number of connection lines and the complexity of wiring layout, and the communication delay can be correspondingly reduced, meeting the communication requirements of low-delay services.
[0084] It should be noted that the switching chip in the present application is a rack bus level CXL switch chip (hereinafter referred to as switch), which has cache coherence maintenance function of CXL in addition to switching and forwarding function. In addition, the connection between any two CXL switch chips is a CXL bus.
[0085] Referring to FIG. 2, the embodiment adopts a direct connection mode with reference to the Dragonfly network (a kind of network topology) topology, and divides the whole system into three layers: a system layer, a group layer (a resource pool layer) and a switch layer. The system layer includes four resource pools, and full connection is realized between the four resource pools, which can improve network interconnectivity and redundancy, and ensure efficient and reliable transmission of data in the whole system. In addition, the whole system supports resource scalability, is suitable for construction of a pooled system, the switch forwarding hop number is at most 3, and the transmission path and the number of nodes are superior to CLOS topology, having superior delay performance. Compared with a device full interconnection topology or a switch full interconnection topology, the connection lines in the example of FIG. 3 are greatly reduced, and the wiring difficulty, deployment difficulty and cost are significantly reduced.
[0086] FIG. 2 shows a heterogeneous acceleration resource pool of GPUs and FPGAs, a CXL extended memory resource pool based on DRAM and NVMe SSD (Non-Volatile Memory Express Solid State Drive).
[0087] The system shown in FIG. 2 realizes full interconnection of pooled resources with limited links, greatly reducing the consumption of multiple devices, improving the energy efficiency ratio, and reducing transmission delay. Second, deploying different types of resource pools facilitates resource management and scheduling, and can reasonably layout application computing power according to the characteristics of different resources to improve acceleration efficiency and improve resource utilization.
[0088] Optionally, the embodiment configures 3 CXL 3.1 switch chips for each resource pool, supporting switch-to-switch connection. 1 upstream port in each switch chip is set as interconnection of different resource pools, a total of 6 PCIe links, bidirectional bandwidth: PCIE 4.0 192GB / s (Gigabytes per second) and PCIE 5.0 384GB / s. Within the resource pool, 2 upstream ports of each switch are set as interconnection of resource devices connected on different switch chips in the same resource pool, and 3 PCIe links can realize full connection of devices within the resource pool. Each CXL switch in FIG. 2 has 10 interfaces, and in fact there can be more interfaces.
[0089] The embodiment uses the number X-Y-Z to represent: the Xth resource pool, the Zth port on the Yth switch. Corresponding to FIG. 2, X takes values 1-4, Y takes values 1-3, and Z takes values 1-10. 1st port of each switch is set as interconnection between resource pools, and the port interconnection table is shown in Table 1 to realize two-by-two interconnection between four pools. 3rd and 4th ports of each switch are set as two-by-two interconnection between the three switches within the resource pool.
[0090] Table 1 Interconnection Table Between Resource Pools
[0091] X represents the link within the resource pool, and the interconnection within the resource pool is shown in Table 2.
[0092] Table 2 Interconnection Table Within Resource Pool
[0093] In this example, 12 CXL switch chips are fully connected, which requires 6+3x4=18 connection lines in total, and each switch chip occupies 3 interfaces. If a full interconnection topology is used for the 12 CXL switch chips, in order to directly interconnect the 12 switch chips two by two, 12x11 / 2=66 connection lines are required, and the interface of each CXL switch is occupied 11. According to this embodiment, the number of occupied interfaces is relatively small, which also means that in the case of limited expansion capability of the switch, a larger scale system interconnection can be realized. Thus, on the basis of interconnection between all devices, the use of connection lines and the occupation of switch port resources are greatly reduced, and the wiring difficulty is reduced.
[0094] In addition, the number of hardware devices connected by the switch, the number of switches in each resource pool, and the number of resource pools can be expanded according to the actual computing demand, and the expansion cost of the three levels increases in turn. Within the memory resource pool, by using the memory expansion function provided by the CXL protocol, devices such as DRAM, NVMe SSD and other storage media can be expanded into system memory through the CXL bus, and the access delay is much lower than the delay required by the RDMA (Remote Direct Memory Access) technology through network access in the traditional separate memory pool.
[0095] On the switch in the FPGA and GPU accelerated computing resource pool, a corresponding interface is left to connect the NIC network card, which can realize connection with Ethernet for cross-rack range communication, and can also support expansion of the system scale.
[0096] In addition, the resource pool inside also supports a variety of hardware device resource topologies, which can be compatible with CLOS topology and DPU (Data Processing Unit) full interconnection topology scheme, providing more possibilities for system performance improvement and innovative applications.
[0097] If a link between two resource pools cannot meet the bandwidth requirement, 2 and 3 ports on each switch can be interconnected to obtain greater parallel bandwidth. Ensure that each switch has a link directly to another three resource pools, as shown in FIG. 3. At this time, the maximum hop count of the system is reduced to 2, and the bandwidth is increased while the average latency of the system is reduced. Correspondingly, the port connection is as shown in Table 3, and a total of 18 links are required.
[0098] Table 3: Interconnection between resource pools in bandwidth expansion mode
[0099] In an example, the switch can be implemented based on a reconfigurable FPGA, and is configured to support CXL protocol and PCIe protocol. As shown in FIG. 4, the switch includes the following functional modules:
[0100] A routing and forwarding engine module is responsible for implementing routing and forwarding of data packets, controlling distribution of traffic, and arbitrating management of bus and interface resources, to ensure efficiency and stability of data transmission. A ternary content addressable memory (TCAM) is configured to store and quickly search table entries such as routing tables (function: fast routing search, store the port with communication connection), has high search efficiency, and can quickly match the destination address of the data packet to determine the next hop path.
[0101] A host-managed device memory decoder (HDM) is present in each upstream port of the switch, and stores volatile and persistent memory capacity of extended memory, and determines the mapping relationship between host physical addresses and device physical addresses. The HDM decoder in the upstream port determines which downstream port is the target of memory access. The address mapping relationship between the host address and the downstream device address is analyzed,
[0102] An on-chip storage module can use an on-chip SRAM (Static Random Access Memory) memory, which has fast read and write speed and low access delay, and can be used as a storage space for caching data packets, forwarding tables, and transit buffers.
[0103] A direct memory access (DMA) engine module is configured to support DMA high-speed data transmission function, and reduces the load of CPU by directly accessing memory.
[0104] A control logic module is responsible for controlling, managing and configuring the resources and functions of the switch, and can optionally include:
[0105] 1) A data packet conversion unit is responsible for converting the format of received data packets, such as converting data packets from CXL protocol format to CCIX protocol format, or editing and modifying the header information of data packets, and implementing cache coherence maintenance functions. CCIX protocol is used between intra-group acceleration computing devices.
[0106] 2) A register unit is responsible for implementing the configuration space of the switch device, and is configured to configure parameters, attributes and states of the port, configure routing table and resource awareness table parameters, and other functions.
[0107] 3) Resource awareness table, configured to monitor system resource usage. The resource utilization table can be stored in the SDRAM (Synchronous Dynamic Random Access Memory) cache on the module, which can quickly read and update data to ensure that the resource utilization can be quickly obtained when needed. For example: the memory / computing device sends the idle memory capacity / idle computing unit number to the control logic module at a certain period, and the control logic module configures the resource awareness table after receiving the port resource information, so as to facilitate dynamic awareness of resource utilization. The control logic module can also send signals to the port module at a preset period to detect whether the link between the port and the external device is normal, and control whether the physical link of the port is connected, to realize dynamic transformation of system connection topology and monitoring and isolation of faults.
[0108] Cache consistency interface module: can be connected to the PCIe interface of the mainboard, providing a data transmission channel between the accelerator and other devices. Including protocol transmission layer, link layer, physical layer, the physical layer is responsible for physical link initialization and control related functions, the link layer is responsible for data link state control and management, and the transmission layer is responsible for message encapsulation and decapsulation.
[0109] Other support modules are also needed on the switch device, such as clock and reset modules to provide clock synchronization and reset functions between devices to ensure the accuracy of data transmission; the power management module is responsible for supplying power to each module of the switch and managing the power consumption of the switch device.
[0110] It should be noted that the resource awareness module and the idle port in the switch support dynamic expansion of resources in the pool and efficient management of pooled resources; at the same time, the same type of resources in the pool are redundant devices, and even if one or several devices fail, their tasks can be taken over by the same type of devices through internal scheduling of the system, improving the reliability of the pooled platform.
[0111] In summary, the embodiment proposes a hierarchical pooling server system interconnection topology scheme based on the CXL 3.1 protocol specification. The first layer is a pooling server. The second layer is a plurality of resource pools, which are interconnected with each other. The third layer is a single switch layer, which connects a plurality of heterogeneous acceleration and storage devices. The devices support a plurality of interconnection modes, such as interconnection between devices through CCIX and the like. The interconnection topology not only reduces the system connection complexity, but also enables interconnection and intercommunication between devices in all resource pools with a small hop count. Through the interconnection architecture, low-latency coherent access of CPU and accelerator heterogeneous computing resources to host memory and extended memory resources in the entire server system can be achieved, and large-scale expansion and resource pooling management and maintenance of the pooling resources are supported. The scheme also supports bandwidth expansion mode, realizes more link and higher bandwidth interconnection scheme between resource pools, and has good system flexibility. The link can be flexibly configured and adjusted according to the requirements. Each switch has a link directly connected to other resource pools, so that the interconnection between resource pools has a redundant path, improving the reliability and fault tolerance of the server system. In addition, a resource perception unit and related control logic are designed in the switch device, supporting dynamic perception of the utilization of the pooling resources. The control logic module supports dynamic configuration of the port link, can realize dynamic adjustment of the system topology, and increases the flexibility of the system topology while supporting fault isolation.
[0112] The pooling server interconnection topology scheme provided by the embodiment can realize high-bandwidth and low-latency interconnection in a high-performance server system, and can be connected to a network through a network card and the like, supporting larger-range and higher-level interconnection. The topology interconnection scheme not only simplifies system deployment, but also reduces cost and power consumption, and is suitable for application in a data center.
[0113] A device interconnection method provided by an embodiment of the present application is described below. The device interconnection method described below can be mutually referred to with other embodiments described herein.
[0114] The embodiment of the present application discloses a device interconnection method applied to the device interconnection system of any of the foregoing embodiments. The method comprises: for any resource pool in the device interconnection system, connecting at least one upstream port of at least one switching chip in the current resource pool to an upstream port of at least one switching chip in another resource pool; connecting at least one downstream port of at least one switching chip in the current resource pool to a hardware resource device; and connecting at least one upstream port of at least one switching chip in the current resource pool to an upstream port of another switching chip in the current resource pool.
[0115] The device interconnection system comprises at least two resource pools, and each resource pool comprises at least one switching chip; the hardware resource devices in the same resource pool are memory devices, hardware acceleration devices or computing devices, and the hardware resource devices in different resource pools are different;
[0116] The at least one switching chip is configured to forward upstream data sent by any one of the upstream ports to a corresponding downstream port according to a destination address of the upstream data, and / or forward downstream data sent by any one of the downstream ports to a corresponding upstream port according to a destination address of the downstream data.
[0117] In an embodiment, the at least one switching chip comprises a routing and forwarding engine module; the routing and forwarding engine module is configured to determine a destination address of upstream data sent by any one of the upstream ports in the current switching chip and / or a destination address of downstream data sent by any one of the downstream ports in the current switching chip according to a routing table.
[0118] In an embodiment, the at least one switching chip comprises a storage module; the storage module is configured to store the routing table, upstream data sent by any one of the upstream ports in the current switching chip and downstream data sent by any one of the downstream ports in the current switching chip.
[0119] In an embodiment, the at least one switching chip comprises a direct memory access engine module; the direct memory access engine module is configured to realize communication between other resource pools connected to any one of the upstream ports in the current switching chip and hardware resource devices connected to any one of the downstream ports in the current switching chip by using a direct memory access technology.
[0120] In an embodiment, the at least one switching chip comprises a control logic module; the control logic module comprises a data packet conversion unit, a register unit and a resource awareness module.
[0121] The data packet conversion unit is configured to perform data format conversion on upstream data sent by any one of the upstream ports in the current switching chip and / or downstream data sent by any one of the downstream ports in the current switching chip.
[0122] The register unit is configured to configure port information of any one of the upstream ports and / or any one of the downstream ports in the current switching chip; the port information comprises port parameters, port attributes and port states.
[0123] The resource awareness module is configured to monitor whether each downstream port in the current switching chip is connected to a hardware resource device and / or monitor resource usage information of the hardware resource device connected to the downstream port in the current switching chip, and generate a corresponding resource awareness table.
[0124] In an embodiment, each downstream port of the at least one switch chip supports the CCIX protocol, and each upstream port of the at least one switch chip supports the CXL protocol; accordingly, the data packet conversion unit is configured to convert upstream data sent by any one of the upstream ports of the current switch chip from the CXL protocol format to the CCIX protocol format; and / or convert upstream data sent by any one of the downstream ports of the current switch chip from the CCIX protocol format to the CXL protocol format.
[0125] In an embodiment, the register unit is configured to configure the routing table and the resource awareness table.
[0126] In an embodiment, the resource awareness unit is configured to update the resource awareness table in real time according to whether each downstream port is connected to a hardware resource device and / or resource usage information of the hardware resource device obtained through real-time monitoring.
[0127] In an embodiment, any hardware resource device is configured to send information about idle resources in itself to the resource awareness unit at a preset period; accordingly, the resource awareness unit updates the resource awareness table according to the information sent by the hardware resource device.
[0128] In an embodiment, the resource awareness unit is configured to periodically detect whether the communication link between each upstream port of the current switch chip and the connected resource pool is connected, and periodically detect whether the communication link between each downstream port of the current switch chip and the connected hardware resource device is connected.
[0129] In an embodiment, each upstream port of the at least one switch chip is provided with a decoder; the decoder is configured to determine the address mapping relationship between the memory of the computing device connected to the upstream port where the current decoder is located and the hardware resource device connected to the downstream port of the current switch chip.
[0130] In an embodiment, the decoder is configured to store the memory capacity information of the computing device connected to the upstream port where the current decoder is located.
[0131] In an embodiment, the at least one switch chip comprises a cache coherence interface module; the cache coherence interface module is configured to provide a cache coherence transmission channel between the hardware resource devices connected to each downstream port of the current switch chip.
[0132] In an embodiment, the at least one switch chip comprises a clock and reset module; the clock and reset module is configured to maintain clock synchronization between devices in the system.
[0133] In an embodiment, the at least one switch chip further comprises a switching function module.
[0134] In an implementation, the at least one switching chip comprises a power management module, and the power management module is configured to provide power supply for the switching chip and the switching machine in each resource pool and manage power consumption of the switching chip and the switching machine in each resource pool.
[0135] In an implementation, each hardware resource device in each resource pool supports a CLOS topology and / or a full interconnection topology.
[0136] In the embodiments, the optional working processes of the modules and units can refer to the corresponding content disclosed in the foregoing embodiments, and will not be described in detail herein.
[0137] It can be seen that the method for device interconnection provided in the embodiments realizes interconnection of devices in a system with as few ports as possible, and downstream devices do not need to be connected two by two, thereby reducing the number of connection lines and the complexity of wiring layout, and communication delay can be reduced accordingly, thereby meeting the communication requirements of low-delay requirement services.
[0138] An electronic device provided by the embodiments of the present application is described below, and the electronic device described below can be referred to the other embodiments described herein. The electronic device can be any functional module of the foregoing system.
[0139] Referring to FIG. 5, the embodiments of the present application disclose an electronic device, comprising:
[0140] The memory 501 is configured to save a computer program.
[0141] The processor 502 is configured to execute the computer program to implement the method disclosed in any of the foregoing embodiments.
[0142] In the embodiments, when the processor executes the computer program saved in the memory, the following steps can be implemented: forwarding upstream data sent by any one upstream port to a corresponding downstream port according to a destination address of the upstream data; and / or forwarding downstream data sent by any one downstream port to a corresponding upstream port according to a destination address of the downstream data.
[0143] In the embodiments, when the processor executes the computer program saved in the memory, the following steps can be implemented: determining a destination address of upstream data sent by any one upstream port in the current switching chip according to the routing table; and / or determining a destination address of downstream data sent by any one downstream port in the current switching chip.
[0144] In the embodiments, when the processor executes the computer program saved in the memory, the following steps can be implemented: storing the routing table, upstream data sent by any one upstream port in the current switching chip, and downstream data sent by any one downstream port in the current switching chip.
[0145] In the embodiment, when the processor executes the computer program stored in the memory, the following steps can be implemented: in a direct memory access technology, communication between other resource pools connected by any one of the upstream ports of the current switch chip and the hardware resource devices connected by any one of the downstream ports of the current switch chip is implemented.
[0146] In the embodiment, when the processor executes the computer program stored in the memory, the following steps can be implemented: data format conversion is performed on the upstream data sent by any one of the upstream ports of the current switch chip and / or the downstream data sent by any one of the downstream ports of the current switch chip.
[0147] In the embodiment, when the processor executes the computer program stored in the memory, the following steps can be implemented: port information of any upstream port and / or any downstream port of the current switch chip is configured; the port information includes: port parameters, port attributes and port states.
[0148] In the embodiment, when the processor executes the computer program stored in the memory, the following steps can be implemented: whether each downstream port of the current switch chip is connected with a hardware resource device and / or resource usage information of the hardware resource device connected by the downstream port of the current switch chip is monitored, and a corresponding resource awareness table is generated.
[0149] In the embodiment, when the processor executes the computer program stored in the memory, the following steps can be implemented: upstream data sent by any one of the upstream ports of the current switch chip is converted from a CXL protocol format to a CCIX protocol format; and / or upstream data sent by any one of the downstream ports of the current switch chip is converted from a CCIX protocol format to a CXL protocol format.
[0150] In the embodiment, when the processor executes the computer program stored in the memory, the following steps can be implemented: a routing table and a resource awareness table are configured.
[0151] In the embodiment, when the processor executes the computer program stored in the memory, the following steps can be implemented: according to whether each downstream port is connected with a hardware resource device and / or resource usage information of the hardware resource device obtained through real-time monitoring, the resource awareness table is updated in real time.
[0152] In the embodiment, when the processor executes the computer program stored in the memory, the following steps can be implemented: idle resource information in the self is sent to the resource awareness unit according to a preset period; accordingly, the following steps can also be implemented: the resource awareness table is updated according to the information sent by the hardware resource device.
[0153] In the embodiment, when the processor executes the computer program stored in the memory, the following steps can be implemented: periodically detecting whether the communication link between each upstream port of the current switching chip and the connected resource pool is connected; periodically detecting whether the communication link between each downstream port of the current switching chip and the connected hardware resource device is connected.
[0154] In the embodiment, when the processor executes the computer program stored in the memory, the following steps can be implemented: determining the address mapping relationship between the memory of the computing device connected to the upstream port where the current decoder is located and the hardware resource device connected to the downstream port in the current switching chip.
[0155] In the embodiment, when the processor executes the computer program stored in the memory, the following steps can be implemented: storing the memory capacity information of the computing device connected to the upstream port where the current decoder is located.
[0156] In the embodiment, when the processor executes the computer program stored in the memory, the following steps can be implemented: providing a cache consistency transmission channel between the hardware resource devices connected to each downstream port in the current switching chip.
[0157] In the embodiment, when the processor executes the computer program stored in the memory, the following steps can be implemented: maintaining the clock synchronization between devices in the system.
[0158] In the embodiment, when the processor executes the computer program stored in the memory, the following steps can be implemented: providing power supply for the switch and the switching chip in each resource pool, and managing the power consumption of the switch and the switching chip in each resource pool.
[0159] Optionally, the embodiment of the present application further provides an electronic device. The electronic device can be a server as shown in FIG. 6, or a terminal as shown in FIG. 7. FIG. 6 and FIG. 7 are structural diagrams of electronic devices according to an exemplary embodiment, and the contents in the figures should not be considered as any limitation on the use range of the present application.
[0160] FIG. 6 is a structural schematic diagram of a server provided by the embodiment of the present application. The server can include at least one processor, at least one memory, a power supply, a communication interface, an input / output interface and a communication bus. The memory is configured to store a computer program, which is loaded and executed by the processor to implement the related steps in the device interconnection disclosed in any of the preceding embodiments.
[0161] In this embodiment, the power supply is configured to provide working voltage for each hardware device on the server; the communication interface is capable of creating data transmission channel between the server and external devices, and the communication protocol followed by the communication interface is any communication protocol applicable to the technical solution of the present application, which is not limited herein; the input / output interface is configured to obtain external input data or output data to the outside world, and the optional interface type can be selected according to the application needs, which is not limited herein.
[0162] In addition, the memory as a carrier for storing resources can be a read-only memory, a random access memory, a magnetic disk or an optical disk, etc., and the resources stored thereon include an operating system, a computer program and data, etc., and the storage mode can be temporary storage or permanent storage.
[0163] The operating system is configured to manage and control each hardware device and computer program on the server, so as to realize the operation and processing of the processor on the data in the memory, and the operating system can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program capable of completing the device interconnection method disclosed in any of the preceding embodiments, the computer program can also include a computer program capable of completing other specific work. In addition to the data including the update information of the application program and the like, the data can also include the developer information of the application program and the like.
[0164] FIG. 7 is a structural schematic diagram of a terminal provided by an embodiment of the present application, which can include but is not limited to a smart phone, a tablet computer, a notebook computer or a desktop computer, etc.
[0165] Generally, the terminal in this embodiment includes a processor and a memory.
[0166] The processor can include one or more processing cores, such as a 4-core processor, an 8-core processor, and the like. The processor can be implemented in at least one of a hardware form of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), a PLA (Programmable Logic Array). The processor can also include a main processor and a coprocessor. The main processor is a processor configured to process data in an awake state, also referred to as a CPU (Central Processing Unit). The coprocessor is a low-power processor configured to process data in a standby state. In some embodiments, the processor can be integrated with a GPU (Graphics Processing Unit). The GPU is configured to be responsible for rendering and drawing of content required to be displayed by the display screen. In some embodiments, the processor can further include an AI (Artificial Intelligence) processor configured to process computing operations related to machine learning.
[0167] The memory can include one or more computer non-volatile readable storage media, which can be non-transitory. The memory can also include a high-speed random access memory and a non-volatile memory such as one or more disk storage devices, flash memory devices. In this embodiment, the memory is at least configured to store the following computer programs, which, after being loaded and executed by the processor, can implement the related steps of the device interconnection method performed by the terminal side disclosed in any of the preceding embodiments. In addition, the resources stored by the memory can also include an operating system and data, and the storage mode can be temporary storage or permanent storage. The operating system can include Windows, Unix, Linux, and the like. The data can include, but is not limited to, update information of application programs.
[0168] In some embodiments, the terminal can further include a display screen, an input / output interface, a communication interface, a sensor, a power supply, and a communication bus.
[0169] Those skilled in the art can understand that the structure shown in FIG. 7 does not constitute a limitation on the terminal, and can include more or fewer components than those shown.
[0170] A non-volatile readable storage medium provided by an embodiment of the present application is described below. The non-volatile readable storage medium described below can be mutually referred to with other embodiments described herein.
[0171] A non-volatile readable storage medium for storing a computer program, wherein the computer program is executed by a processor to implement the device interconnection method disclosed in the foregoing embodiments. The non-volatile readable storage medium is a computer-readable non-volatile storage medium, which is a carrier for storing resources, and can be a read-only memory, a random access memory, a magnetic disk, an optical disk, or the like. The resources stored on the non-volatile readable storage medium include an operating system, a computer program, data, and the like, and the storage manner can be temporary storage or permanent storage.
[0172] A computer program product is introduced below, and the computer program product described below can be mutually referred to with other embodiments described herein.
[0173] A computer program product includes a computer program / instruction, which is executed by a processor to implement the steps of the device interconnection method disclosed above.
[0174] The embodiments in the specification are described in a progressive manner, and each embodiment focuses on the difference from other embodiments. The same or similar parts of each embodiment can be referred to each other.
[0175] The steps of the method or algorithm described in combination with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a compact disc read-only memory (CD-ROM), or any other form of non-volatile storage medium known in the art.
[0176] Optional examples are applied herein to illustrate the principles and implementation manners of the present application. The above description of the embodiments is only used to help understand the method and core idea of the present application; meanwhile, for those skilled in the art, according to the idea of the present application, the optional implementation manners and application ranges can be changed. In summary, the content of the specification should not be understood as a limitation of the present application.
Claims
A device interworking system characterized by, The application relates to a method for realizing a multi-resource-pool system, comprising the following steps: At least two resource pools are provided, each resource pool comprising at least one switching chip; At least one downstream port of the at least one switching chip is arranged to be connected with a hardware resource device; the hardware resource devices in the same resource pool are memory devices, hardware acceleration devices or computing devices, and the hardware resource devices in different resource pools are different; At least one upstream port of the at least one switching chip is arranged to be connected with an upstream port of at least one switching chip in another resource pool; at least one upstream port of the at least one switching chip is arranged to be connected with an upstream port of another switching chip in the resource pool where the current switching chip is located; The at least one switching chip is arranged to forward upstream data sent by any one upstream port to a corresponding downstream port according to the destination address of the upstream data; and / or forward downstream data sent by any one downstream port to a corresponding upstream port according to the destination address of the downstream data. The at least one switching chip comprises a routing and forwarding engine module; The device interworking system according to claim 1, characterized in that, The routing and forwarding engine module is arranged to determine the destination address of upstream data sent by any one upstream port in the current switching chip and / or the destination address of downstream data sent by any one downstream port in the current switching chip according to a routing table. The at least one switching chip comprises a storage module; The device interworking system according to claim 1, characterized in that, The storage module is arranged to store the routing table, upstream data sent by any one upstream port in the current switching chip and downstream data sent by any one downstream port in the current switching chip. The at least one switching chip comprises a direct memory access engine module; The device interworking system according to claim 1, characterized in that, The direct memory access engine module is arranged to realize communication between other resource pools connected with any one upstream port in the current switching chip and hardware resource devices connected with any one downstream port in the current switching chip by using a direct memory access technology. The at least one switching chip comprises a control logic module; The device interworking system according to claim 1, characterized in that, The control logic module comprises a data packet conversion unit, a register unit and a resource awareness module; The data packet conversion unit is arranged to perform data format conversion on upstream data sent by any one upstream port in the current switching chip and / or downstream data sent by any one downstream port in the current switching chip; The register unit is arranged to configure port information of any one upstream port and / or any one downstream port in the current switching chip; the port information comprises port parameters, port attributes and port states; The resource awareness module is arranged to monitor whether each downstream port in the current switching chip is connected with a hardware resource device and / or monitor resource usage information of the hardware resource device connected with the downstream port in the current switching chip, and generate a corresponding resource awareness table. Each downstream port of the at least one switching chip supports a first protocol, and each upstream port of the at least one switching chip supports a second protocol. The device interworking system according to claim 5, characterized in that, Correspondingly, the data packet conversion unit is configured to convert upstream data sent by any one of the upstream ports of the current switch chip from the second protocol format to the first protocol format; and / or convert upstream data sent by any one of the downstream ports of the current switch chip from the first protocol format to the second protocol format. The device interworking system according to claim 5, characterized in that, The register unit is configured to configure a routing table and a resource awareness table. The device interworking system according to claim 5, characterized in that, The resource awareness unit is configured to update the resource awareness table in real time according to whether each downstream port is connected with a hardware resource device and / or resource usage information of the hardware resource device obtained through real-time monitoring. The device interworking system according to claim 5, characterized in that, Any hardware resource device is configured to send idle resource information in the hardware resource device to the resource awareness unit at a preset period; Correspondingly, the resource awareness unit updates the resource awareness table according to the information sent by the hardware resource device. The device interworking system according to claim 5, characterized in that, The resource awareness unit is configured to periodically detect whether a communication link between each upstream port of the current switch chip and a connected resource pool is connected; and periodically detect whether a communication link between each downstream port of the current switch chip and a connected hardware resource device is connected. The device interworking system according to claim 1, characterized in that, Each upstream port of the at least one switch chip is provided with a decoder. The decoder is configured to determine an address mapping relationship between a memory of a computing device connected to the upstream port where the current decoder is located and a hardware resource device connected to a downstream port of the current switch chip. The device interworking system according to claim 11, characterized in that, The decoder is configured to store memory capacity information of a computing device connected to the upstream port where the current decoder is located. The device interworking system according to claim 1, characterized in that, The at least one switch chip comprises a cache coherency interface module. The cache coherency interface module is configured to provide a cache coherency transmission channel between hardware resource devices connected to each downstream port of the current switch chip. The device interworking system according to claim 1, characterized in that, The at least one switch chip comprises a clock and reset module. The clock and reset module is configured to maintain clock synchronization between devices in the system. The device interworking system according to any one of claims 1 to 14, characterized in that, The device interconnection system further comprises a switching function module. The device interworking system according to claim 15, characterized in that, The at least one switch chip comprises a power management module. The power management module is configured to provide power supply for the switching function module and the switch chips in each resource pool, and manage power consumption of the switching function module and the switch chips in each resource pool. The device interworking system according to any one of claims 1 to 14, characterized in that, Each hardware resource device in each resource pool supports a CLOS topology and / or a full interconnection topology. A device interconnection method, characterized by, The device interconnection system is applied to any one of claims 1 to 17, comprising: For any resource pool in the device interconnection system, at least one upstream port of at least one switch chip in the current resource pool is connected to an upstream port of at least one switch chip in another resource pool; at least one downstream port of at least one switch chip in the current resource pool is connected to a hardware resource device; and at least one upstream port of at least one switch chip in the current resource pool is connected to an upstream port of another switch chip in the current resource pool; The device interconnection system comprises at least two resource pools, each resource pool comprising at least one switch chip; hardware resource devices in the same resource pool are memory devices, hardware acceleration devices or computing devices, and hardware resource devices in different resource pools are different; The at least one switching chip is configured to: forward upstream data sent by any one of the upstream ports to a corresponding downstream port according to a destination address of the upstream data; and / or forward downstream data sent by any one of the downstream ports to a corresponding upstream port according to a destination address of the downstream data. An electronic device, characterized by comprising: Comprising: a memory configured to store a computer program; a processor configured to execute the computer program to implement the method of claim 18. A non-volatile readable storage medium characterized by A computer program for saving, wherein the computer program is executed by a processor to implement the method of claim 18. A computer program product comprising computer programs / instructions, characterized in that, The computer program / instructions are executed by a processor to implement the method of claim 18.
Citation Information
Patent Citations
A hardware reconfiguration system and method
CN109240832A
Multi-accelerator card heterogeneous server and resource link reconstruction method
CN117687956A
Interconnection pooling system topology management device and method based on PCIe Switch
CN117834447A
Virtual resource control and distribution
US20190138363A1