Computing system, control method and device, equipment, medium and switch

By using the first switch to interconnect the southward interface of the acceleration card in a multi-machine and multi-card cluster, and using the controller to configure the southward address and routing information, the problem of the long communication path of the acceleration card is solved, and communication efficiency and system scalability are improved.

CN120378429APending Publication Date: 2025-07-25LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510558647.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, the accelerator card communication path and low communication efficiency of a multi-machine and multi-card cluster are problems, especially in the accelerator card interconnection across host domains, which requires forwarding through a central processor, resulting in inefficiency.

Method used

The first switch is used to realize the southward interface interconnection of the accelerator card, and the southward address is configured for the accelerator card through the first controller, and the routing information is configured according to the connection relationship between the accelerator card and the bus bridge, so that the bus bridge forwards data, realizing direct communication between the accelerator card, avoiding northward interface forwarding and path specified by the upper-level software.

Benefits of technology

It improves the communication efficiency of accelerated card interconnection in multi-machine and multi-card clusters, improves the scalability of the computing system, and provides easy to popularize and cost-effective solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378429A_ABST
    Figure CN120378429A_ABST
Patent Text Reader

Abstract

The invention discloses a computing system, a control method, a control device, equipment, a medium and a switch, and relates to the technical field of computers, in a multi-machine multi-card cluster comprising a plurality of computing equipment, a first switch is adopted to realize interconnection between southbound interfaces of different acceleration cards, and the first switch comprises a first controller and a bus bridge, the bus bridge is connected with a southbound interface of the acceleration card, the first controller is used for distributing a southbound address for the acceleration card and configuring routing information of the bus bridge, so that the acceleration card carries out data forwarding based on the southbound address, and the bus bridge carries out data forwarding among southbound interfaces of different acceleration cards according to the routing information. According to the method and the device, any cross-host accelerator card does not need to be forwarded through a northbound interface during communication, and does not need to be appointed by upper software, so that the problems of relatively long communication path and communication blockage in related technologies are solved, and the communication efficiency of interconnection of the accelerator cards in a multi-host multi-card cluster is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a computing system, a control method, a device, a device, a medium and a switch. Background Art

[0002] With the development of artificial intelligence technology, the demand for computing power has increased significantly. It is not only necessary to deploy acceleration cards on a single server to improve the computing power of a single machine, but also necessary to interconnect a large number of servers to form a multi-machine multi-card cluster. For the interconnection of acceleration cards across hosts when deploying a multi-machine multi-card cluster, there are problems of long communication paths and low communication efficiency in related technologies.

[0003] How to improve the communication efficiency of the southbound interconnection of acceleration cards in a multi-machine multi-card cluster is a technical problem that needs to be solved by those skilled in the art. Summary of the Invention

[0004] The present invention provides a computing system, a control method, a device, a device, a medium and a switch to at least solve the problem of low communication efficiency in the interconnection of acceleration cards across hosts in related technologies.

[0005] The present invention provides a computing system, including: a plurality of computing devices and a first switch; The first switch includes a first controller and a bus bridge; The computing device is provided with an interface for connecting an acceleration card, and the southbound interfaces of different acceleration cards located on the same computing device and the southbound interfaces of different acceleration cards located on different computing devices are interconnected through the bus bridge; The first controller is used to configure a southbound address for the acceleration card, and configure routing information of the bus bridge according to the connection relationship between the acceleration card and the bus bridge and the southbound address of the acceleration card; The bus bridge is used to forward data between the southbound interfaces of different acceleration cards according to the routing information.

[0006] The present invention provides a switch, including: a first controller and a bus bridge; The southbound interfaces of different acceleration cards located on the same computing device and the southbound interfaces of different acceleration cards located on different computing devices are interconnected through the bus bridge; The first controller is used to configure a southbound address for the acceleration card, and configure routing information of the bus bridge according to the connection relationship between the acceleration card and the bus bridge and the southbound address of the acceleration card; The bus bridge is used to forward data between the southbound interfaces of different acceleration cards according to the routing information.

[0007] The present invention provides a control method for a computing system, including: Configuring a southbound address for an acceleration card of a computing device in the computing system; Configuring routing information of the bus bridge according to the connection relationship between the acceleration card and the bus bridge and the southbound address of the acceleration card, so that the bus bridge forwards data between southbound interfaces of different acceleration cards according to the routing information; Wherein, the southbound interfaces between different acceleration cards located in the same computing device and the southbound interfaces between different acceleration cards located in different computing devices are interconnected through the bus bridge.

[0008] The present invention further provides a control device for a computing system, including: A first configuration module, configured to configure a southbound address for an acceleration card of a computing device in the computing system; A second configuration module, configured to configure routing information of the bus bridge according to the connection relationship between the acceleration card and the bus bridge and the southbound address of the acceleration card, so that the bus bridge forwards data between southbound interfaces of different acceleration cards according to the routing information; Wherein, the southbound interfaces between different acceleration cards located in the same computing device and the southbound interfaces between different acceleration cards located in different computing devices are interconnected through the bus bridge.

[0009] The present invention further provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any one of the above control methods for a computing system when executing the computer program.

[0010] The present invention further provides a computer-readable storage medium, in which a computer program is stored, and wherein the computer program implements the steps of any one of the above control methods for a computing system when executed by a processor.

[0011] The present invention further provides a computer program product, including a computer program, and the computer program implements the steps of any one of the above control methods for a computing system when executed by a processor.

[0012] Through the present invention, in a multi-machine and multi-card cluster including multiple computing devices, a first switch is used to interconnect the southbound interfaces of different acceleration cards. The first switch includes a first controller and a bus bridge. The southbound interfaces of different acceleration cards located in the same computing device and the southbound interfaces of different acceleration cards located in different computing devices are interconnected through the bus bridge. The first controller is used to configure southbound addresses for the acceleration cards, and configure the routing information of the bus bridge according to the connection relationship between the acceleration cards and the bus bridge and the southbound addresses of the acceleration cards, so as to realize the global southbound address configuration of the acceleration cards based on the first controller. Thus, the acceleration cards can directly communicate with other acceleration cards in the same computing device or acceleration cards in other computing devices based on the southbound addresses, and the bus bridge forwards data between the southbound interfaces of different acceleration cards according to the routing information, without passing through the forwarding path from the northbound interface of the acceleration card to the central processing unit, without the need for upper-layer software to specify the forwarding path, and there is no problem of long communication paths and communication congestion caused by other acceleration cards forwarding data between non-point-to-point interconnected acceleration cards in the point-to-point interconnection scheme, thereby improving the communication efficiency of the acceleration card interconnection in the multi-machine and multi-card cluster. In addition, compared with the point-to-point interconnection scheme of acceleration cards, the present invention also improves the scalability of the computing system, can flexibly construct a multi-host supernode system, and provides an easy-to-popularize and cost-effective solution for realizing a multi-card supernode system. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] To more clearly illustrate the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0014] Figure 1 FIG. is a schematic diagram of point-to-point interconnection of southbound interfaces of acceleration cards; Figure 2 FIG. is a schematic structural diagram of a computing system provided by an embodiment of the present invention; Figure 3 FIG. is a flowchart of global southbound address configuration in a computing system provided by an embodiment of the present invention; Figure 4 FIG. is a schematic structural diagram of another computing system provided by an embodiment of the present invention; Figure 5 For Figure 4 FIG. is a schematic diagram of southbound interface data forwarding of an acceleration card in the shown computing system. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0016] It should be noted that in the description of the present invention, the terms "including", "comprising" or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present invention are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0017] To enable those skilled in the art of the present technology to better understand the solution of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0018] Here, some key terms used in the embodiments of the present invention will be explained first.

[0019] With the rapid development of artificial intelligence (AI) technology, the demand for computing power has increased exponentially. Especially in typical AI applications such as generative artificial intelligence, large models, and deep learning, the heterogeneous acceleration server whole-system also faces many challenges such as improving computing power efficiency and system performance, and continuous research and innovation are needed to support the continuous development of high-performance computing power infrastructure.

[0020] Interconnection between acceleration cards refers to the connection mechanism for high-speed data interaction and collaborative work among multiple acceleration cards in environments such as data centers or high-performance computing. With the wide application of acceleration cards in fields such as artificial intelligence, scientific computing, and graphics rendering, the performance of a single acceleration card gradually becomes difficult to meet the requirements of large-scale and complex tasks. Through the interconnection between acceleration cards, a powerful acceleration card cluster or multi-acceleration card system can be built to realize the integration and expansion of computing resources and improve the processing capacity and efficiency of the overall system.

[0021] Accelerator card interconnection has a wide range of application scenarios. For example, in the training of deep learning models, such as large-scale image recognition and natural language processing models, the amount of data and computational volume are huge. Through the interconnection between accelerator cards, multiple accelerator cards can process different batches of data or different layers of neural network calculations in parallel. For instance, when training a deep neural network with hundreds of millions of parameters, multiple interconnected accelerator cards can simultaneously calculate the gradients during the forward and backward propagation processes, accelerating the model convergence speed. In the inference phase, for applications with high real-time requirements, such as video analysis or object detection in autonomous driving, multi-accelerator card interconnection can increase the processing frame rate, ensuring that the system can respond in a timely manner.

[0022] The interconnection solutions between accelerator cards are mainly divided into two types: northbound interconnection of accelerator cards and southbound interconnection of accelerator cards.

[0023] In the computer field, the northbound interface is defined as an interface that provides access and management capabilities to upper-layer systems (such as operators and network management platforms), and is usually used for data reporting, status monitoring, or resource scheduling. For example, in the Software-Defined Networking (SDN) architecture, the northbound interface allows service applications to call underlying network resources through the controller. The southbound interface is defined as an interface that sends configuration commands or management commands to lower-layer devices (such as switches and routers), focusing on device control. For example, the SDN controller issues flow tables to switches through the southbound interface (such as OpenFlow).

[0024] Therefore, after installing accelerator cards such as Graphics Processing Unit (GPU) in the server, the communication interface used by the accelerator card to connect to the host side is usually defined as the northbound interface of the accelerator card, and other communication interfaces of the accelerator card are defined as the southbound interface of the accelerator card.

[0025] The northbound interface of the accelerator card is used to connect to the local central processing unit. In the northbound interconnection solution of accelerator cards, for the interconnection of accelerator cards across host domains, since the accelerator cards on different hosts are assigned addresses by the host where they are located, neither the host nor the accelerator card can directly communicate with other hosts and accelerator cards on other hosts based on local addresses. Therefore, address conversion between different host domains is usually implemented by a switch. In the cross-host domain communication of accelerator cards, the accelerator card sends data to the local central processing unit, and the central processing unit sends it to the switch. The switch uses the non-transparent bridge method to implement address conversion between different hosts, and then forwards the data to the target host and the target accelerator card.

[0026] It can be seen that in such an interconnection solution, accelerator card interconnection, especially the interconnection of accelerator cards across host domains, all requires passing through the central processing unit. The forwarding path is long and involves complex address conversion, resulting in low communication efficiency.

[0027] The southbound interconnection solution between acceleration cards refers to a solution that directly connects the southbound interfaces of different acceleration cards to address the problem of low communication efficiency in the northbound interconnection solution of acceleration cards. However, building a system based on the southbound interconnection between acceleration cards often requires dedicated hardware devices (such as high-speed interconnection chips, high-end servers, etc.) and complex software configurations, which results in higher costs. For example, in related technologies, implementing the southbound interconnection of acceleration cards requires the support of specific acceleration card models, and additional design and verification work are also required for integration on the server platform.

[0028] On the other hand, the southbound interconnection system of acceleration cards adopted in related technologies often uses point-to-point interconnection, which has great limitations and performance bottlenecks in communication.

[0029] Figure 1 It is a schematic diagram of the point-to-point interconnection of the southbound interfaces of acceleration cards.

[0030] Figure 1 It shows the architecture of the interconnection of the southbound interfaces of acceleration cards between two computing devices (denoted as Host 1 and Host 2). Each host is connected to 8 acceleration cards, and these 8 acceleration cards are divided into two groups, and there is an interconnection between the acceleration cards within the group. In addition, high-efficiency interconnection is achieved between the acceleration cards across hosts through the southbound interfaces. Specifically, between Host 1 and Host 2, the acceleration cards with the same number are connected through the southbound interfaces, indicating the interconnection of the southbound interfaces of the acceleration cards, thus realizing the interconnection communication between hosts.

[0031] It can be seen that Figure 1 in, the communication between the acceleration cards among multiple hosts is point-to-point. For example, when the acceleration card 1 of Host 1 communicates with the acceleration card 2 of Host 2, since there is no directly defined interconnection communication link, it is necessary for the upper-layer software to define the specific communication path, such as it can be defined as "Host 1 acceleration card 1 → Host 2 acceleration card 1 → Host 2 acceleration card 2" or "Host 1 acceleration card 1 → Host 1 acceleration card 2 → Host 2 acceleration card 2", which increases the workload of the upper-layer software when initiating communication and requires the upper-layer software to select the best communication path. In addition, this point-to-point interconnection solution will also cause a large bandwidth bottleneck. For example, when the acceleration card 1 of Host 1 communicates with the acceleration cards 1, 2, 5, and 6 of Host 2, it may all require the path "Host 1 acceleration card 1 → Host 2 acceleration card 1", resulting in a large communication blockage. And the scalability of this point-to-point interconnection solution is poor. As Figure 1 shown in the dual-host interconnection demonstrated in, it is very difficult to extend to multi-host interconnection, and it will also bring more bandwidth bottlenecks.

[0032] To solve the problem of low communication efficiency in the interconnection of acceleration cards across hosts in related technologies, the present invention provides a computing system, a control method, a device, equipment, a medium, and a switch. In a multi-machine multi-card cluster including multiple computing devices, a first switch is used to implement the interconnection between the southbound interfaces of different acceleration cards. The first switch includes a first controller and a bus bridge. The southbound interfaces of different acceleration cards located in the same computing device and the southbound interfaces of different acceleration cards located in different computing devices are interconnected through the bus bridge. The first controller is used to configure a southbound address for the acceleration card and configure the routing information of the bus bridge according to the connection relationship between the acceleration card and the bus bridge and the southbound address of the acceleration card, so as to realize the global southbound address configuration of the acceleration card based on the first controller. Thus, the acceleration card can directly communicate with other acceleration cards in the computing device where it is located or acceleration cards of other computing devices based on the southbound address, and the bus bridge forwards data between the southbound interfaces of different acceleration cards according to the routing information. It neither needs to pass through the forwarding path from the northbound interface of the acceleration card to the central processing unit, nor does it need the upper-layer software to specify the forwarding path, and there is no problem that the communication path is long and communication congestion occurs due to the need for other acceleration cards to forward data between non-point-to-point interconnected acceleration cards in the point-to-point interconnection scheme, thereby improving the communication efficiency of the interconnection of acceleration cards in the multi-machine multi-card cluster. In addition, compared with the point-to-point interconnection scheme of acceleration cards, the present invention also improves the scalability of the computing system, can flexibly construct a multi-host supernode system, and provides an easy-to-popularize and cost-effective solution for realizing a multi-card supernode system.

[0033] Figure 2 It is a schematic structural diagram of a computing system provided by an embodiment of the present invention.

[0034] As Figure 2 shown, the computing system provided by the present invention may include multiple computing devices and a first switch; the first switch includes a first controller and a bus bridge; the computing device is provided with an interface for connecting an acceleration card, and the southbound interfaces of different acceleration cards located in the same computing device and the southbound interfaces of different acceleration cards located in different computing devices are interconnected through the bus bridge. The first controller is used to configure a southbound address for the acceleration card and configure the routing information of the bus bridge according to the connection relationship between the acceleration card and the bus bridge and the southbound address of the acceleration card. The bus bridge is used to forward data between the southbound interfaces of different acceleration cards according to the routing information.

[0035] In the embodiment of the present invention, the computing device may be a server.

[0036] It should be noted that an acceleration card is a hardware device used to improve the performance of a computing device and accelerate the data transfer rate. It is usually installed in the Peripheral Component Interconnect Express (PCIe) slot of the computing device and connected to the motherboard of the computing device. The acceleration card includes a circuit structure for implementing relevant data operations, which can specifically be presented as: a Graphics Processing Unit (GPU), a Field-Programmable Gate Array (FPGA), etc.

[0037] The acceleration card is usually plugged into the PCIe interface on the motherboard or the backplane of the computing device through a gold finger. This PCIe interface is a northbound interface used for communication between the acceleration card and the local host unit. To achieve cross-host domain interconnection of the acceleration card, an interface for the acceleration card to support interconnection with a remote acceleration card is required. In the embodiments of the present invention, an acceleration card configured with a southbound interface is used to achieve cross-host domain interconnection of the acceleration card. Then, the southbound interface of the acceleration card can be connected to the first switch through the Peripheral Component Interconnect Express. Specifically, the acceleration card adopted in the embodiments of the present invention can be provided with a Peripheral Component Interconnect Express switching module and a southbound interface. The first end of the Peripheral Component Interconnect Express switching module is connected to the gold finger of the acceleration card, the second end of the Peripheral Component Interconnect Express switching module is connected to the controller of the acceleration card (i.e., the computing core of the acceleration card), and the third end of the Peripheral Component Interconnect Express switching module is connected to the southbound interface of the acceleration card, so as to divide the PCIe signal introduced from the gold finger into a signal for the southbound interface for interconnection with a remote acceleration card.

[0038] As Figure 2 shown, the southbound interface of the acceleration card is connected to the first switch. Through the first switch, data forwarding can be achieved between the southbound interfaces of different acceleration cards within the same computing device, and data forwarding can also be achieved between the southbound interfaces of different acceleration cards in different computing devices. Then, the bus bridge forwards data between the southbound interfaces of different acceleration cards according to the routing information, which may include: the host unit initiates a communication request for a target acceleration card at a remote end, and the acceleration card of the computing device where the host unit is located sends the data packet of the communication request to the bus bridge, so that the bus bridge forwards the data packet to the target acceleration card according to the routing information.

[0039] Applying the acceleration card southbound interface interconnection architecture provided by the embodiments of the present invention, acceleration cards located in the same computing device and in different computing devices can all achieve many-to-many direct connection. Compared with the traditional acceleration card interconnection solution based on the northbound interface of the acceleration card, there is no need to forward through the host unit anymore. Compared with the traditional southbound interface point-to-point interconnection solution, there is no need for the upper-layer software to specify the forwarding path, and there will be no bandwidth bottleneck caused by multiple requests occupying the same link.

[0040] The first switch provided by the present invention is composed of a first controller and a bus bridge. One first switch may include one or more bus bridges. If multiple bus bridges are included, the multiple bus bridges may form a tree structure, a star structure, a ring structure, etc. to meet different topology requirements. Taking the tree structure as an example, the bus bridge at the bottom layer is directly connected to the southbound interface of the acceleration card, and the bus bridges at other layers are connected to the upstream interfaces of the bus bridges at the lower layer through the downstream interfaces until the bus bridge at the top layer.

[0041] It should be noted that in the embodiments of the present invention, the communication interface of the bus bridge facing the acceleration card side is denoted as the downstream interface, and the communication interface on the other side of the bus bridge is denoted as the upstream interface.

[0042] A complete structure formed by the above multiple bus bridges can be denoted as a switching node. One first switch may include one or more switching nodes.

[0043] To achieve a larger expansion of the computing system, multiple first switches for interconnecting the southbound interfaces of acceleration cards of multiple computing devices may also be used.

[0044] Figure 2 The architecture of interconnecting two computing devices shows the solution applying the embodiments of the present invention. Taking the interconnection of two computing devices (denoted as host 1 and host 2) as an example, each host is connected with 8 acceleration cards. These 8 acceleration cards are divided into two groups, and there is an interconnection between the acceleration cards within the group. In addition, the interconnection of the southbound interfaces of the acceleration cards is realized by the first switch. All the acceleration cards installed in the computing device can be connected to the first switch, or only some of the acceleration cards can be connected to the first switch.

[0045] In the embodiments of the present invention, a bus bridge is applied to realize data forwarding between the southbound interfaces of acceleration cards. The bus bridge (PCIe Bridge) of the PCIe bus is an intermediate node of the PCIe bus topology structure, responsible for connecting different types of PCIe bus links and realizing transmission between different buses. The main components of the bus bridge include:

[0046] Physical Layer: Responsible for the transmission of physical signals, including sending and receiving high-speed serial computer expansion bus signals.

[0047] Link Layer: Responsible for initializing, training, and maintaining the link to ensure the reliability of data transmission.

[0048] Transaction Layer: Responsible for handling transaction layer protocols, including read / write requests, configuration space access, etc.

[0049] Configuration Space: Each bus bridge has a configuration space that contains the identification of the device.

[0050] Information, status registers, and configuration registers.

[0051] Routing Table: Used to store routing information and determine the transmission path of data packets.

[0052] Address Translator: Performs address translation to ensure that data packets are correctly routed to the target device.

[0053] After the first switch is powered on, the first controller is used to initialize the PCIe system of the first switch, start enumerating all connected PCIe devices, including acceleration cards and bus bridges, and read the configuration space of the devices. The first controller sets the link parameters and configuration registers of the devices according to the configuration space information of the devices. The first controller activates the PCle link to make the devices ready to receive and send data.

[0054] To solve the cross-host domain communication problem of acceleration cards, in the embodiment of the present invention, the first controller uniformly configures southbound addresses for acceleration cards located in different computing devices, so that each acceleration card in the computing system uniquely corresponds to a section of southbound addresses, and the southbound addresses correspond to the addresses in the storage space of the acceleration card.

[0055] Then, in the embodiment of the present invention, the first controller configuring southbound addresses for acceleration cards may include: the first controller obtains the unique identifier of the acceleration card in the computing system and the storage configuration information of the acceleration card, allocates southbound addresses for the acceleration card according to the unique identifier of the acceleration card and the storage configuration information of the acceleration card, and writes the corresponding southbound addresses into the acceleration card. Among them, the storage configuration information of the acceleration card may be the video memory size of the acceleration card. For example, if the video memory size of each acceleration card is 32GB, then the first controller may allocate addresses 0~32GB for acceleration card 1, allocate addresses 32GB~64GB for acceleration card 2... and so on.

[0056] In an embodiment of the present invention, the first controller writing the southbound address to the acceleration card correspondingly may include: the first controller writing the southbound address of the acceleration card to the base address register of the acceleration card. The base address register (BAR) of the acceleration card is a component of the PCI configuration space, and is used to define the size of the configuration space required by the PCI device and the occupied space. After the first controller writes the allocated southbound address range of the acceleration card to the base address register of the acceleration card, it can help the host operating system understand the address space of each acceleration card and avoid address conflicts.

[0057] In an embodiment of the present invention, the bus bridge forwarding data between the southbound interfaces of different acceleration cards according to the routing information may include: the acceleration card reporting the southbound address to the host unit of the computing device where it is located; the host units of different computing devices interacting with the southbound addresses of the acceleration cards to share the global southbound address table of the computing system; the acceleration card determining the target southbound address of the target acceleration card according to the global southbound address table, and sending the data packet carrying the target southbound address to the bus bridge through the southbound interface, so that the bus bridge forwards the data packet to the target acceleration card. Specifically, after the first controller allocates the southbound address for the acceleration card, the host unit of the computing device where the acceleration card is located can obtain the southbound address of the acceleration card through the Host OS Client, and summarize the southbound addresses of the local acceleration cards and interact with the host units of other computing devices to realize the construction of the global southbound address table of the computing system. Thus, the host units and acceleration cards of each computing device in the computing system can perform many-to-many data communication based on the southbound interface of the acceleration card based on the global southbound address table.

[0058] Based on this, in some application scenarios of the embodiment of the present invention, when the acceleration card needs to send data to the target acceleration card, it can determine the target southbound address in the target acceleration card based on the global southbound address table, and send the data packet carrying the target southbound address to the bus bridge through the southbound interface, and the bus bridge forwards it to the target acceleration card. In other application scenarios of the embodiment of the present invention, when the host unit needs to send data to the target acceleration card, it can determine the target southbound address in the target acceleration card based on the global southbound address table, and send the data packet carrying the target southbound address to the acceleration card through the northbound interface of the acceleration card, and the acceleration card sends the data packet from the southbound interface to the bus bridge, and the bus bridge forwards it to the target acceleration card.

[0059] In an embodiment of the present invention, since the first controller needs to allocate a southbound address for the acceleration card according to the unique identifier of the acceleration card and the storage configuration information of the acceleration card, this information can be obtained from a node capable of communicating with each computing device in the computing system, and this node can be the management node of the computing system. In a computing device cluster, a management node is usually set up to monitor and manage the devices in the cluster. This management node can be used for power-on and power-off management of computing devices, obtaining hardware status information of computing devices (including status information of acceleration cards) by communicating with the management controllers on the computing devices, and so on.

[0060] Then, the computing system provided by the embodiment of the present invention may further include a management node. The first controller configuring a southbound address for the acceleration card may include: the first controller allocating a southbound address for the acceleration card according to the unique identifier of the acceleration card in the computing system and the storage configuration information of the acceleration card received from the management node, and correspondingly writing the southbound address into the acceleration card. This management node is used to communicate with the computing device to obtain the identifier information of the acceleration card on the computing device and the storage configuration information of the acceleration card, and allocate a unique identifier for the acceleration card in the computing system according to the identifier information of the acceleration card.

[0061] This management node may also be connected to the management controller of the first switch to perform status management on the first switch. The management node performing status management on the first switch may include at least one of power-on and power-off management of the first switch, status monitoring of the bus bridge, and parameter configuration of the first switch.

[0062] The "management controller" in the embodiment of the present invention may refer to a management controller board, which includes bus controllers, logic programmable units, memories and other components in addition to the out-of-band monitoring master controller. The "out-of-band management controller" in the embodiment of the present invention may also refer to an out-of-band monitoring master controller on a baseboard management controller board, which may be a single-core processor or a multi-core processor. In the embodiment of the present invention, the out-of-band management controller or the out-of-band monitoring master controller on the out-of-band management controller board may use an ARM processor, on which the operating system of the out-of-band management controller runs. In the computing device and the first switch, the management controller may be installed on the motherboard or mainboard of the device, and use but not limited to the Intelligent Platform Management Interface (IPMI) protocol to monitor the status of the hardware in the device by monitoring the sensors in the device. The management controller can use an integrated circuit bus (such as an Inter-Integrated Circuit (I2C)) or an Intelligent Platform Management BUS (IPMB) to communicate with modules inside the device, such as the South Bridge (Platform Controller Hub, PCH), memory (such as Dual-Inline-Memory-Modules (DIMM)), power supply, etc. The management controller can also connect to sensors in the device through the integrated circuit bus or the intelligent platform management bus, and monitor the status of the hardware in the device through the sensors, such as temperature, humidity, power supply voltage, fan speed, communication parameters and operating system (OS) functions, etc., and process when any of these variables exceeds the specified range. The processing methods may include but are not limited to: recording logs, triggering the reset of abnormal components and other abnormal solutions, sending remote alarm information to the management node to notify the operation and maintenance personnel to handle the abnormality, etc.

[0063] like Figure 2 As shown, the first switch can be connected to the communication module of the management node to exchange information with the management node. The management node can obtain the address information of the accelerator card from the computing device. The first controller can obtain the address information of the accelerator card from the management node to configure the bus bridge. The management node can synchronize the routing table obtained from the first controller to the computing device.

[0064] In some other optional implementations of the embodiments of the present invention, the operation and maintenance personnel may also write the unique identifier of the accelerator card and the storage configuration information of the accelerator card into the first controller, or the operation and maintenance personnel may directly write the southbound address of each accelerator card into the first controller.

[0065] The first controller configures the routing information of the bus bridge according to the connection relationship between the acceleration card and the bus bridge and the southbound address of the acceleration card, which may include: the first controller configures the address range of the bus bridge according to the southbound address of the acceleration card connected to the bus bridge. Specifically, the registers of the configured bus bridge are used to configure the starting address (Addresss Base) and the ending address (Addresss Limit) of the acceleration card corresponding to the bus bridge, so that the bus bridge forwards the data packets of the acceleration card according to the local address range.

[0066] In the embodiments of the present invention, in addition to the connection of the southbound interface, the northbound interface of the acceleration card may be connected to the host unit of the computing device where it is located. In some alternative embodiments of the embodiments of the present invention, the acceleration card is further used to report the southbound address of the acceleration card to the host unit of the computing device where it is located. On this basis, the bus bridge forwards data between the southbound interfaces of different acceleration cards according to the routing information, which may include: the host unit initiates a communication request to the target acceleration card according to the target southbound address of the remote target acceleration card, and the acceleration card of the computing device where the host unit is located sends the data packet of the communication request to the bus bridge, so that the bus bridge forwards the data packet to the target acceleration card according to the routing information.

[0067] In some alternative embodiments of the embodiments of the present invention, different acceleration cards located in the same computing device are interconnected through a bridge board. As Figure 2 shown, within host 1, acceleration card 1, acceleration card 2, acceleration card 5, and acceleration card 6 can be directly interconnected. If the communication interfaces of the acceleration cards are insufficient, the inter-card interfaces of multiple acceleration cards located in the same computing device can be connected to the bridge board, and the inter-card interconnection within the machine of each acceleration card can be realized through the bridge board.

[0068] Figure 3 It is a flowchart of global southbound address configuration in a computing system provided by an embodiment of the present invention.

[0069] In a computing system including four computing devices (host 1 to host 4), taking host 1 as an example, as Figure 3 shown, the process of global southbound address configuration in the computing system may include:

[0070] ① The management node interacts with the first controller of the first switch, and the cluster management software sends the unique identifier of the acceleration card to the switch management software.

[0071] ② The first controller configures the routing information of the bus bridge of the first switch based on the switch management software.

[0072] ③ The first controller writes the southbound address of the acceleration card into the acceleration card based on the switch management software.

[0073] ④The host operating system client of the computing device obtains the southbound address of the local acceleration card.

[0074] ⑤The management node obtains the southbound addresses of the acceleration cards on each computing device collected by the host operating system clients of each computing device based on the cluster management software.

[0075] ⑥The management node summarizes the southbound addresses of each acceleration card on each computing device based on the cluster management software, constructs a global southbound address table, and updates the global southbound address table to each computing device.

[0076] Based on the above embodiments, the embodiments of the present invention further illustrate the configuration method of the bus bridge.

[0077] In the embodiments of the present invention, the first controller configures the routing information of the bus bridge according to the connection relationship between the acceleration card and the bus bridge and the southbound address of the acceleration card, including: the first controller configures the routing information of the bus bridge according to the southbound address of the acceleration card connected to the downstream interface of the bus bridge.

[0078] In some optional embodiments of the embodiments of the present invention, the first controller configures the routing information of the bus bridge according to the southbound address of the acceleration card connected to the downstream interface of the bus bridge, which may include: the first controller configures the routing information of the nth layer bus bridge according to the southbound address of the acceleration card directly connected to the nth layer bus bridge; for other layer bus bridges, the routing information of the i-th layer bus bridge is configured according to the routing information of the (i + 1)-th layer bus bridge connected to the i-th layer bus bridge; where i [1,n], both i and n are positive integers. Among them, the nth layer bus bridge is the bus bridge directly connected to the southbound interface of the acceleration card, and the first layer bus bridge is the bus bridge farthest from the southbound interface of the acceleration card.

[0079] The address ranges of other layer bus bridges can be controlled separately by the first controller or configured in a recursive form.

[0080] In some optional embodiments of the embodiments of the present invention, for other layer bus bridges, configuring the routing information of the current layer bus bridge according to the routing information of the next layer bus bridge may include: the first controller configures the routing information of the i-th layer bus bridge according to the routing information of the (i + 1)-th layer bus bridge connected to the i-th layer bus bridge.

[0081] In some other alternative embodiments of the embodiments of the present invention, for other layer bus bridges, configuring the routing information of the i-th layer bus bridge according to the routing information of the (i + 1)-th layer bus bridge connected to the i-th layer bus bridge may further include: the i-th layer bus bridge configures the local routing information according to the routing information of the (i + 1)-th layer bus bridge received from the downlink interface.

[0082] That is to say, when the first switch includes multiple layer bus bridges, the first controller can determine the address range of the bottom layer bus bridge according to the direct connection relationship between the bottom layer bus bridge and the southbound interface of the acceleration card, and configure the register of the bottom layer bus bridge with this. Then, starting from the bottom layer bus bridge, after each layer bus bridge determines its local address range, it transmits it to the upper layer bus bridge through the uplink interface. The current layer bus bridge can then determine its local address range according to the address range of the lower layer bus bridge received from the downlink interface, so as to complete the configuration of all layer bus bridges through a recursive form.

[0083] Through the several methods introduced above, the configuration of the multi-layer bus bridge from bottom to top can be achieved.

[0084] Figure 4 It is a schematic structural diagram of another computing system provided by the embodiments of the present invention.

[0085] As Figure 4 shown, taking the interconnection of two computing devices (denoted as host 1 and host 2) as an example, a three-layer bus bridge is adopted in the first controller. Among them, the bottom layer, that is, the third layer, includes a total of 8 L3 bus bridges (L3-1, L3-2, L3-3, L3-4, L3-5, L3-6, L3-7, L3-8), with a total of 16 downlink ports, connecting a total of 16 acceleration cards of the two computing devices. The second layer includes 2 L2 bus bridges (L2-1, L2-2), with a total of 8 downlink ports, connecting 8 L3 bus bridges. The top layer, that is, the first layer, includes 1 L1 bus bridge (L1-1), connecting two L2 bus bridges.

[0086] The first controller can be connected only to the L1 bus bridge or to each bus bridge separately.

[0087] Then Figure 4 The link routing establishment process of the acceleration card southbound interface interconnection architecture of the shown computing system includes:

[0088] According to its physical hardware design, the southbound interface of the acceleration card can apply for a memory space of a specified size. For example, the acceleration card 1 of host 1 can apply for the address space from 0x1_0000_0000 to 0x2_0000_0000, and the address space applied by the acceleration card 5 of host 1 is from 0x2_0000_0000 to 0x3_0000_0000.

[0089] Set the address range (including the start address Addresss Base and the end address Addresss Limit) on the L3-1 bus bridge to limit the memory address range of the L3-1 bus bridge to 0x1_0000_0000 to 0x3_0000_0000. Similarly, the L3-2 bus bridge can be set to 0x3_0000_0000 to 0x5_0000_0000 (for acceleration card 2 and acceleration card 6 of host 1), the L3-3 bus bridge can be set to 0x5_0000_0000 to 0x7_0000_0000 (for acceleration card 3 and acceleration card 7 of host 1), the L3-4 bus bridge can be set to 0x7_0000_0000 to 0x9_0000_0000 (for acceleration card 4 and acceleration card 8 of host 1), and the address ranges of the remaining four L3 bus bridges can be set in sequence.

[0090] After that, set the address range on the L2 bus bridge. The address range of the L2-1 bus bridge can be set to 0x1_0000_0000 to 0x9_0000_0000, and this address range covers the L3-1 bus bridge, the L3-2 bus bridge, the L3-3 bus bridge, and the L3-4 bus bridge. Similarly, the L2-2 bus bridge can set the address range.

[0091] Finally, set the address range on the L1-1 bus bridge to 0x1_0000_0000 to 0x10_0000_0000, and this address range covers all the acceleration cards in the computing system through the three-layer bus bridge.

[0092] In some other alternative embodiments of the embodiment of the present invention, the first controller configures the routing information of the bus bridge according to the southbound address of the acceleration card connected to the downstream interface of the bus bridge, and may further include: the first controller configures the routing information of the first-layer bus bridge according to the southbound addresses of the acceleration cards connected to the first switch; the first controller configures the routing information of the other-layer bus bridges according to the topological structure of the bus bridges in the first switch; wherein, the downstream interface of the nth-layer bus bridge in the first switch is directly connected to the southbound interface of the acceleration card, and n is a positive integer.

[0093] That is to say, in addition to the above-described bottom-up bus bridge configuration method, the first controller can directly perform a top-down bus bridge configuration or a parallel bus bridge configuration according to the topology of the bus bridge, thereby improving the configuration efficiency of the bus bridge.

[0094] Based on the above embodiments, the embodiments of the present invention further illustrate the manner in which the bus bridge forwards data between the southbound interfaces of different acceleration cards according to the routing information.

[0095] In the embodiments of the present invention, the bus bridge forwards data between the southbound interfaces of different acceleration cards according to the routing information, which may include: after receiving a data packet of the acceleration card, the bus bridge performs address resolution on the data packet to obtain the target southbound address of the data packet; the bus bridge forwards the data packet according to the local routing information and the target southbound address.

[0096] In a specific implementation, if the target address information of the data packet belongs to the local address range of the bus bridge, the bus bridge can directly forward the data packet to the target acceleration card according to the downstream port corresponding to the target address information.

[0097] When the first switch includes a multi-layer bus bridge structure composed of multiple bus bridges, the target address information of the data packet received by the bus bridge may not belong to the local address range. At this time, the bus bridge forwards the data packet to other bus bridges for further forwarding until it is forwarded to the bus bridge belonging to its address range, and is forwarded to the target acceleration card through the downstream port of the bus bridge.

[0098] Then, the bus bridge forwards the data packet according to the local routing information and the target southbound address, which may include: if the target southbound address is within the address range of the local routing information, the data packet is sent to the target acceleration card through the corresponding downstream interface; if the target southbound address is not within the local address range, the data packet is sent to the upper-layer bus bridge through the upstream interface.

[0099] When the multi-layer bus bridge is in a tree structure, after receiving a data packet, if the bus bridge resolves that the target address information of the data packet belongs to the local address range, it can directly send the data packet to the target acceleration card through the corresponding downstream interface. If the target address information is not within the local address range, it means that the target address information of the acceleration card belongs to the address range of other parallel bus bridges. Therefore, at this time, the bus bridge sends the data packet to the upper-level bus bridge through the upstream interface for forwarding to the bus bridge with the target address information in its address range.

[0100] In addition to parsing the target address information, in the embodiments of the present invention, the bus bridge can also be used to parse a data packet to obtain the transaction tag information of the data packet, and determine the transaction type of the data packet according to the transaction tag information of the data packet to perform corresponding processing.

[0101] To improve the communication efficiency of different transaction types, the bus bridge can also be used to parse and obtain the transaction tag information to perform corresponding processing. In some optional embodiments of the embodiments of the present invention, the bus bridge determines the transaction type of the data packet according to the transaction tag information of the data packet to perform corresponding processing, which may include: if the transaction type of the data packet is a non-issue transaction, the bus bridge reserves resources for the completed data packet corresponding to the data packet locally.

[0102] Based on the southbound interface interconnection architecture of the accelerator card proposed in the embodiments of the present invention, the scenarios where the accelerator card communication is a read request and a write request are described below respectively.

[0103] When the bus bridge receives a data packet of a read request from accelerator card A, it parses the header information of the data packet to determine the port and address of the target accelerator card of the request, so as to forward the data packet to the target accelerator card. The target accelerator card then returns the data to be read to accelerator card A through the bus bridge.

[0104] When the bus bridge receives a data packet of a write request from accelerator card A, it parses the header information of the data packet to determine the port and address of the target accelerator card of the request, and the data to be written, and sends the write request to the target accelerator card. After the target accelerator card completes the write operation, it sends an acknowledgment signal to the bus bridge. After receiving the acknowledgment signal, the bus bridge forwards it to accelerator card A to indicate that the write operation has been successfully completed. To improve the write efficiency, the target accelerator card can send the acknowledgment signal to the bus bridge immediately after confirming that the data to be written has been written to the computing device where it is located, so that accelerator card A can know that the write operation has been successfully completed, and then the target accelerator card moves the data to be written from the system memory to the local memory.

[0105] Figure 5 For Figure 4 A schematic diagram of data forwarding of the southbound interface of an accelerator card in the computing system shown.

[0106] Taking Figure 4 the communication between accelerator card 1 of host 1 and accelerator card 8 of host N in the southbound interface interconnection architecture of the accelerator card shown as an example for description.

[0107] Host 1 initiates a communication request between accelerator card 1 and accelerator card 8 of host N, and the target address of the communication is 0xf8000_0000.

[0108] The data packet of the communication request is sent from the acceleration card 1 of host 1 to the L3-1 bus bridge. If the target address information 0xf8000_0000 of the data packet parsed by the L3-1 bus bridge does not belong to the local address range (0x1_0000_0000~0x3_0000_0000), the data packet will be sent to the upper-level bus bridge, i.e., the L2-1 bus bridge, through the upstream interface.

[0109] If the target address information 0xf8000_0000 of the data packet parsed by the L2-1 bus bridge does not belong to the local address range (0x1_0000_0000~0x9_0000_0000), the data packet will be sent to the upper-level bus bridge, i.e., the L1-1 bus bridge, through the upstream interface.

[0110] If the target address information 0xf8000_0000 of the data packet parsed by the L1-1 bus bridge belongs to the local address range (0x1_0000_0000~0x10_0000_0000) and within the address range of the L2-2 bus bridge (0x9_0000_0000~0x10_0000_0000) corresponding to another downstream interface, the data packet will be sent to the L2-2 bus bridge through the corresponding downstream interface.

[0111] If the target address information 0xf8000_0000 of the data packet parsed by the L2-2 bus bridge belongs to the local address range (0x9_0000_0000~0x10_0000_0000) and within the address range of the L3-3 bus bridge (0x_e0000_0000~0x10_0000_0000) connected to one of the downstream interfaces, the data packet will be sent to the L3-3 bus bridge through the corresponding downstream interface.

[0112] If the target address information 0xf8000_0000 of the data packet parsed by the L3-3 bus bridge belongs to the local address range (0x_e0000_0000~0x10_0000_0000), it will be sent to the acceleration card 8 of host N, completing a communication process from the acceleration card 1 of host 1 to the acceleration card 8 of host N.

[0113] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0114] An embodiment of the present invention further provides a switch, which may include: a first controller and a bus bridge; the southbound interfaces of different acceleration cards located in the same computing device and the southbound interfaces of different acceleration cards located in different computing devices are interconnected through the bus bridge; the first controller is used to configure a southbound address for the acceleration card, and configure the routing information of the bus bridge according to the connection relationship between the acceleration card and the bus bridge and the southbound address of the acceleration card; the bus bridge is used to forward data between the southbound interfaces of different acceleration cards according to the routing information.

[0115] In an embodiment of the present invention, the first controller configures a southbound address for the acceleration card, which may include: the first controller obtains the unique identifier of the acceleration card in the computing system and the storage configuration information of the acceleration card, allocates a southbound address for the acceleration card according to the unique identifier of the acceleration card and the storage configuration information of the acceleration card, and writes the southbound address correspondingly into the acceleration card.

[0116] In an embodiment of the present invention, the first controller writes the southbound address correspondingly into the acceleration card, which may include: the first controller writes the southbound address of the acceleration card into the base address register of the acceleration card.

[0117] In an embodiment of the present invention, the bus bridge forwards data between the southbound interfaces of different acceleration cards according to the routing information, which may include: the acceleration card reports the southbound address to the host unit of the computing device where it is located; the host units of different computing devices exchange the southbound addresses of the acceleration cards to share the global southbound address table of the computing system; the acceleration card determines the target southbound address of the target acceleration card according to the global southbound address table, and sends the data packet carrying the target southbound address to the bus bridge through the southbound interface, so that the bus bridge forwards the data packet to the target acceleration card.

[0118] In an embodiment of the present invention, the first controller configures the routing information of the bus bridge according to the connection relationship between the acceleration card and the bus bridge and the southbound address of the acceleration card, which may include: the first controller configures the routing information of the bus bridge according to the southbound address of the acceleration card connected to the downstream interface of the bus bridge.

[0119] In some optional implementation manners of an embodiment of the present invention, the first controller configures the routing information of the bus bridge according to the southbound address of the acceleration card connected to the downstream interface of the bus bridge, which may include: the first controller configures the routing information of the nth layer bus bridge according to the southbound address of the acceleration card directly connected to the nth layer bus bridge; for other layer bus bridges, the routing information of the i-th layer bus bridge is configured according to the routing information of the (i + 1)-th layer bus bridge connected to the i-th layer bus bridge; where i [1,n], both i and n are positive integers.

[0120] Among them, for other - layer bus bridge - devices, configuring the routing information of the i - th layer bus bridge - device according to the routing information of the (i + 1)-th layer bus bridge - device connected to the i - th layer bus bridge - device may include: The i - th layer bus bridge - device configures its local routing information according to the routing information of the (i + 1)-th layer bus bridge - device received from the downstream interface.

[0121] Alternatively, for other - layer bus bridge - devices, configuring the routing information of the current - layer bus bridge - device according to the routing information of the next - layer bus bridge - device may include: The first controller configures the routing information of the i - th layer bus bridge - device according to the routing information of the (i + 1)-th layer bus bridge - device connected to the i - th layer bus bridge - device.

[0122] In some other alternative embodiments of the embodiments of the present invention, the first controller configures the routing information of the bus bridge - device according to the south - bound address of the acceleration card connected to the downstream interface of the bus bridge - device, and may further include: The first controller configures the routing information of the first - layer bus bridge - device according to the south - bound addresses of the acceleration cards connected to the first switch; The first controller configures the routing information of other - layer bus bridge - devices according to the topological structure of the bus bridge - devices in the first switch; where the downstream interface of the n - th layer bus bridge - device in the first switch is directly connected to the south - bound interface of the acceleration card, and n is a positive integer.

[0123] In the embodiments of the present invention, the first controller configures the routing information of the bus bridge - device according to the connection relationship between the acceleration card and the bus bridge - device and the south - bound address of the acceleration card, and may include: The first controller configures the address range of the bus bridge - device according to the south - bound address of the acceleration card connected to the bus bridge - device.

[0124] In the embodiments of the present invention, the first controller configures the south - bound address for the acceleration card, and may include: The first controller assigns a south - bound address to the acceleration card according to the unique identifier of the acceleration card in the computing system received from the management node and the storage configuration information of the acceleration card, and writes the south - bound address correspondingly into the acceleration card. The management node is used to communicate with the computing device to obtain the identification information of the acceleration card on the computing device and the storage configuration information of the acceleration card, and assigns a unique identifier for the acceleration card in the computing system according to the identification information of the acceleration card.

[0125] In the embodiments of the present invention, the bus bridge - device forwards data between the south - bound interfaces of different acceleration cards according to the routing information, and may include: After receiving a data packet from the acceleration card, the bus bridge - device performs address resolution on the data packet to obtain the destination south - bound address of the data packet; The bus bridge - device forwards the data packet according to the local routing information and the destination south - bound address.

[0126] In an embodiment of the present invention, the bus bridge forwards data packets according to local routing information and a target southbound address, which may include: if the target southbound address is within the address range of the local routing information, the data packet is sent to the target acceleration card through the corresponding downstream interface; if the target southbound address is not within the local address range, the data packet is sent to the upper-layer bus bridge through the upstream interface.

[0127] In an embodiment of the present invention, the bus bridge is further configured to parse the data packet to obtain the transaction tag information of the data packet, and determine the transaction type of the data packet according to the transaction tag information of the data packet to perform corresponding processing.

[0128] In an embodiment of the present invention, the bus bridge determines the transaction type of the data packet according to the transaction tag information of the data packet to perform corresponding processing, which may include: if the transaction type of the data packet is a non-issued transaction, the bus bridge reserves resources for the completed data packet corresponding to the data packet locally.

[0129] For the specific implementation manner of the switch provided in the embodiment of the present invention, reference may be made to the first switch in the above embodiment, which will not be elaborated here.

[0130] The embodiment of the present invention further provides a control method for a computing system, including: configuring a southbound address for an acceleration card of a computing device in the computing system; configuring the routing information of the bus bridge according to the connection relationship between the acceleration card and the bus bridge and the southbound address of the acceleration card, so that the bus bridge forwards data between the southbound interfaces of different acceleration cards according to the routing information; wherein, the southbound interfaces between different acceleration cards located in the same computing device and the southbound interfaces between different acceleration cards located in different computing devices are interconnected through the bus bridge.

[0131] For the specific implementation manner of the control method for the computing system provided in the embodiment of the present invention, reference may be made to the embodiment of the above computing system, which will not be elaborated here.

[0132] The embodiment of the present invention further provides a control device for a computing system, including: a first configuration module, configured to configure a southbound address for an acceleration card of a computing device in the computing system; a second configuration module, configured to configure the routing information of the bus bridge according to the connection relationship between the acceleration card and the bus bridge and the southbound address of the acceleration card, so that the bus bridge forwards data between the southbound interfaces of different acceleration cards according to the routing information; wherein, the southbound interfaces between different acceleration cards located in the same computing device and the southbound interfaces between different acceleration cards located in different computing devices are interconnected through the bus bridge.

[0133] For the specific implementation manner of the control device for the computing system provided in the embodiment of the present invention, reference may be made to the embodiment of the above computing system, which will not be elaborated here.

[0134] An embodiment of the present invention further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in the embodiment of the control method of the above-mentioned computing system.

[0135] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in the embodiment of the control method of the above-mentioned computing system when running.

[0136] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical disks that can store computer programs.

[0137] An embodiment of the present invention further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in the embodiment of the control method of the above-mentioned computing system are implemented.

[0138] An embodiment of the present invention further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the embodiment of the control method of the above-mentioned computing system are implemented.

[0139] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.

[0140] The above has introduced in detail a computing system and its control method, device, equipment, medium, and switch provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.

Claims

1. A computing system, characterized in that, Including: Multiple computing devices and a first switch; The first switch includes a first controller and a bus bridge; An interface for connecting an acceleration card is provided on the computing device, and the southbound interfaces of different acceleration cards on the same computing device and the southbound interfaces of different acceleration cards on different computing devices are interconnected through the bus bridge; The first controller is used to configure a southbound address for the acceleration card and configure the routing information of the bus bridge according to the connection relationship between the acceleration card and the bus bridge and the southbound address of the acceleration card; The bus bridge is used to forward data between the southbound interfaces of different acceleration cards according to the routing information.

2. The computing system according to claim 1, wherein The first controller configuring a southbound address for the acceleration card includes: The first controller obtains the unique identifier of the acceleration card in the computing system and the storage configuration information of the acceleration card, allocates a southbound address for the acceleration card according to the unique identifier of the acceleration card and the storage configuration information of the acceleration card, and writes the southbound address correspondingly into the acceleration card.

3. The computing system according to claim 2, wherein The first controller writing the southbound address correspondingly into the acceleration card includes: The first controller writes the southbound address of the acceleration card into the base address register of the acceleration card.

4. The computing system according to claim 1, wherein The bus bridge forwarding data between the southbound interfaces of different acceleration cards according to the routing information includes: The acceleration card reports the southbound address to the host unit of the computing device where it is located; The host units of different computing devices exchange the southbound addresses of the acceleration cards to share the global southbound address table of the computing system; The acceleration card determines the target southbound address of the target acceleration card according to the global southbound address table, and sends a data packet carrying the target southbound address to the bus bridge through the southbound interface, so that the bus bridge forwards the data packet to the target acceleration card.

5. The computing system according to claim 1, wherein The first controller configuring the routing information of the bus bridge according to the connection relationship between the acceleration card and the bus bridge and the southbound address of the acceleration card includes: The first controller configures the routing information of the bus bridge according to the southbound address of the acceleration card connected to the downstream interface of the bus bridge.

6. The computing system according to claim 5, wherein The first controller configuring the routing information of the bus bridge according to the southbound address of the acceleration card connected to the downstream interface of the bus bridge includes: The first controller configures the routing information of the nth layer bus bridge according to the southbound address of the acceleration card directly connected to the nth layer bus bridge; For the bus bridges of other layers, the routing information of the ith layer bus bridge is configured according to the routing information of the (i + 1)th layer bus bridge connected to the ith layer bus bridge; where i [1, n], both i and n are positive integers.

7. The computing system according to claim 6, wherein For the bus bridges of other layers, the routing information of the ith layer bus bridge is configured according to the routing information of the (i + 1)th layer bus bridge connected to the ith layer bus bridge, including: The ith layer bus bridge configures the local routing information according to the routing information of the (i + 1)th layer bus bridge received from the downstream interface.

8. The computing system according to claim 6, wherein For the bus bridge of other layers, configure the routing information of the bus bridge of the current layer according to the routing information of the bus bridge of the next layer, including: The first controller configures the routing information of the i-th layer bus bridge according to the routing information of the (i + 1)-th layer bus bridge connected to the i-th layer bus bridge.

9. The computing system according to claim 5, wherein The first controller configures the routing information of the bus bridge according to the southbound address of the acceleration card connected to the downlink interface of the bus bridge, including: The first controller configures the routing information of the first layer bus bridge according to the southbound addresses of the acceleration cards connected to the first switch. The first controller configures the routing information of the bus bridges of other layers according to the topology structure of the bus bridges in the first switch. Among them, the downlink interface of the n-th layer bus bridge in the first switch is directly connected to the southbound interface of the acceleration card, and n is a positive integer.

10. The computing system according to claim 1, wherein The first controller configures the routing information of the bus bridge according to the connection relationship between the acceleration card and the bus bridge and the southbound address of the acceleration card, including: The first controller configures the address range of the bus bridge according to the southbound address of the acceleration card connected to the bus bridge.

11. The computing system according to claim 1, wherein It also includes a management node; The first controller configures the southbound address for the acceleration card, including: The first controller allocates a southbound address for the acceleration card according to the unique identifier of the acceleration card in the computing system received from the management node and the storage configuration information of the acceleration card, and writes the southbound address correspondingly into the acceleration card. The management node is used to communicate with the computing device to obtain the identification information of the acceleration card on the computing device and the storage configuration information of the acceleration card, and allocate a unique identifier for the acceleration card in the computing system according to the identification information of the acceleration card.

12. The computing system according to claim 11, wherein The management node is also connected to the management controller of the first switch to perform status management on the first switch.

13. The computing system according to claim 12, wherein The management node performs status management on the first switch, including: At least one of power-on and power-off management of the first switch, status monitoring of the bus bridge, and parameter configuration of the first switch.

14. The computing system according to claim 1, wherein The bus bridge forwards data between the southbound interfaces of different acceleration cards according to the routing information, including: After receiving the data packet of the acceleration card, the bus bridge performs address resolution on the data packet to obtain the target southbound address of the data packet. The bus bridge forwards the data packet according to the local routing information and the target southbound address.

15. The computing system according to claim 14, wherein The bus bridge forwards the data packet according to the local routing information and the target southbound address, including: If the target southbound address is within the address range of the local routing information, the data packet is sent to the target acceleration card through the corresponding downlink interface; if the target southbound address is not within the local address range, the data packet is sent to the bus bridge of the upper layer through the uplink interface.

16. The computing system according to claim 14, wherein The bus bridge is also used to parse the data packet to obtain the transaction tag information of the data packet, and determine the transaction type of the data packet according to the transaction tag information of the data packet to perform corresponding processing.

17. The computing system according to claim 16, wherein The bus bridge determines the transaction type of the data packet according to the transaction tag information of the data packet to perform corresponding processing, including: If the transaction type of the data packet is a non-issued transaction, the bus bridge reserves resources for the completed data packet corresponding to the data packet locally.

18. A switch, characterized in that, Including: A first controller and a bus bridge; The southbound interfaces of different acceleration cards located in the same computing device and the southbound interfaces of different acceleration cards located in different computing devices are interconnected through the bus bridge; The first controller is used to configure a southbound address for the acceleration card, and configure the routing information of the bus bridge according to the connection relationship between the acceleration card and the bus bridge and the southbound address of the acceleration card; The bus bridge is used to forward data between the southbound interfaces of different acceleration cards according to the routing information.

19. The switch according to claim 18, characterized in that, The first controller configures a southbound address for the acceleration card, including: The first controller obtains the unique identifier of the acceleration card in the computing system and the storage configuration information of the acceleration card, allocates a southbound address for the acceleration card according to the unique identifier of the acceleration card and the storage configuration information of the acceleration card, and writes the southbound address correspondingly into the acceleration card.

20. The switch according to claim 19, wherein The first controller writes the southbound address correspondingly into the acceleration card, including: The first controller writes the southbound address of the acceleration card into the base address register of the acceleration card.

21. The switch according to claim 18, wherein The bus bridge forwards data between the southbound interfaces of different acceleration cards according to the routing information, including: The acceleration card reports the southbound address to the host unit of the computing device where it is located; The host units of different computing devices exchange the southbound addresses of the acceleration cards to share the global southbound address table of the computing system; The acceleration card determines the target southbound address of the target acceleration card according to the global southbound address table, and sends a data packet carrying the target southbound address to the bus bridge through the southbound interface, so that the bus bridge forwards the data packet to the target acceleration card.

22. The switch according to claim 18, characterized in that, The first controller configures the routing information of the bus bridge according to the connection relationship between the acceleration card and the bus bridge and the southbound address of the acceleration card, including: The first controller configures the routing information of the bus bridge according to the southbound address of the acceleration card connected to the downstream interface of the bus bridge.

23. The switch according to claim 22, wherein The first controller configures the routing information of the bus bridge according to the southbound address of the acceleration card connected to the downstream interface of the bus bridge, including: The first controller configures the routing information of the nth layer bus bridge according to the southbound address of the acceleration card directly connected to the nth layer bus bridge; For other layer bus bridges, configure the routing information of the ith layer bus bridge according to the routing information of the (i + 1)th layer bus bridge connected to the ith layer bus bridge; where i ∈ [1, n], and both i and n are positive integers.

24. The switch according to claim 23, characterized in that, For the bus bridge of other layers, configuring the routing information of the bus bridge of the i-th layer according to the routing information of the bus bridge of the (i + 1)-th layer connected to the bus bridge of the i-th layer includes: The bus bridge of the i-th layer configures the local routing information according to the routing information of the bus bridge of the (i + 1)-th layer received from the downstream interface.

25. The switch according to claim 18, wherein The first controller configures the routing information of the bus bridge according to the connection relationship between the acceleration card and the bus bridge and the southbound address of the acceleration card, including: The first controller configures the address range of the bus bridge according to the southbound address of the acceleration card connected to the bus bridge.

26. The switch according to claim 18, wherein The first controller configures the southbound address for the acceleration card, including: The first controller allocates a southbound address for the acceleration card according to the unique identifier of the acceleration card in the computing system received from the management node and the storage configuration information of the acceleration card, and writes the corresponding southbound address into the acceleration card. The management node is used to communicate with the computing device to obtain the identifier information of the acceleration card on the computing device and the storage configuration information of the acceleration card, and allocate a unique identifier for the acceleration card in the computing system according to the identifier information of the acceleration card.

27. The switch according to claim 18, characterized in that, The bus bridge forwards data between the southbound interfaces of different acceleration cards according to the routing information, including: After receiving the data packet of the acceleration card, the bus bridge performs address resolution on the data packet to obtain the target southbound address of the data packet. The bus bridge forwards the data packet according to the local routing information and the target southbound address.

28. The switch according to claim 27, wherein The bus bridge forwards the data packet according to the local routing information and the target southbound address, including: If the target southbound address is within the address range of the local routing information, the data packet is sent to the target acceleration card through the corresponding downstream interface; if the target southbound address is not within the local address range, the data packet is sent to the bus bridge of the upper layer through the upstream interface.

29. A control method for a computing system, characterized in that, Including: Configuring a southbound address for the acceleration card of the computing device in the computing system; Configuring the routing information of the bus bridge according to the connection relationship between the acceleration card and the bus bridge and the southbound address of the acceleration card, so that the bus bridge forwards data between the southbound interfaces of different acceleration cards according to the routing information; Among them, the southbound interfaces between different acceleration cards located on the same computing device and the southbound interfaces between different acceleration cards located on different computing devices are interconnected through the bus bridge.

30. A control device for a computing system, characterized in that, Including: The first configuration module is used to configure a southbound address for the acceleration card of the computing device in the computing system; The second configuration module is used to configure the routing information of the bus bridge according to the connection relationship between the acceleration card and the bus bridge and the southbound address of the acceleration card, so that the bus bridge forwards data between the southbound interfaces of different acceleration cards according to the routing information; Among them, the southbound interfaces of different acceleration cards located in the same computing device and the southbound interfaces of different acceleration cards located in different computing devices are interconnected through the bus bridge.

31. An electronic device, characterized in that, Comprising: A memory for storing a computer program; A processor for implementing the steps of the control method of the computing system as described in claim 29 when executing the computer program.

32. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the control method of the computing system as described in claim 29 when executed by a processor.

33. A computer program product, comprising a computer program, characterized in that, The computer program implements the steps of the control method of the computing system as described in claim 29 when executed by a processor.

Citation Information

Cited By

  • Switching unit, computing system and multi-node system

    CN121441865A