Methods, systems, and computing devices for emulating multi-card interconnect

By establishing associations between multiple peer-to-peer network ports and logical addresses between the graphics processor card and the switch, the limited capacity of existing simulation platforms is solved, enabling efficient multi-card interconnection simulation and saving hardware resources and costs.

CN121210232BActive Publication Date: 2026-03-31SHANGHAI BIREN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing simulation platforms have limited capacity when performing simulation verification of interconnects between graphics processor cards, resulting in low simulation verification efficiency.

Method used

By connecting multiple peer-to-peer network ports of a graphics processing unit (GPU) card to a switch and associating each peer-to-peer network port with multiple different logical addresses, the switch resolves the source and destination logical addresses of read/write requests and forwards the requests to the corresponding peer-to-peer network ports to simulate an interconnection scenario of multiple GPU cards.

Benefits of technology

It improves the efficiency of simulation verification, significantly saves hardware resources and costs, and can simulate the interconnection of multiple graphics processor cards simultaneously, thus significantly improving simulation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121210232B_ABST
    Figure CN121210232B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, system and computing device for simulating multi-card interconnection. The method comprises: connecting a plurality of peer-to-peer network ports of a graphics processor card to a switch respectively, and associating each of the plurality of peer-to-peer network ports with a plurality of different logical addresses for simulating ports of a plurality of graphics processor cards; parsing, by the switch, a received read-write request to obtain source logical address information and a destination logical address carried by the read-write request; and based on the obtained source logical address information and the destination logical address, forwarding the received read-write request to a peer-to-peer network port associated with the destination logical address for testing interconnection of the simulated graphics processor card. The technical solution of the present disclosure can simulate a scenario of interconnection of a plurality of graphics processor cards to improve the efficiency of simulation verification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this disclosure generally relate to the field of graphics processor card emulation technology, and more specifically to a method, system and computing device for emulating multi-card interconnection. Background Technology

[0002] Tasks such as training large models with hundreds of billions of parameters and processing large-scale datasets involve enormous computational and data volumes. For example, when simulating and verifying graphics processing unit (GPU) chips, such as constructing interconnects between GPU cards and using simulation platforms, current simulation platforms have limited capacity, resulting in low efficiency for simulation verification. Summary of the Invention

[0003] To address the aforementioned issues, this disclosure provides a method, system, and computing device for simulating multi-card interconnection, which can equivalently simulate a scenario where multiple graphics processor cards are interconnected, in order to perform simulation.

[0004] According to a first aspect of this disclosure, a method for simulating multi-GPU interconnection is provided, the method comprising: connecting a plurality of peer-to-peer (P2P) ports of a graphics processing unit (GPU) card to a switch, and associating each of the plurality of peer-to-peer ports with a plurality of different logical addresses to simulate ports of a plurality of GPU cards; the switch parsing a received read / write request to obtain source logical address information and destination logical address carried in the read / write request; and forwarding the received read / write request to the peer-to-peer port associated with the destination logical address based on the obtained source logical address information and destination logical address, for testing interconnection of the simulated GPU cards.

[0005] In some embodiments, the destination logical address indicates a peer network port identifier and a queuepair identifier, wherein the peer network port identifier is used to indicate the peer network port associated with the source of the read / write request, and the queuepair identifier is used to indicate the queuepair used for the data transmission of the read / write request.

[0006] In some embodiments, associating each of a plurality of peer network ports with a plurality of different logical addresses includes: associating each peer network port with a plurality of queue pairs, and configuring each queue pair with a different logical address to simulate different logical communication links, wherein each queue pair is associated with a queue pair identifier, and each queue pair is associated with other peer network ports besides the stated peer network port.

[0007] In some embodiments, associating each of a plurality of peer network ports with a plurality of different logical addresses includes: configuring each peer network port with an emulated graphics processor identifier, the emulated graphics processor identifier being associated with a unique Internet Protocol address (IP address) and Media Access Control bit address (MAC address); binding each peer network port with a plurality of queue pairs, and configuring each queue pair with a different Internet Protocol address and Media Access Control bit address.

[0008] In some embodiments, forwarding the received read / write request to the peer network port associated with the destination logical address based on the obtained source logical address information and destination logical address includes: the switch forwarding the received read / write request to the peer network port associated with the destination logical address according to pre-configured routing information based on the obtained source logical address information and destination logical address.

[0009] In some embodiments, the method further includes: simultaneously sending read / write requests to the switch based on at least two queue pairs among a plurality of queue pairs associated with the same peer-to-peer network port; and / or simultaneously sending read / write requests to the switch based on at least two peer-to-peer network ports among a plurality of peer-to-peer network ports.

[0010] According to a second aspect of this disclosure, a system for simulating multi-GPU interconnects is provided. The system includes: a graphics processing unit (GPU) card, a plurality of peer-to-peer network ports of the GPU card being connected to a switch, each of the plurality of peer-to-peer network ports being associated with a plurality of different logical addresses for simulating ports of the plurality of GPU cards; and a switch that parses received read / write requests to obtain source logical address information and destination logical address information carried in the read / write requests; and forwards the received read / write requests to the peer-to-peer network port associated with the destination logical address based on the obtained source logical address information and destination logical address information, for testing interconnects of the simulated GPU cards.

[0011] In some embodiments, the number of graphics processor cards is at least two, and each of the at least two graphics processor cards includes multiple peer-to-peer network ports.

[0012] According to a third aspect of this disclosure, a computing device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of the first aspect of this disclosure.

[0013] According to a fourth aspect of this disclosure, a computer-readable storage medium is provided that stores a computer program, which, when executed by a machine, performs the method of the first aspect of this disclosure.

[0014] According to a fifth aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a machine, performs the method of the first aspect of this disclosure.

[0015] According to the technical solution of this disclosure, multiple peer-to-peer network ports of a graphics processing unit (GPU) card are connected to a switch, and each of the multiple peer-to-peer network ports is associated with multiple different logical addresses to simulate the ports of multiple GPU cards. The switch parses the received read / write requests to obtain the source and destination logical address information carried in the read / write requests. Based on the obtained source and destination logical address information, the switch forwards the received read / write requests to the peer-to-peer network port associated with the destination logical address for testing the interconnection of the simulated GPU cards. The technical solution of this disclosure can equivalently simulate the scenario of interconnection of multiple GPU cards for simulation, thereby improving the efficiency of simulation verification.

[0016] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0017] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements.

[0018] Figure 1 A schematic diagram of a platform for interconnect emulation of graphics processor cards is shown.

[0019] Figure 2 A schematic diagram of a system for simulating a multi-card interconnect according to an embodiment of the present disclosure is shown.

[0020] Figure 3 A schematic diagram of a system for simulating a multi-card interconnect according to an embodiment of the present disclosure is shown.

[0021] Figure 4 A flowchart illustrating a method for simulating a multi-card interconnect according to an embodiment of this disclosure is shown.

[0022] Figure 5A schematic block diagram of an example electronic device is shown, illustrating a method for processing a target object that can be used to implement embodiments of the present disclosure. Detailed Implementation

[0023] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0024] The term "comprising" and its variations as used herein signify open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "one example embodiment" and "one embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0025] As described earlier, for example, when performing simulation verification for graphics processing unit (GPU) chips, such as constructing interconnects between GPU cards, a simulation platform is used for simulation verification. However, current simulation platforms have limited capacity, resulting in low efficiency for simulation verification.

[0026] Figure 1 A schematic diagram of a platform for interconnect emulation of graphics processor cards is shown. For example, refer to... Figure 1 For example, a platform for interconnecting and simulating graphics processor cards (GPUs) is constructed using a controller 114, a first graphics processor card (GPU) 101, a second GPU 103, and a switch 104. The first GPU 101 is interconnected with the switch 104 via its port 1, and the second GPU 103 is interconnected with the switch 104 via its port 2. For example, the controller 114 is communicatively connected to both the first GPU 101 and the second GPU 103. The controller 114 sends simulation commands to both GPUs 101 and 103. One of the GPUs 101 and 103 acts as a source GPU, and the other as a destination GPU. The source GPU sends read / write requests to the switch 104, with the destination GPU as the destination for read / write access.

[0027] It should be noted that Figure 1The platform shown can only build an interconnect between two graphics processing units (GPUs). Clearly, this platform is inefficient for simulation verification of GPUs, as it cannot simultaneously handle complete model parameters and the massive amounts of intermediate data generated during training. The required simulation verification cycle is lengthy, resulting in low efficiency.

[0028] In summary, the shortcomings of traditional solutions are that the required simulation verification cycle is lengthy and the efficiency of simulation verification is low.

[0029] To at least partially address one or more of the aforementioned problems and other potential issues, exemplary embodiments of this disclosure propose a technical solution related to a method, system, and computing device for simulating multi-GPU interconnects. In this technical solution, multiple peer-to-peer network ports of a graphics processing unit (GPU) card are connected to a switch, and each of these peer-to-peer network ports is associated with multiple different logical addresses to simulate the ports of multiple GPU cards. The switch parses received read / write requests to obtain the source and destination logical address information carried in the requests. Based on the obtained source and destination logical address information, the switch forwards the received read / write requests to the peer-to-peer network port associated with the destination logical address for testing the interconnection of the simulated GPU cards. This technical solution enables multiple peer-to-peer network ports of a GPU card to interconnect via a switch, effectively simulating a scenario of multiple GPU cards interconnected, thus improving the efficiency of simulation verification.

[0030] The following description, in conjunction with the accompanying drawings, describes the embodiments of this disclosure.

[0031] Figure 2 A schematic diagram of a system 100 for simulating a multi-card interconnect according to an embodiment of the present disclosure is shown. It should be understood that the system 100 may also include additional components not shown and / or the components shown may be omitted, and the scope of the present disclosure is not limited in this respect.

[0032] System 100 includes a graphics processing card 102 and a switch 104. Multiple peer-to-peer network ports of the graphics processing card 102 are connected to the switch 104, and each of the multiple peer-to-peer network ports is associated with multiple different logical addresses to emulate the ports of multiple graphics processing cards.

[0033] The switch 104 parses the received read / write requests to obtain the source and destination logical address information carried in the read / write requests; and based on the obtained source and destination logical address information, forwards the received read / write requests to the peer network port associated with the destination logical address for testing the interconnection of the simulated graphics processor card 102.

[0034] The graphics processor card 102 includes, for example, multiple peer-to-peer network ports such as peer-to-peer network port P1, peer-to-peer network port P2, ..., peer-to-peer network port Pn.

[0035] For example, multiple peer-to-peer network ports of graphics processor card 102 are connected to switch 104, and each of the multiple peer-to-peer network ports is associated with multiple different logical addresses to emulate the ports of multiple graphics processor cards 102.

[0036] The graphics processing card 102 can, for example, send emulation information via any one of its multiple peer-to-peer network ports (e.g., referred to as the "source peer-to-peer network port"), including read / write requests for the destination peer-to-peer network port. The destination peer-to-peer network port is, for example, a peer-to-peer network port other than the source peer-to-peer network port. That is, the graphics processing card 102 sends emulation information from the source peer-to-peer network port to the destination peer-to-peer network port to perform relevant read / write operations on the destination peer-to-peer network port for emulating ports of multiple graphics processing cards 102.

[0037] For example, associating each of multiple peer network ports with multiple different logical addresses includes: associating each peer network port with multiple queue pairs, and configuring each queue pair with a different logical address to simulate different logical communication links, wherein each queue pair is associated with a queue pair identifier, and each queue pair is associated with other peer network ports besides the stated peer network port.

[0038] That is, a single peer network port is associated with multiple queue pairs. For example, peer network port P1 is associated with multiple queue pairs. These multiple queue pairs associated with peer network port P1 are, for example, associated with multiple related peer network ports, which are, for example, other peer network ports (e.g., peer network port P2...peer network port Pn, etc.) among the multiple peer network ports possessed by the graphics processor card 102, besides the peer network port P1. For example, when peer network port P1 is the source peer network port, peer network ports P2...peer network port Pn, etc., can, for example, be the destination peer network ports. For example, peer network port P1 is associated with (n-1) queue pairs, which correspond to peer network ports P2...peer network port Pn, respectively. It should be understood that queue pairs can serve as dedicated channels for data transmission. For example, corresponding queue pairs can be used to enable the source peer network port (peer network port P1) to transmit data to the destination peer network port via the queue pair.

[0039] It should be understood that queue pairs are, for example, core communication units in remote direct memory access (RDMA) technology. As dedicated channels for data transmission, queue pairs can be associated with the address information of corresponding ports (e.g., the corresponding destination peer network ports). For example, different IP addresses and MAC addresses can be configured for different queue pairs to enable communication between different data links.

[0040] The destination logical address indicates the peer network port identifier (e.g., represented by P2P-port-ID) and the queue pair identifier (e.g., represented by QP-ID). The P2P-port-ID indicates the peer network port associated with the source of the read / write request, and the QP-ID indicates the queue pair used for the data transmission of the read / write request. It should be understood that the peer network port associated with the source of the read / write request is, for example, the "source peer network port" mentioned above, i.e., the peer network port that issued the read / write request (e.g., the read / write request message). The queue pair used to perform the data transmission for the read / write request can be determined through the QP-ID.

[0041] For example, associating each of multiple peer-to-peer network ports with multiple different logical addresses includes configuring each peer-to-peer network port with an emulated graphics processor identifier (e.g., represented by a GUI-ID), which is associated with a unique Internet Protocol address (IP address) and Media Access Control (MAC) address. Furthermore, binding multiple queue pairs to each peer-to-peer network port, and configuring each queue pair with a different Internet Protocol address (IP address) and Media Access Control (MAC) address.

[0042] It is worth noting that the unique Internet Protocol address (IP address) and Media Access Control (MAC) address associated with the emulated graphics processor identifier can be used to identify the address of the peer network port corresponding to the emulated graphics processor identifier within the network, serving as a logical identifier for the peer network port. Furthermore, the Internet Protocol address (IP address) and Media Access Control (MAC) address configured for each queue pair can, for example, characterize the address of the destination peer network port that the queue pair points to within the network. Therefore, addressing can be performed based on the Internet Protocol address (IP address) and Media Access Control (MAC) address configured for the queue pair to enable data transmission to the destination peer network port via the queue pair.

[0043] It should be understood that the peer network port identifier and queue pair identifier, indicated by the destination logical address of the request message, can collaboratively define a complete data path. Specifically, the peer network port identifier (P2P-port-ID) indicates the peer network port associated with the source of the read / write request. That is, based on the P2P-port-ID, the physical outgoing port of the data can be determined; specifically, the P2P-port-ID indicates which peer network port of the graphics processor card 102 the requested data (e.g., the read / write request message) should be sent to the switch from. The queue pair identifier (QP-ID) specifies the queue pair used for this data transmission. This queue pair is configured with a dedicated source and destination address, for example, during the initialization phase. The source address includes, for example, the source IP address and the source MAC address. The source address is associated with, for example, the physical identifier of the peer network port associated with the source of the read / write request, which includes, for example, the address of the peer network port (e.g., IP address and MAC address). The destination address, for example, precisely points to the logical access point of the destination, thus clearly defining the target location where the data needs to flow. It should be understood that this logical access point is, for example, the peer network port pointed to by the destination address (i.e., the destination peer network port). It should be understood that the queue pair identifier indicated by the destination logical address of the request message points, for example, to the destination peer network port, and the peer network port identifier indicated by the destination logical address of the request message points, for example, to the source peer network port, so that when the destination peer network port returns data in response to a read / write request, the source peer network port is clearly identified.

[0044] It is worth noting that in some embodiments, the graphics processor card 102 sends a read / write request to the switch 104 based on a queue pair associated with a peer-to-peer network port. It should be understood that this read / write request (i.e., the request message) carries source logical address information and a destination logical address. The source logical address information, for example, identifies the address of the source peer-to-peer network port in the network. The destination logical address indicates a peer-to-peer network port identifier and a queue pair identifier, wherein the peer-to-peer network port identifier indicates the peer-to-peer network port associated with the source of the read / write request, and the queue pair identifier indicates the queue pair used for the data transmission of the read / write request.

[0045] For example, if the peer-to-peer network port P1 of the graphics processor card 102 is used as the source peer-to-peer network port, then the peer-to-peer network ports P2...Pn of the graphics processor card 102 can all be used as destination peer-to-peer network ports. Correspondingly, if the peer-to-peer network port P2 of the graphics processor card 102 is used as the source peer-to-peer network port, then the peer-to-peer network ports P1...Pn of the graphics processor card 102 can all be used as destination peer-to-peer network ports. The interconnection between each source peer-to-peer network port and one of its corresponding destination peer-to-peer network ports can be used to test the interconnection of the simulated graphics processor cards. This interconnection between the source peer-to-peer network port and one of its corresponding destination peer-to-peer network ports can be considered as a pair of interconnections between the source graphics processor card and the destination graphics processor card, so as to test the interconnection simulation between the source and destination graphics processor cards. Therefore, based on system 100, various combinations of interconnections between source peer-to-peer network ports and their corresponding destination peer-to-peer network ports can be realized, thus enabling the simulation of multi-card interconnections.

[0046] In some embodiments, the graphics processor card 102 may simultaneously send read / write requests to the switch, for example, based on at least two of a plurality of queue pairs associated with the same peer-to-peer network port.

[0047] For example, the peer-to-peer network port P1 of the graphics processor card 102, as the source peer-to-peer network port, can simultaneously send read and write requests to the switch using at least two of the multiple queue pairs associated with it. That is, it can simultaneously perform interconnection simulation between at least two source graphics processor card pairs and destination graphics processor card pairs, which can significantly improve efficiency.

[0048] In some embodiments, read / write requests are simultaneously sent to the switch based on at least two of the multiple peer-to-peer network ports. For example, peer-to-peer port P1 of graphics processor card 102, acting as a source peer-to-peer port, simultaneously sends read / write requests to the switch using at least one of its associated queue pairs; simultaneously, peer-to-peer port P2 of graphics processor card 102, acting as a source peer-to-peer port, simultaneously sends read / write requests to the switch using at least one of its associated queue pairs. Therefore, interconnection simulation between at least two source and destination graphics processor card pairs can be performed simultaneously, significantly improving efficiency.

[0049] Then, at switch 104, based on the obtained source logical address information and destination logical address, the received read / write request is forwarded to the peer network port associated with the destination logical address. For example, the switch forwards the received read / write request to the peer network port associated with the destination logical address according to the pre-configured routing information based on the obtained source logical address information and destination logical address.

[0050] For example, switch 104 parses the received read / write request (e.g., read / write request message) to obtain the source and destination logical address information carried in the read / write request. Then, based on the obtained source and destination logical address information, switch 104 forwards the received read / write request to the peer network port associated with the destination logical address according to pre-configured routing information, thereby simulating, for example, the interconnection of two graphics processing cards.

[0051] As previously described, the multiple peer-to-peer network ports of the graphics processor card 102 can be configured to form various pairing combinations of source-end peer-to-peer network ports and destination-end peer-to-peer network ports. Each combination can simulate, for example, the interconnection of two graphics processor cards. The various pairing combinations can be used sequentially to test the interconnection of the simulated graphics processor cards, or at least two of the pairing combinations can be used concurrently to test the interconnection of the simulated graphics processor cards.

[0052] Therefore, the technical solutions of the embodiments of this disclosure can equivalently simulate a scenario where multiple graphics processing unit (GPU) cards are interconnected, thereby improving the efficiency of simulation verification. Furthermore, for example, a single GPU card can be used to equivalently simulate a scenario where multiple GPU cards are interconnected, significantly saving hardware resources and costs.

[0053] In some embodiments, based on system 100, on-network computing verification of the graphics processing card (GPU) can be performed, for example. GPU 102 sends a Load Data Request (LDR) to switch 104. The destination address of this request message contains an identification field `mcId` (multicast group identifier), which specifies the multicast group to which the read request belongs. Switch 104 can quickly resolve all destination peer network ports corresponding to the multicast group (i.e., which peer network ports to which the data read request needs to be initiated) using a pre-configured `mcId - port mapping table`. Switch 104 forwards the LDR request in a broadcast manner to all destination peer network ports within the corresponding multicast group based on the multicast information. For the destination peer network ports of GPU 102 that receive the request, they read the data according to the address and reply with data to switch 104. Switch 104 stores the received data and performs reduction calculations after collecting all the reduction data. Then, switch 104 sends the reduction result to the LDR requesting end.

[0054] Figure 3 A schematic diagram of a system 200 for simulating a multi-card interconnect according to an embodiment of the present disclosure is shown. It should be understood that the system 200 may also include additional components not shown and / or the components shown may be omitted, and the scope of the present disclosure is not limited in this respect.

[0055] System 200 includes at least two graphics processing units (GPUs) and a switch 104. The at least two GPUs include, for example, a first GPU 121 and a second GPU 122. The first GPU 121 includes, for example, multiple peer-to-peer network ports such as peer-to-peer network port P11, peer-to-peer network port P12, ..., peer-to-peer network port P1n. The second GPU 122 includes, for example, multiple peer-to-peer network ports such as peer-to-peer network port P21, peer-to-peer network port P22, ..., peer-to-peer network port P2n.

[0056] In some embodiments, the first graphics processor card 121 includes, for example, at least two dies, including, for example, a first die 1211 and a second die 1212. The first die 1211 includes multiple peer-to-peer network ports, and the second die 1212 includes multiple peer-to-peer network ports. In some embodiments, one of the first die 1211 and the second die 1212 is, for example, a master unit, and the other of the first die 1211 and the second die 1212 is, for example, a slave unit.

[0057] It is worth noting that, taking the peer network port P11 of the first graphics processor card 121 as an example, the multiple queue pairs associated with the peer network port P11 can point to the peer network ports P12...P1n of the first graphics processor card 121, and the peer network ports P21, P22...P2n of the second graphics processor card 122, that is, all other peer network ports other than the current peer network port.

[0058] When peer network port P11 is used as the source peer network port, all other peer network ports in the first graphics processor card 121 and the second graphics processor card 122, except for the source peer network port, can be used as the destination peer network ports.

[0059] Figure 4 A flowchart of a method 300 for simulating a multi-card interconnect according to an embodiment of the present disclosure is shown. Method 300 can be used, for example, in either system 100 or system 200. It should be understood that method 300 may also include other steps. Method 300 can be performed, for example, based on either system 100 or system 200, or at electronic device 500.

[0060] In step 302, the multiple peer-to-peer network ports of the graphics processor card are connected to the switch respectively, and each of the multiple peer-to-peer network ports is associated with multiple different logical addresses to emulate the ports of multiple graphics processor cards.

[0061] In step 304, the switch parses the received read / write request to obtain the source logical address information and the destination logical address carried in the read / write request.

[0062] In step 306, based on the obtained source logical address information and destination logical address, the received read / write request is forwarded to the peer network port associated with the destination logical address for testing the interconnection of the simulated graphics processor card.

[0063] The destination logical address indicates the peer network port identifier (e.g., represented by P2P-port-ID) and the queue pair identifier (e.g., represented by QP-ID). The P2P-port-ID indicates the peer network port associated with the source of the read / write request, and the QP-ID indicates the queue pair used for the data transmission of the read / write request. It should be understood that the peer network port associated with the source of the read / write request is, for example, the "source peer network port" mentioned above, i.e., the peer network port that issued the read / write request (e.g., the read / write request message). The queue pair used to perform the data transmission for the read / write request can be determined through the QP-ID.

[0064] For example, associating each of multiple peer-to-peer network ports with multiple different logical addresses includes configuring each peer-to-peer network port with an emulated graphics processor identifier (e.g., represented by a GUI-ID), which is associated with a unique Internet Protocol address (IP address) and Media Access Control (MAC) address. Furthermore, binding multiple queue pairs to each peer-to-peer network port, and configuring each queue pair with a different Internet Protocol address (IP address) and Media Access Control (MAC) address.

[0065] In some embodiments, associating each of a plurality of peer network ports with a plurality of different logical addresses includes: configuring each peer network port with an emulated graphics processor identifier associated with a unique Internet Protocol (IP) address and a Media Access Control (MAC) address; binding each peer network port to a plurality of queue pairs, and configuring each queue pair with a different IP address and a MAC address.

[0066] In some embodiments, forwarding the received read / write request to the peer network port associated with the destination logical address based on the obtained source logical address information and destination logical address includes: the switch forwards the received read / write request to the peer network port associated with the destination logical address according to pre-configured routing information based on the obtained source logical address information and destination logical address.

[0067] In some embodiments, the graphics processor card 102 may simultaneously send read / write requests to the switch, for example, based on at least two of a plurality of queue pairs associated with the same peer-to-peer network port.

[0068] For example, referring to system 100, the peer-to-peer network port P1 of graphics processor card 102 serves as the source peer-to-peer network port. It uses at least two of the multiple queue pairs associated with it to simultaneously send read and write requests to the switch. That is, it can simultaneously perform interconnection simulation between at least two source graphics processor card pairs and destination graphics processor card pairs, which can significantly improve efficiency.

[0069] In some embodiments, read / write requests are simultaneously sent to the switch based on at least two of a plurality of peer-to-peer network ports. For example, referring to system 100, peer-to-peer network port P1 of graphics processor card 102, acting as a source peer-to-peer network port, simultaneously sends read / write requests to the switch using at least one of its associated queue pairs; simultaneously, peer-to-peer network port P2 of graphics processor card 102, acting as a source peer-to-peer network port, simultaneously sends read / write requests to the switch using at least one of its associated queue pairs. Therefore, interconnection simulation between at least two source and destination graphics processor card pairs can be performed simultaneously, significantly improving efficiency.

[0070] Then, at switch 104, based on the obtained source logical address information and destination logical address, the received read / write request is forwarded to the peer network port associated with the destination logical address. For example, the switch forwards the received read / write request to the peer network port associated with the destination logical address according to the pre-configured routing information based on the obtained source logical address information and destination logical address.

[0071] For example, switch 104 parses the received read / write request (e.g., read / write request message) to obtain the source and destination logical address information carried in the read / write request. Then, based on the obtained source and destination logical address information, switch 104 forwards the received read / write request to the peer network port associated with the destination logical address according to pre-configured routing information, thereby simulating, for example, the interconnection of two graphics processing cards.

[0072] As previously described, referring to system 100, the multiple peer-to-peer network ports of graphics processor card 102 can be configured to form various source-end peer-to-peer network port pairing interconnections, each of which can simulate, for example, the interconnection of two graphics processor cards. The various pairing interconnections can be used sequentially to test the interconnection of the simulated graphics processor cards, or at least two pairing interconnections can be used concurrently to test the interconnection of the simulated graphics processor cards.

[0073] Therefore, the technical solutions of the embodiments of this disclosure can equivalently simulate a scenario where multiple graphics processing unit (GPU) cards are interconnected, thereby improving the efficiency of simulation verification. Furthermore, for example, a single GPU card can be used to equivalently simulate a scenario where multiple GPU cards are interconnected, significantly saving hardware resources and costs.

[0074] In some embodiments, based on system 100 and referring to method 300, on-network computing verification of the graphics processing card (GPU) can be performed, for example. GPU 102 sends a Load Data Request (LDR) to switch 104, for example. The destination address of this request message contains an identification field `mcId` (multicast group identifier), which specifies the multicast group to which the read request belongs. Switch 104 can quickly resolve all destination peer network ports corresponding to the multicast group (i.e., which peer network ports to which the data read request needs to be initiated) using a pre-configured `mcId - port mapping table`. Switch 104 forwards the LDR request in a broadcast manner to all destination peer network ports within the corresponding multicast group based on the multicast information. For the destination peer network ports of GPU 102 that receive the request, they read the data according to the address and reply with data to switch 104. Switch 104 stores the received data and performs reduction calculations after collecting all the reduction data. Then, switch 104 sends the reduction result to the LDR requesting end.

[0075] It is worth noting that the technical solutions of the embodiments disclosed herein can be applied, for example, to the pre-silicon verification field of high-performance computing chips for graphics processing units (GPUs), and are deeply adapted to related simulation platforms such as EMUs (Intelligent Management Units) and FPGAs (Field-Programmable Gate Arrays). For example, using a high-performance switch as the core interconnection hub, the interconnection topology and collaborative logic of multiple graphics processing unit (GPU) cards can be simulated. This solution can not only generate multi-card interconnection scenarios several times the actual carrying capacity of the platform through flexible topology virtualization technology, but also accurately complete the comprehensive verification of core capabilities such as multi-card collaborative scheduling and data transmission, achieving comprehensive verification of core capabilities such as multi-card collaborative scheduling, data transmission, and on-network computing.

[0076] Figure 5 A schematic block diagram of an example electronic device 500 for processing a target object, which can be used to implement embodiments of the present disclosure, is shown. As shown, the electronic device 500 includes a central processing unit (i.e., CPU 501), which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (i.e., ROM 502) or loaded from storage unit 508 into random access memory (i.e., RAM 503). Various programs and data required for the operation of the electronic device 500 may also be stored in RAM 503. The CPU 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output interfaces (i.e., I / O interfaces 505) are also connected to bus 504.

[0077] Multiple components in electronic device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, microphone, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows electronic device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0078] The various processes and handling described above, such as method 300, can be executed by CPU 501. For example, in some embodiments, method 300 can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by CPU 501, one or more actions of method 300 described above can be performed.

[0079] This disclosure relates to methods, apparatus, systems, electronic devices, computer-readable storage media, and / or computer program products. A computer program product may include computer-readable program instructions for performing various aspects of this disclosure.

[0080] In some embodiments, the method 300 described above can be implemented as a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of this disclosure.

[0081] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0082] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge computing devices. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to computer-readable storage media within the respective computing / processing device.

[0083] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0084] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0085] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0086] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0087] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0088] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

[0089] The above are merely optional embodiments of this disclosure and are not intended to limit this disclosure. Various modifications and variations can be made to this disclosure by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for emulating multi-card interconnection, characterized by, The method comprises: connecting a plurality of peer-to-peer network ports of a graphics processor card to a switch respectively, and associating each of the plurality of peer-to-peer network ports with a plurality of different logical addresses for emulating ports of a plurality of graphics processor cards; the switch parses the received read-write request to obtain source logical address information and a destination logical address carried by the read-write request; and based on the obtained source logical address information and the destination logical address, forwarding the received read-write request to the peer-to-peer network port associated with the destination logical address for testing the interconnection of the emulated graphics processor card; the destination logical address indicates a peer-to-peer network port identifier.

2. The method of claim 1, wherein, The destination logical address also indicates a queue pair identifier, the peer-to-peer network port identifier is used to indicate a peer-to-peer network port associated with a source of the read-write request, and the queue pair identifier is used to indicate a queue pair used for data transmission of the read-write request.

3. The method of claim 1, wherein, Associating each of the plurality of peer-to-peer network ports with a plurality of different logical addresses comprises: associating each peer-to-peer network port with a plurality of queue pairs, and configuring each queue pair with a different logical address to simulate a different logical communication link, each queue pair in the plurality of queue pairs is associated with a queue pair identifier, and each queue pair in the plurality of queue pairs is associated with other peer-to-peer network ports except the each peer-to-peer network port respectively.

4. The method of claim 1, wherein, Associating each of the plurality of peer-to-peer network ports with a plurality of different logical addresses comprises: configuring each peer-to-peer network port with an emulated graphics processor identifier, the emulated graphics processor identifier is associated with a unique internet protocol address and a media access control bit address; binding a plurality of queue pairs for each peer-to-peer network port, and configuring each queue pair with a different internet protocol address and a media access control bit address.

5. The method of claim 1, wherein, Based on the obtained source logical address information and the destination logical address, forwarding the received read-write request to the peer-to-peer network port associated with the destination logical address comprises: the switch forwards the received read-write request to the peer-to-peer network port associated with the destination logical address according to preconfigured routing information based on the obtained source logical address information and the destination logical address.

6. The method of claim 3, wherein, Further comprising: sending read-write requests to the switch based on at least two queue pairs in a plurality of queue pairs associated with the same peer-to-peer network port simultaneously; and / or sending read-write requests to the switch based on at least two peer-to-peer network ports in the plurality of peer-to-peer network ports simultaneously.

7. A system for emulating multi-card interconnect, the system comprising: Comprise: a graphics processor card, a plurality of peer-to-peer network ports of the graphics processor card are connected to a switch respectively, each of the plurality of peer-to-peer network ports is associated with a plurality of different logical addresses for emulating ports of a plurality of graphics processor cards; and The switch is configured to, for a received read or write request, parse the read or write request to obtain source logical address information and destination logical address information carried by the read or write request, and forward the received read or write request to a peer network port associated with the destination logical address based on the obtained source logical address information and destination logical address, for testing an interconnect of the emulated graphics processor card.

8. The system of claim 7, wherein, The number of graphics processor cards is at least two, and each of the at least two graphics processor cards comprises a plurality of peer network ports.

9. A computing device, comprising: Comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer readable storage medium, the computer program being executed by a machine to perform the method of any one of claims 1-6.

11. A computer program product, characterised in that, A computer program is stored on the computer readable storage medium, the computer program being executed by a machine to perform the method of any one of claims 1-6.

Citation Information

Patent Citations

  • Multi-card cascaded simulation unit model and communication method thereof

    CN119179664A

  • Host internal network delay diagnosis method based on loopback test

    CN119788571A