Methods, systems, and computing devices for emulating multi-card interconnects

By connecting the peer-to-peer network ports of graphics processor cards to a switch and resolving and forwarding logical addresses, the interconnection scenario of multiple graphics processor cards is simulated, which solves the problem of limited capacity of existing simulation platforms, improves simulation verification efficiency, and saves hardware resources.

CN121210232AActive Publication Date: 2025-12-26SHANGHAI BIREN TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511786649.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2025-12-26
Estimated Expiration
2045-11-28

AI Technical Summary

Technical Problem

Existing simulation platforms have limited capacity when performing simulation verification of interconnects between graphics processor cards, resulting in low simulation verification efficiency.

Method used

By connecting multiple peer-to-peer network ports of a graphics processing unit (GPU) card to a switch and associating each peer-to-peer network port with multiple different logical addresses, the switch resolves the source and destination logical addresses of read/write requests and forwards the requests to the corresponding peer-to-peer network ports to simulate an interconnection scenario of multiple GPU cards.

Benefits of technology

It improves the efficiency of simulation verification, significantly saves hardware resources and costs, and can simulate the interconnection of multiple graphics processor cards simultaneously, thus significantly improving simulation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121210232A_ABST
    Figure CN121210232A_ABST
Patent Text Reader

Abstract

The invention relates to a method, a system and a computing device for emulating multi-card interconnection. The method comprises the following steps: respectively connecting a plurality of peer-to-peer network ports of the graphics processor cards with a switch, and associating each peer-to-peer network port in the plurality of peer-to-peer network ports with a plurality of different logic addresses so as to simulate ports of the plurality of graphics processor cards; the switch analyzes the received read-write request so as to obtain source end logic address information and a destination end logic address carried by the read-write request; and forwarding the received read-write request to a peer-to-peer network port associated with the destination end logic address based on the obtained source end logic address information and the destination end logic address so as to test interconnection about the simulated graphics processor card. According to the technical scheme provided by the invention, a scene in which a plurality of graphics processor cards are interconnected can be equivalently simulated, so that simulation is facilitated, and the efficiency of simulation verification can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present disclosure generally relate to the field of graphics processor card emulation, and more particularly to a method, system and computing device for emulating multi-card interconnection. BACKGROUND

[0002] For tasks such as training of a trillion-level parameter large model and processing of a large-scale data set, there are huge amounts of computation and data. For example, when performing simulation verification on a graphics processor chip, for example, interconnection between graphics processor cards is constructed, and a simulation platform is used to perform simulation verification. For example, the current simulation platform has limited capacity, and the efficiency of simulation verification is low. SUMMARY

[0003] To solve the above problems, the present disclosure provides a method, system and computing device for emulating multi-card interconnection, which can equivalently simulate a scenario of interconnection between multiple graphics processor cards for simulation.

[0004] According to a first aspect of the present disclosure, a method for emulating multi-card interconnection is provided, the method comprising: connecting a plurality of peer to peer (P2P) ports of a graphics processor card to a switch respectively, and associating each of the plurality of P2P ports with a plurality of different logical addresses for emulating ports of a plurality of graphics processor cards; parsing, by the switch, a read-write request received to obtain source logical address information and a destination logical address carried by the read-write request; and based on the obtained source logical address information and the destination logical address, forwarding the received read-write request to a P2P port associated with the destination logical address for testing interconnection of the emulated graphics processor card.

[0005] In some embodiments, the destination logical address indicates a P2P port identifier and a queue pair identifier, the P2P port identifier being used to indicate a P2P port associated with a source of the read-write request, and the queue pair identifier being used to indicate a queue pair used for data transmission of the read-write request.

[0006] In some embodiments, associating each of the plurality of P2P ports with a plurality of different logical addresses comprises: associating each P2P port with a plurality of queue pairs, and configuring each queue pair with a different logical address to simulate a different logical communication link, each queue pair of the plurality of queue pairs being associated with a queue pair identifier, and each queue pair of the plurality of queue pairs being associated with other P2P ports than the each P2P port respectively.

[0007] In some embodiments, associating each of the plurality of peer-to-peer network ports with a plurality of different logical addresses comprises: configuring each peer-to-peer network port with a simulated graphics processor identification associated with a unique Internet Protocol (IP) address and a Media Access Control (MAC) address; and binding a plurality of queue pairs to each peer-to-peer network port, and configuring each queue pair with a different Internet Protocol (IP) address and a Media Access Control (MAC) address.

[0008] In some embodiments, forwarding the received read-write request to the peer-to-peer network port associated with the destination logical address based on the obtained source logical address information and the destination logical address comprises: switching the received read-write request to the peer-to-peer network port associated with the destination logical address based on the obtained source logical address information and the destination logical address according to pre-configured routing information.

[0009] In some embodiments, the method further comprises: sending read-write requests to the switch simultaneously based on at least two of the plurality of queue pairs associated with the same peer-to-peer network port; and / or sending read-write requests to the switch simultaneously based on at least two of the plurality of peer-to-peer network ports.

[0010] According to a second aspect of the present disclosure, there is provided a system for emulating multi-card interconnect, the system comprising: a graphics processor card, a plurality of peer-to-peer network ports of the graphics processor card being connected to a switch respectively, each of the plurality of peer-to-peer network ports being associated with a plurality of different logical addresses for emulating ports of a plurality of graphics processor cards; and the switch, for parsing a received read-write request to obtain source logical address information and a destination logical address carried by the read-write request; and forwarding the received read-write request to a peer-to-peer network port associated with the destination logical address based on the obtained source logical address information and the destination logical address for testing interconnects with respect to the emulated graphics processor card.

[0011] In some embodiments, the number of graphics processor cards is at least two, each of the at least two graphics processor cards comprising a plurality of peer-to-peer network ports.

[0012] According to a third aspect of the present disclosure, there is provided a computing device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of the first aspect of the present disclosure.

[0013] According to a fourth aspect of the present disclosure, a computer readable storage medium is provided, having stored thereon a computer program which, when executed by a machine, performs the method of the first aspect of the present disclosure.

[0014] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a machine, performs the method of the first aspect of the present disclosure.

[0015] According to the technical solution of the present disclosure, a plurality of peer-to-peer network ports of a graphics processor card are connected to a switch respectively, and each of the plurality of peer-to-peer network ports is associated with a plurality of different logical addresses for simulating the ports of a plurality of graphics processor cards; the switch parses a received read-write request to obtain source logical address information and a destination logical address carried by the read-write request; and based on the obtained source logical address information and the destination logical address, the received read-write request is forwarded to the peer-to-peer network port associated with the destination logical address for testing the interconnection of the simulated graphics processor card. The technical solution of the present disclosure can simulate the scenario of interconnection of a plurality of graphics processor cards to perform simulation, and can improve the efficiency of simulation verification.

[0016] It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0017] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail the following embodiments with reference to the attached drawings. In the drawings, the same or similar reference numerals refer to the same or similar elements.

[0018] Figure 1 A schematic diagram of a platform for interconnection simulation for graphics processor cards is shown.

[0019] Figure 2 A schematic diagram of a system for simulating multi-card interconnection according to an embodiment of the present disclosure is shown.

[0020] Figure 3 A schematic diagram of a system for simulating multi-card interconnection according to an embodiment of the present disclosure is shown.

[0021] Figure 4 A flowchart of a method for simulating multi-card interconnection according to an embodiment of the present disclosure is shown.

[0022] Figure 5A schematic block diagram of an example electronic device is shown, illustrating a method for processing a target object that can be used to implement embodiments of the present disclosure. Detailed Implementation

[0023] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0024] The term "comprising" and its variations as used herein signify open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "one example embodiment" and "one embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0025] As described earlier, for example, when performing simulation verification for graphics processing unit (GPU) chips, such as constructing interconnects between GPU cards, a simulation platform is used for simulation verification. However, current simulation platforms have limited capacity, resulting in low efficiency for simulation verification.

[0026] Figure 1 A schematic diagram of a platform for interconnect emulation of graphics processor cards is shown. For example, refer to... Figure 1 For example, a platform for interconnecting and simulating graphics processor cards (GPUs) is constructed using a controller 114, a first graphics processor card (GPU) 101, a second GPU 103, and a switch 104. The first GPU 101 is interconnected with the switch 104 via its port 1, and the second GPU 103 is interconnected with the switch 104 via its port 2. For example, the controller 114 is communicatively connected to both the first GPU 101 and the second GPU 103. The controller 114 sends simulation commands to both GPUs 101 and 103. One of the GPUs 101 and 103 acts as a source GPU, and the other as a destination GPU. The source GPU sends read / write requests to the switch 104, with the destination GPU as the destination for read / write access.

[0027] It should be noted that Figure 1In the shown platform, only the interconnection between two graphic processor cards can be constructed. It is obvious that the platform cannot simultaneously carry the complete model parameters and the massive intermediate data generated in the training process during the simulation verification for the graphic processor card, and the simulation verification period is long and the simulation verification efficiency is low.

[0028] In summary, the traditional scheme has the problem of long simulation verification period and low simulation verification efficiency.

[0029] To at least partially solve one or more of the above problems and other potential problems, example embodiments of the present disclosure propose a technical scheme related to a method, system and computing device for simulating multi-card interconnection. In the technical scheme of the present disclosure, a plurality of peer-to-peer network ports of a graphic processor card are respectively connected to a switch, and each of the plurality of peer-to-peer network ports is associated with a plurality of different logical addresses for simulating the ports of a plurality of graphic processor cards; the switch parses a received read-write request to obtain source logical address information and a destination logical address carried by the read-write request; and based on the obtained source logical address information and the destination logical address, the received read-write request is forwarded to the peer-to-peer network port associated with the destination logical address for testing the interconnection of the simulated graphic processor card. The technical scheme of the present disclosure can form an interconnection between a plurality of peer-to-peer network ports of a graphic processor card via a switch, simulate a scenario of interconnection of a plurality of graphic processor cards, and improve the efficiency of simulation verification.

[0030] The technical scheme of the embodiments of the present disclosure will be described below with reference to the accompanying drawings.

[0031] Figure 2 A schematic diagram of a system 100 for simulating multi-card interconnection according to an embodiment of the present disclosure is shown. It should be understood that the system 100 can also include additional components not shown and / or can omit the components shown, and the scope of the present disclosure is not limited in this regard.

[0032] The system 100 includes a graphic processor card 102 and a switch 104. A plurality of peer-to-peer network ports of the graphic processor card 102 are respectively connected to the switch 104, and each of the plurality of peer-to-peer network ports is associated with a plurality of different logical addresses for simulating the ports of a plurality of graphic processor cards.

[0033] The switch 104 parses the received read / write request to obtain source logical address information and destination logical address carried by the read / write request, and forwards the received read / write request to a peer network port associated with the destination logical address based on the obtained source logical address information and destination logical address, for testing the interconnection of the emulated graphics processor card 102.

[0034] The graphics processor card 102 includes a plurality of peer network ports, such as a peer network port P1, a peer network port P2, a peer network port Pn, and the like.

[0035] For example, the plurality of peer network ports of the graphics processor card 102 are respectively connected to the switch 104, and each of the plurality of peer network ports is associated with a plurality of different logical addresses, for emulating the ports of the plurality of graphics processor cards 102.

[0036] The graphics processor card 102 may, for example, send information about emulation via any one of the plurality of peer network ports thereof (for example, referred to as a “source peer network port”), wherein the information about emulation includes, for example, a read / write request about a destination peer network port. The destination peer network port is, for example, other than the source peer network port. That is, the graphics processor card 102 sends the information about emulation to the destination peer network port via the source peer network port, so as to perform relevant read / write operations on the destination peer network port, for emulating the ports of the plurality of graphics processor cards 102.

[0037] For example, associating each of the plurality of peer network ports with a plurality of different logical addresses includes: associating each of the plurality of peer network ports with a plurality of queue pairs, and configuring each of the queue pairs with a different logical address, for simulating different logical communication links, wherein each of the plurality of queue pairs is associated with a queue pair identifier, and each of the plurality of queue pairs is respectively associated with other peer network ports than the each of the plurality of peer network ports.

[0038] That is, one peer-to-peer network port is associated with multiple queue pair, and the multiple queue pairs associated with the peer-to-peer network port. For example, the peer-to-peer network port P1 is associated with multiple queue pairs. The multiple queue pairs associated with the peer-to-peer network port P1 are associated with multiple related peer-to-peer network ports, for example, the peer-to-peer network ports P2, P3, …, Pn, etc. other than the peer-to-peer network port (i.e. the peer-to-peer network port P1) in the multiple peer-to-peer network ports of the graphics processor card 102. For example, the peer-to-peer network port P1 as a source peer-to-peer network port, the peer-to-peer network ports P2, P3, …, Pn, etc. can be destination peer-to-peer network ports. For example, the peer-to-peer network port P1 is associated with (n-1) queue pairs, which correspond to the peer-to-peer network ports P2, P3, …, Pn, etc. respectively. It should be understood that the queue pair can be a dedicated channel for data transmission. For example, the source peer-to-peer network port (the peer-to-peer network port P1) can transmit data to the destination peer-to-peer network port via the corresponding queue pair.

[0039] It should be understood that the queue pair is, for example, a core communication unit in the remote direct memory access (RDMA) technology. The queue pair as a dedicated channel for data transmission can associate the address information of the corresponding port (e.g. the corresponding destination peer-to-peer network port). For example, different IP addresses and MAC addresses can be configured for different queue pairs to realize communication between different data links.

[0040] Wherein, the destination logical address indicates the peer-to-peer network port identification (e.g. represented by P2P-port-ID) and the queue pair identification (e.g. represented by QP-ID). Wherein, the peer-to-peer network port identification P2P-port-ID is used to indicate the peer-to-peer network port associated with the source of the read-write request, and the queue pair identification QP-ID is used to indicate the queue pair used for data transmission of the read-write request. It should be understood that the peer-to-peer network port associated with the source of the read-write request is, for example, the "source peer-to-peer network port" described above, i.e. the peer-to-peer network port that issues the read-write request (e.g. the read-write request message). Through the queue pair identification QP-ID, the queue pair used for data transmission of the read-write request can be determined.

[0041] For example, causing each of the plurality of peer-to-peer network ports to be associated with a plurality of different logical addresses comprises configuring each peer-to-peer network port with an emulated graphics processor identification (e.g. characterized by a GUI-ID) that is associated with a unique Internet Protocol address (IP address) and Media Access Control bit address (MAC address). Also, causing each peer-to-peer network port to be bound to a plurality of queue pairs, and causing each queue pair to be configured with a different Internet Protocol address (IP address) and Media Access Control bit address (MAC address).

[0042] It is worth noting that the unique Internet Protocol address (IP address) and Media Access Control bit address (MAC address) associated with the emulated graphics processor identification can be used to identify the address of the peer-to-peer network port corresponding to the emulated graphics processor identification in the network, which is a logical identification of the peer-to-peer network port. Also, the Internet Protocol address (IP address) and Media Access Control bit address (MAC address) configured for each queue pair can be used to characterize the address of the destination peer-to-peer network port pointed to by the queue pair in the network. Thus, the destination peer-to-peer network port can be addressed based on the Internet Protocol address (IP address) and Media Access Control bit address (MAC address) configured for the queue pair in order to enable transmission of data to the destination peer-to-peer network port via the queue pair.

[0043] It should be appreciated that based on the peer-to-peer port identification and the queue pair identification indicated by the destination logical address of the request message, the peer-to-peer port identification and the queue pair identification can cooperatively define a complete data path. Among them, the peer-to-peer port identification P2P-port-ID is used to indicate the peer-to-peer port associated with the source of the read-write request, that is, based on the peer-to-peer port identification P2P-port-ID, the physical outflow port of the data can be determined, that is, the peer-to-peer port identification P2P-port-ID indicates which peer-to-peer port of the graphics processor card 102 the data (for example, the read-write request message) of the request needs to be sent to the switch. The queue pair identification QP-ID is used to specify the queue pair used for this data transmission. Among them, the queue pair is configured with a dedicated source address and a destination address in the initialization stage. The source address, for example, includes a source IP address and a source MAC address. The source address is associated with the physical identification of the peer-to-peer port associated with the source of the read-write request, which includes the address (for example, the IP address and the MAC address) of the peer-to-peer port. The destination address is accurately directed to the logical access point of the destination, so as to clearly indicate the target position where the data needs to flow in. It should be appreciated that the logical access point is, for example, the peer-to-peer port (i.e., the destination peer-to-peer port) pointed to by the destination address. It should be appreciated that the queue pair identification indicated by the destination logical address of the request message is, for example, directed to the destination peer-to-peer port, and the peer-to-peer port identification indicated by the destination logical address of the request message is, for example, directed to the source peer-to-peer port, so that when the destination peer-to-peer port returns the data in response to the read-write request, the source peer-to-peer port is clear.

[0044] It should be noted that in some embodiments, the graphics processor card 102 sends a read-write request to the switch 104 based on the queue pair associated with the peer-to-peer port. It should be appreciated that the read-write request (i.e., the request message) carries source logical address information and destination logical address. The source logical address information is used to identify the address of the source peer-to-peer port in the network, for example. The destination logical address indicates the peer-to-peer port identification and the queue pair identification, the peer-to-peer port identification is used to indicate the peer-to-peer port associated with the source of the read-write request, and the queue pair identification is used to indicate the queue pair used for data transmission of the read-write request.

[0045] For example, peer-to-peer network port PI of graphics processor card 102 as a source peer-to-peer network port, peer-to-peer network port P2... peer-to-peer network port Pn of graphics processor card 102 can all be destination peer-to-peer network ports. Correspondingly, peer-to-peer network port P2 of graphics processor card 102 as a source peer-to-peer network port, peer-to-peer network port PI... peer-to-peer network port Pn of graphics processor card 102 can all be destination peer-to-peer network ports. The interconnection between each source peer-to-peer network port and a corresponding destination peer-to-peer network port can be used to test the interconnection of the emulated graphics processor card. The interconnection between the source peer-to-peer network port and the corresponding destination peer-to-peer network port can be regarded as the interconnection of a pair of source graphics processor card and destination graphics processor card, so as to be used to test the interconnection emulation of the pair of source graphics processor card and destination graphics processor card. Thus, based on system 100, various combinations of interconnections between source peer-to-peer network ports and corresponding destination peer-to-peer network ports can be implemented, and thus can be used to emulate multi-card interconnections.

[0046] In some embodiments, graphics processor card 102 sends read-write requests to the switch based on at least two queue pairs associated with the same peer-to-peer network port, for example.

[0047] For example, peer-to-peer network port PI of graphics processor card 102 as a source peer-to-peer network port, peer-to-peer network port P2... peer-to-peer network port Pn of graphics processor card 102 can all be destination peer-to-peer network ports. Correspondingly, peer-to-peer network port P2 of graphics processor card 102 as a source peer-to-peer network port, peer-to-peer network port PI... peer-to-peer network port Pn of graphics processor card 102 can all be destination peer-to-peer network ports. The interconnection between each source peer-to-peer network port and a corresponding destination peer-to-peer network port can be used to test the interconnection of the emulated graphics processor card. The interconnection between the source peer-to-peer network port and the corresponding destination peer-to-peer network port can be regarded as the interconnection of a pair of source graphics processor card and destination graphics processor card, so as to be used to test the interconnection emulation of the pair of source graphics processor card and destination graphics processor card. Thus, based on system 100, various combinations of interconnections between source peer-to-peer network ports and corresponding destination peer-to-peer network ports can be implemented, and thus can be used to emulate multi-card interconnections.

[0048] In some embodiments, graphics processor card 102 sends read-write requests to the switch based on at least two queue pairs associated with the same peer-to-peer network port, for example.

[0049] Then, at the switch 104, based on the obtained source logical address information and the destination logical address, the received read-write request is forwarded to the peer-to-peer network port associated with the destination logical address. For example, the switch forwards the received read-write request to the peer-to-peer network port associated with the destination logical address according to pre-configured routing information based on the obtained source logical address information and the destination logical address.

[0050] For example, the switch 104 parses the received read-write request (e.g., read-write request packet) to obtain the source logical address information and the destination logical address carried in the read-write request. Then, the switch 104 forwards the received read-write request to the peer-to-peer network port associated with the destination logical address according to pre-configured routing information based on the obtained source logical address information and the destination logical address, thereby emulating the interconnection of two graphics processor cards, for example.

[0051] As described above, the plurality of peer-to-peer network ports of the graphics processor card 102 can be connected to form a plurality of paired combination interconnection modes of source peer-to-peer network ports and destination peer-to-peer network ports, each of which can emulate the interconnection of two graphics processor cards, for example. The plurality of paired combination interconnection modes formed can be used in turn for testing the interconnection of the emulated graphics processor card, or at least two of the paired combination interconnection modes can be used concurrently for testing the interconnection of the emulated graphics processor card.

[0052] Therefore, the technical solution of the embodiments of the present disclosure can simulate the interconnection of multiple graphics processor cards to perform emulation, thereby improving the efficiency of emulation verification. In addition, for example, one graphics processor card can be used to simulate the interconnection of multiple graphics processor cards to perform emulation, thereby significantly saving hardware resources and costs.

[0053] In some embodiments, based on the system 100, in-network computing verification with respect to a graphics processor card can be performed. The graphics processor card 102, for example, sends a load data request (LDR) to the switch 104. The request message contains an identification field mcld (multicast group identification) in the destination address, which is used to explicitly indicate the multicast group to which the current read request belongs. The switch 104 can quickly resolve all destination peer-to-peer network ports corresponding to the multicast group (i.e., which peer-to-peer network ports to initiate a data read request to) through a pre-configured “mcld - port mapping table”. The switch 104 forwards the LDR request to all destination peer-to-peer network ports in the corresponding multicast group in a broadcast manner according to the multicast information. For the destination peer-to-peer network port of the graphics processor card 102 that receives the request, it reads data according to the address and replies the data to the switch 104. The switch 104 stores the received data, and when all the reduction data is collected, performs reduction calculation. Then, the switch 104 sends the reduction result to the LDR request end.

[0054] Figure 3 A schematic diagram of a system 200 for emulating multi-card interconnection according to embodiments of the present disclosure is shown. It should be understood that the system 200 can also include additional components not shown and / or can omit components shown, without limitation in this regard.

[0055] The system 200 includes at least two graphics processor cards, a switch 104. The at least two graphics processor cards include, for example, a first graphics processor card 121 and a second graphics processor card 122. The first graphics processor card 121 includes, for example, a plurality of peer-to-peer network ports, such as a peer-to-peer network port P11, a peer-to-peer network port P12, …, and a peer-to-peer network port P1n. The second graphics processor card 122 includes, for example, a plurality of peer-to-peer network ports, such as a peer-to-peer network port P21, a peer-to-peer network port P22, …, and a peer-to-peer network port P2n.

[0056] In some embodiments, the first graphics processor card 121 includes, for example, at least two dies, including a first die 1211 and a second die 1212. The first die 1211 includes a plurality of peer-to-peer network ports, and the second die 1212 includes a plurality of peer-to-peer network ports. In some embodiments, one of the first die 1211 and the second die 1212 is a master, and the other one of the first die 1211 and the second die 1212 is a slave.

[0057] It is worth mentioning that, taking the peer-to-peer network port P11 of the first graphics processor card 121 as an example, the plurality of queue pairs associated with the peer-to-peer network port P11 can be directed to the peer-to-peer network ports P12, …, P1n of the first graphics processor card 121 and the peer-to-peer network ports P21, P22, …, P2n of the second graphics processor card 122, i.e., each peer-to-peer network port other than the current peer-to-peer network port.

[0058] When the peer-to-peer network port P11 is a source peer-to-peer network port, the other peer-to-peer network ports in the first graphics processor card 121 and the second graphics processor card 122 than the source peer-to-peer network port can be destination peer-to-peer network ports.

[0059] Figure 4 A flow chart of a method 300 for emulating multi-card interconnection according to an embodiment of the present disclosure is shown. The method 300 can be used to control any one of the system 100, the system 200, for example. It should be understood that the method 300 can also include other steps. The method 300 can be executed based on any one of the system 100, the system 200, for example, or can be executed at the electronic device 500.

[0060] At step 302, a plurality of peer-to-peer network ports of a graphics processor card are connected to a switch respectively, and each of the plurality of peer-to-peer network ports is associated with a plurality of different logical addresses for emulating ports of a plurality of graphics processor cards.

[0061] At step 304, the switch parses a received read-write request to obtain source logical address information and a destination logical address carried by the read-write request.

[0062] At step 306, based on the obtained source logical address information and the destination logical address, the received read-write request is forwarded to a peer-to-peer network port associated with the destination logical address for testing interconnection of the emulated graphics processor card.

[0063] Wherein, the destination logical address indicates a peer-to-peer network port identification (e.g., represented by P2P-port-ID) and a queue pair identification (e.g., represented by QP-ID). Wherein, the peer-to-peer network port identification P2P-port-ID is used to indicate a peer-to-peer network port associated with a source of the read-write request, and the queue pair identification QP-ID is used to indicate a queue pair used for data transmission of the read-write request. It should be understood that the peer-to-peer network port associated with the source of the read-write request is, for example, the "source peer-to-peer network port" described above, i.e., the peer-to-peer network port that issues the read-write request (e.g., read-write request message). Through the queue pair identification QP-ID, the queue pair used for data transmission of the read-write request can be determined.

[0064] For example, associating each of the plurality of peer-to-peer network ports with a plurality of different logical addresses includes configuring each peer-to-peer network port with an emulated graphics processor identification (e.g., characterized by a GUI-ID) associated with a unique Internet Protocol address (IP address) and Media Access Control bit address (MAC address). Also, binding a plurality of queue pairs to each peer-to-peer network port, and configuring each queue pair with a different Internet Protocol address (IP address) and Media Access Control bit address (MAC address).

[0065] In some embodiments, associating each of the plurality of peer-to-peer network ports with a plurality of different logical addresses includes configuring each peer-to-peer network port with an emulated graphics processor identification associated with a unique Internet Protocol address and Media Access Control bit address; binding a plurality of queue pairs to each peer-to-peer network port, and configuring each queue pair with a different Internet Protocol address and Media Access Control bit address.

[0066] In some embodiments, forwarding the received read-write request to the peer-to-peer network port associated with the destination logical address based on the obtained source logical address information and the destination logical address includes: forwarding, by the switch, the received read-write request to the peer-to-peer network port associated with the destination logical address according to pre-configured routing information based on the obtained source logical address information and the destination logical address.

[0067] In some embodiments, the graphics processor card 102 sends the read-write request to the switch based on at least two of the plurality of queue pairs associated with the same peer-to-peer network port, for example.

[0068] For example, referring to the system 100, the peer-to-peer network port P1 of the graphics processor card 102, as a source peer-to-peer network port, sends the read-write request to the switch based on at least two of the plurality of queue pairs associated therewith, i.e., at least two source graphics processor card-to-destination graphics processor card pairs can be interconnected simultaneously, which can significantly improve efficiency.

[0069] In some embodiments, the read-write request is sent to the switch simultaneously based on at least two of the plurality of peer-to-peer network ports. For example, referring to the system 100, the peer-to-peer network port P1 of the graphics processor card 102 serves as a source peer-to-peer network port, and at least one of the plurality of queue pairs associated with the peer-to-peer network port P1 is used to send the read-write request to the switch simultaneously; meanwhile, the peer-to-peer network port P2 of the graphics processor card 102 serves as a source peer-to-peer network port, and at least one of the plurality of queue pairs associated with the peer-to-peer network port P2 is used to send the read-write request to the switch simultaneously. Therefore, the interconnection emulation between at least two pairs of source graphics processor card and destination graphics processor card can be performed simultaneously, and the efficiency can be improved significantly.

[0070] Then, at the switch 104, the received read-write request is forwarded to the peer-to-peer network port associated with the destination logical address based on the obtained source logical address information and the destination logical address. For example, the switch forwards the received read-write request to the peer-to-peer network port associated with the destination logical address according to the preconfigured routing information based on the obtained source logical address information and the destination logical address.

[0071] For example, the switch 104 parses the received read-write request (e.g., read-write request packet) to obtain the source logical address information and the destination logical address carried in the read-write request. Then, the switch 104 forwards the received read-write request to the peer-to-peer network port associated with the destination logical address according to the preconfigured routing information based on the obtained source logical address information and the destination logical address, thereby emulating the interconnection of, for example, two graphics processor cards.

[0072] As described above, referring to the system 100, the plurality of peer-to-peer network ports of the graphics processor card 102 can be combined in a plurality of pairing combinations of source peer-to-peer network port and destination peer-to-peer network port, and each combination can emulate the interconnection of, for example, two graphics processor cards. The plurality of pairing combinations can be used sequentially to test the interconnection of the emulated graphics processor cards, or at least two of the pairing combinations can be used concurrently to test the interconnection of the emulated graphics processor cards.

[0073] Therefore, the technical solution of the embodiments of the present disclosure can simulate the scenario of the interconnection of a plurality of graphics processor cards to perform emulation, and the efficiency of the emulation verification can be improved. In addition, one graphics processor card can be used to simulate the scenario of the interconnection of a plurality of graphics processor cards to perform emulation, and the hardware resources and costs can be saved significantly.

[0074] In some embodiments, based on the system 100, with reference to the method 300, in-network computing verification about the graphics processor card can be performed, for example. The graphics processor card 102 sends a data read request (LDR) to the switch 104, for example. The destination address of the request message contains an identification field mcId (multicast group identification), which is used to specify the multicast group to which this read request belongs. The switch 104 can quickly resolve all destination peer-to-peer network ports corresponding to the multicast group (i.e., which peer-to-peer network ports to initiate data read requests to) through the pre-configured “mcId - port mapping table”. The switch 104 forwards the LDR request to all destination peer-to-peer network ports in the corresponding multicast group in a broadcast manner according to the multicast information. For the destination peer-to-peer network port of the graphics processor card 102 that receives the request, it reads data according to the address and replies data to the switch 104. The switch 104 stores the received data, and when all the reduction data is collected, it performs reduction calculation. Then, the switch 104 sends the reduction result to the LDR request end.

[0075] It is worth noting that the technical solutions of the embodiments of the present disclosure can be applied to the field of silicon pre-verification of high-performance computing chips of graphics processors, for example, and are deeply adapted to EMU (intelligent management unit), FPGA (field programmable gate array), and other related simulation platforms. For example, a high-performance switch is used as the core interconnection hub to simulate the interconnection topology and coordination logic of multiple graphics processor cards. This solution not only generates multiple card interconnection scenarios that are several times larger than the actual carrying capacity of the platform through flexible topology virtualization technology, but also accurately completes comprehensive verification of core capabilities such as multi-card coordination scheduling and data transmission of the chip, thereby achieving comprehensive verification of core capabilities such as multi-card coordination scheduling, data transmission, and in-network computing.

[0076] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement a method for processing a target object of an embodiment of the present disclosure is shown. As shown, the electronic device 500 includes a central processing unit (i.e., CPU 501), which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (i.e., ROM 502) or loaded from a storage unit 508 into a random access memory (i.e., RAM 503). In the RAM 503, various programs and data required for the operation of the electronic device 500 can also be stored. The CPU 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output interface (i.e., I / O interface 505) is also connected to the bus 504.

[0077] A number of components in the electronic device 500 are connected to the I / O interface 505, including an input unit 506, such as a keyboard, a mouse, a microphone, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, a magneto-optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0078] The various processes and processes described above, such as the method 300, can be performed by the CPU 501. For example, in some embodiments, the method 300 can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the CPU 501, one or more actions of the method 300 described above can be performed.

[0079] The present disclosure relates to methods, apparatuses, systems, electronic devices, computer-readable storage media, and / or computer program products. The computer program product can include computer readable program instructions for executing various aspects of the present disclosure.

[0080] In some embodiments, the method 300 described above can be implemented as a computer program product. The computer program product can include a computer readable storage medium having computer readable program instructions loaded thereon for executing various aspects of the present disclosure.

[0081] Computer readable storage media can be tangible storage media which can retain and store instructions for use by an instruction execution device. Computer readable storage media can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer readable storage media include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.

[0082] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0083] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0084] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0085] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other data storage device. When the computer readable program instructions are loaded into the computer and other programmable data processing apparatus, a series of operational steps are implemented that provide processes such that the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0086] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0087] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0088] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

[0089] The above are merely optional embodiments of this disclosure and are not intended to limit this disclosure. Various modifications and variations can be made to this disclosure by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for emulating multi-card interconnection, characterized by, The method comprises: connecting a plurality of peer-to-peer network ports of a graphics processor card to a switch respectively, and associating each of the plurality of peer-to-peer network ports with a plurality of different logical addresses for emulating ports of a plurality of graphics processor cards; the switch parsing a received read-write request to obtain source logical address information and a destination logical address carried by the read-write request; and forwarding the received read-write request to a peer-to-peer network port associated with the destination logical address based on the obtained source logical address information and the destination logical address for testing interconnection of the emulated graphics processor card.

2. The method of claim 1, wherein, The destination logical address indicates a peer-to-peer network port identifier and a queue pair identifier, the peer-to-peer network port identifier being used to indicate a peer-to-peer network port associated with a source of the read-write request, and the queue pair identifier being used to indicate a queue pair used for data transmission of the read-write request.

3. The method of claim 1, wherein, The association of each of the plurality of peer-to-peer network ports with a plurality of different logical addresses comprises: associating each of the peer-to-peer network ports with a plurality of queue pairs, and configuring each of the queue pairs with a different logical address to simulate a different logical communication link, each of the plurality of queue pairs being associated with a queue pair identifier, and each of the plurality of queue pairs being associated with other peer-to-peer network ports than the each of the peer-to-peer network ports respectively.

4. The method of claim 1, wherein, The association of each of the plurality of peer-to-peer network ports with a plurality of different logical addresses comprises: configuring each of the peer-to-peer network ports with an emulated graphics processor identifier, the emulated graphics processor identifier being associated with a unique Internet Protocol address and a Media Access Control bit address; binding a plurality of queue pairs for each of the peer-to-peer network ports, and configuring each of the queue pairs with a different Internet Protocol address and a Media Access Control bit address.

5. The method of claim 1, wherein, The forwarding of the received read-write request to the peer-to-peer network port associated with the destination logical address based on the obtained source logical address information and the destination logical address comprises: the switch forwarding the received read-write request to the peer-to-peer network port associated with the destination logical address based on the obtained source logical address information and the destination logical address according to preconfigured routing information.

6. The method of claim 3, wherein, Further comprising: sending read-write requests to the switch based on at least two of a plurality of queue pairs associated with a same peer-to-peer network port simultaneously; and / or sending read-write requests to the switch based on at least two of a plurality of peer-to-peer network ports simultaneously. Comprise:

7. A system for emulating multi-card interconnect, the system comprising: a graphics processor card, a plurality of peer-to-peer network ports of the graphics processor card being connected to a switch respectively, each of the plurality of peer-to-peer network ports being associated with a plurality of different logical addresses for emulating ports of a plurality of graphics processor cards; and ​ The switch is configured to, for a received read or write request, parse the read or write request to obtain source logical address information and destination logical address information carried by the read or write request, and forward the received read or write request to a peer network port associated with the destination logical address based on the obtained source logical address information and destination logical address, for testing an interconnect of the emulated graphics processor card.

8. The system of claim 7, wherein, The number of graphics processor cards is at least two, and each of the at least two graphics processor cards comprises a plurality of peer network ports.

9. A computing device, comprising: Comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer readable storage medium, the computer program being executed by a machine to perform the method of any one of claims 1-6.

11. A computer program product, characterised in that, A computer program is stored on the computer readable storage medium, the computer program being executed by a machine to perform the method of any one of claims 1-6.

Citation Information

Patent Citations

  • Simulation verification system and method for remote direct memory access simulation verification

    CN116028292A

  • PCIe switch chip pre-silicon simulation system

    CN117556754A

  • Multi-card cascaded simulation unit model and communication method thereof

    CN119179664A

  • Host internal network delay diagnosis method based on loopback test

    CN119788571A

  • Data routing and multiplexing architecture to support serial links and advanced relocation of emulation models

    US10860763B1