GPU Die - to - Die Data Transfer Modeling Method, System, Device, and Readable Storage Medium
By configuring an emulator for each GPU core particle under the SST framework and using SST Link connection, the simulation of data transmission between GPUs is achieved, solving the problem that GPGPU-Sim cannot simulate direct communication between GPUs, and improving the performance and design efficiency of multi-GPU systems.
Patent Information
- Application Number
- CN202510130950.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-02-06
AI Technical Summary
The existing GPU simulator GPGPU-Sim cannot effectively simulate the direct communication function between GPUs, resulting in limited data transmission bandwidth and large delay, affecting the performance and efficiency of multi-GPU systems.
Using the SST framework and GPGPU-Sim simulator, by configuring the simulator for each GPU core and abstracting it into SST components, using SST Link for connection and event transmission, simulating data transmission between GPUs, and supporting CUDA's cudaMemcpyPeer function.
It realizes efficient simulation modeling of multi-GPU systems, improves system design efficiency, reduces verification costs, and supports performance testing and verification of multi-GPU systems.
Smart Images

Figure CN119576674B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer structure simulation, and particularly to a method, system, device and readable storage medium for modeling data transmission between GPU dies. Background Art
[0002] With the development of artificial intelligence applications, in the currently designed systems of servers and computing centers, GPUs have become the main components providing computing power, and there are more systems using multiple GPUs as the main computing force. However, for multiple GPUs within the same PCIe node, if the calculation results or data of GPU0 need to be transmitted to GPU1, the communication between the two GPUs completely depends on the CPU, that is, CPU0 first transmits the data to the CPU, and then the CPU transmits the data to GPU0. At this time, it can be seen that the data transmission bandwidth is limited by the CPU bandwidth, and due to two copies, the data latency is also large. In order to improve the utilization efficiency of multiple GPUs and the overall performance of the system, GPU manufacturers provide the function of direct communication between GPUs. For example, NVIDIA has proposed the GPU peer to peer (P2P) technology. When users need to use this P2P technology, they need to use a series of API interfaces provided by CUDA. However, the currently most commonly used GPU simulator, GPGPU-Sim, does not support the simulation of these CUDA functions. Therefore, there is an urgent need for a modeling method for data transmission between GPUs to be able to simulate the data transmission between GPUs in a multi-GPU system. Summary of the Invention
[0003] Therefore, the present invention provides a method, system, device and readable storage medium for modeling data transmission between GPU dies, and realizes the simulation of data transmission between GPUs in a multi-GPU die system based on the SST framework and the GPGPU-Sim simulator.
[0004] According to the design solution provided by the present invention, on the one hand, a method for modeling data transmission between GPU dies is provided, including:
[0005] Configuring a GPU simulator for each GPU die so that each GPU die realizes the die function in an independent GPU simulator respectively, and the simulator executes the GPU die operation logic using GPGPU-Sim;
[0006] Each GPU simulator abstracts each GPU die simulator into a corresponding SST component based on the SST structure simulation toolkit, connects each GPU die simulator through SST Link, uses SST Link to send SST events and receives SST events through the event processor in the peer SST component to simulate the data transmission between GPU dies.
[0007] As the method for modeling data transmission between GPU dies of the present invention, further, each GPU die simulator is abstracted into a corresponding SST component, including:
[0008] Instantiate the GPU die simulator using the SST component code, create and inherit the GPU type of the SST component, and create a clock function and an external Link interface of the SST component. The clock function is used to manage the clocks of all SST components in the SST structure simulation toolkit;
[0009] Register an event handling function for the Link of each SST component corresponding to the GPU die simulator. The event handling function is used to handle SST events passed to the corresponding component of the GPU die simulator through the SST Link.
[0010] As the method for modeling data transmission between GPU dies of the present invention, further, use the SST Link to send SST events to simulate the data transmission between GPU dies, including:
[0011] Take Ariel as the master device connected to multiple GPU simulators to detect the specified function calls in the program using the Ariel PIN stamping tool. The specified function is the cudaMemcpyPeer function;
[0012] When the specified function call is detected, convert the specified function call into an SST event and send it to the corresponding SST component of the GPU simulator;
[0013] After the corresponding component of the GPU simulator receives the SST event, use the cudaMemcpy function from Device To Host to extract the SST event data and store the data in the SST component, and use the cudaMemcpy function from Host To Device to send the data to the SST component of the peer GPU simulator.
[0014] On the other hand, the present invention also provides a system for modeling data transmission between GPU dies, including: a die simulation configuration module and a component instantiation module, where
[0015] The die simulation configuration module is used to configure a GPU simulator for each GPU die so that each GPU die can implement the die function in an independent GPU simulator. The simulator uses GPGPU-Sim to execute the GPU die operation logic;
[0016] A component instantiation module is used for each GPU simulator to abstract each GPU die simulator into a corresponding SST component based on the SST structure simulation toolkit, connect the GPU die simulators through SST Link, send SST events using SST Link, and receive SST events through the event processor in the peer SST component to simulate the data transmission between GPU dies.
[0017] On the other hand, the present invention also provides an electronic device, including:
[0018] At least one processor, and a memory coupled to the at least one processor;
[0019] Wherein, the memory stores a computer program, and the computer program can be executed by the at least one processor to implement the method as described above.
[0020] On yet another aspect, the present invention also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed, it can implement the method as described above.
[0021] Advantages of the present invention:
[0022] The present invention uses the master simulator to detect the CUDA functions that call GPU-to-GPU communication in the user program, adds an SST component in the SST framework to instantiate GPGPU-sim to achieve data transmission between GPU simulators, and uses multiple GPGPU-Sim simulators to simulate the cudaMemcpyPeer function of GPU-to-GPU data transmission, enabling chip designers to perform simulation modeling on a system with multiple GPUs during the chip design stage to test and verify the performance of the system, accelerate the system design time, reduce the system verification cost, and have good application prospects in the field of chip design and verification testing. Description of the Drawings
[0023] Figure 1 Schematic diagram of the modeling process for data transmission between GPU dies in the embodiment;
[0024] Figure 2 Schematic diagram of GPGPU-Sim accessing the SST framework in the embodiment;
[0025] Figure 3 Schematic diagram of Ariel as the master simulator of GPGPU-Sim in the embodiment;
[0026] Figure 4 Schematic diagram of running multiple GPGPU-Sim simulators under the SST framework in the embodiment. Detailed Embodiments
[0027] To make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and technical solutions.
[0028] For the modeling and simulation of data transmission between GPUs in a multi-GPU system, an embodiment of the present invention is described with reference to Figure 1 As shown, a method for modeling data transmission between GPU dies is provided, including:
[0029] S101. Configure a GPU simulator for each GPU die so that each GPU die can implement die functions in an independent GPU simulator respectively, and the simulator uses GPGPU-Sim to execute the GPU die operation logic;
[0030] S102. Based on the SST structure simulation toolkit, each GPU simulator abstracts each GPU die simulator into a corresponding SST component, connects each GPU die simulator through SST Link, and uses SST Link to send SST events and receives SST events through the event processor in the peer SST component to simulate the data transmission between GPU dies.
[0031] Specifically, abstracting each GPU die simulator into a corresponding SST component can be designed to include:
[0032] Instantiate the GPU die simulator using the SST component code, create and inherit the GPU type of the SST component, and create a clock function and an external Link interface of the SST component. The clock function is used to manage the clocks of all SST components in the SST structure simulation toolkit;
[0033] Register an event handling function for the Link of each GPU die simulator corresponding SST component. The event handling function is used to process the SST events transmitted to the corresponding component of the GPU die simulator through SST Link.
[0034] As Figure 2 shown, in the SST framework, each individual simulator is abstracted into an SST component. SST components are connected through SST Link, and an SST component can send SST events to the peer SST component through SST Link. Create an SST component of GPGPU-Sim, and this component will instantiate a GPGPU-Sim simulator. This component contains an SST event processor used to receive events from SST. After receiving the SST event, this component will interpret it and convert it into a C / C++ function call to GPGPU-Sim.
[0035] Among them, using the SST Link to send SST events to simulate data transmission between GPU dies may include:
[0036] Taking the local host as the master device connected to multiple GPU simulators to detect specified function calls in the program using the master device, where the specified function is the cudaMemcpyPeer function;
[0037] When a specified function call is detected, convert the specified function call into an SST event and send it to the corresponding SST component of the GPU simulator;
[0038] After the corresponding component of the GPU simulator receives the SST event, use the cudaMemcpy function from Device To Host to extract the SST event data and store the data in the SST component, and use the cudaMemcpy function from Host To Device to send the data to the SST component of the peer GPU simulator.
[0039] The GPU is a slave device and requires a master device to control it. Generally, this master device is a CPU. In the GPGPU-Sim simulator, it uses the local host as the master device. As Figure 3 shown, in the SST framework, there is an x86 CPU simulator: Ariel. In the embodiments of this case, this CPU simulator can be used as the master simulator. Ariel can detect CUDA function calls in the program through Intel's PIN instrumentation tool. When it detects a CUDA function call, it can be converted into an SST event and sent to the SST component of GPGPU-Sim to achieve the call of GPGPU-Sim. Through the Ariel simulator, the use of the cudaMemcpyPeer function by the user can be detected, and multiple SST components of GPGPU-Sim can be supported to connect to the Ariel component.
[0040] Specifically, for data transmission between GPU simulators, such as Figure 4As shown, the GPGPU-Sim simulator itself does not directly support data transfer functions such as cudaMemcpyPeer for Device To Device data transfer between GPUs. It only supports cudaMemcpy for Host To Device and Device To Host. In the implementation of this case, through the Ariel simulator, users are supported to use the cudaMemcpyPeer function in the program. After detecting the cudaMemcpyPeer function, the Ariel simulator sends an SST event to the SST component of GPGPU-Sim created in the present invention. After receiving the event, the SST component of GPGPU-Sim first uses the cudaMemcpy function of Device To Host for the GPU component of src to extract the data, but does not further send it to the Ariel component as the HOST. Instead, it stores the data in the SST component of GPGPU-Sim, and then uses the cudaMemcpy function of Host To Device for the SST component of dest GPGPU-Sim to send the data to another GPGPU-Sim simulator, so as to achieve data transfer between GPUs by running multiple GPGPU-sim simulators under the SST framework.
[0041] Furthermore, based on the above method, an embodiment of the present invention also provides a GPU die-to-die data transfer modeling system, including: a die simulation configuration module and a component instantiation module, where,
[0042] The die simulation configuration module is used to configure a GPU simulator for each GPU die, so that each GPU die can implement die functions in an independent GPU simulator respectively. The simulator uses GPGPU-Sim to execute the GPU die operation logic;
[0043] The component instantiation module is used to abstract each GPU die simulator into a corresponding SST component based on the SST structure simulation toolkit for each GPU simulator, connect each GPU die simulator through SST Link, send SST events through SST Link and receive SST events through the event processor in the peer SST component, so as to simulate the data transfer between GPU dies.
[0044] When instantiating the SST component of the GPU simulator, add the SST component code of GPGPU-Sim to the sst-element\src\sst\elements directory, create a GPU type that inherits from SST:Compenet, create a clockTic function through which the SST framework uniformly manages the clocks of all SST components, create a link interface for the GPGPU-Sim component to the outside, and register an event handling function for the link of the GPGPU-Sim component to handle the SST events passed to the GPGPU-Sim component through this link.
[0045] In the process of the master simulator detection program calling the CUDA function for GPU communication, the monitoring of the cudaMemcpyPeer function can be added to the InstrumentRoutine function in the fesimple.cc file of the front-end PIN tool of the Ariel simulator. After detecting the call of cudaMemcpyPeer, use the RTN_Replace function to replace it with the call of the instrumented function mapped_cudaMemcpyPeer. In the mapped_cudaMemcpyPeer function, encapsulate the cuda message and send it to the Ariel backend code. In the processNextEvent function of the Ariel backend arielcore.cc, add a case for the cudaMemcpyPeer event type and send the event to the SST component of GPGPU-Sim of the GPU in src of cudaMemcpyPeer through the SST Link.
[0046] The specific process of implementing data transmission between GPU simulators using the GPGPU-Sim component under the SST framework can be described as the following steps:
[0047] Step 1: Add a case for the cudaMemcpyPeer event type to the handleCudaCall function in the gpuInterface.cc file of the SST component of GPGPU-Sim.
[0048] Step 2: In the cudaMemcpyPeer event handling function, perform a Device To Host cudaMemcpy call on the GPGPU-Sim simulator of src.
[0049] Step 3: After the GPGPU-Sim simulator of src finishes execution, transfer the data back to the SST component of src GPGPU-Sim.
[0050] Step 4: After receiving the data, the SST component of src GPGPU-Sim does not send it to the Host, but encapsulates the data into the cudaMemcpyPeer event.
[0051] Step 5: Send the event to the SST component of dest GPGPU-Sim through the SST Link.
[0052] Step 6: After receiving the event, the SST component of dest GPGPU-Sim performs a HostToDevice cudaMemcpy call on the GPGPU-Sim of dest.
[0053] Step 7: The GPGPU-Sim simulator of dest performs the data copy operation.
[0054] Step 8: Thus, the operation of directly transferring data between two GPGPU-Sim simulators through the cudaMemcpyPeer function without passing through the Host is completed.
[0055] Running multiple GPGPU-Sim simulators based on the SST framework to support users in using the cudaMemcpyPeer function to simulate the behavior of directly transferring data between multiple GPU simulators, so as to verify the performance of the multi-GPU system, can greatly improve the system design efficiency and effectively reduce the test and verification cost.
[0056] Furthermore, an embodiment of the present invention also provides an electronic device, including:
[0057] At least one processor, and a memory coupled to the at least one processor;
[0058] Wherein, the memory stores a computer program, and the computer program can be executed by the at least one processor to implement the method as described above.
[0059] Furthermore, an embodiment of the present invention also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed, the method as described above can be implemented.
[0060] Unless otherwise specifically stated, the relative steps, numerical expressions and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0061] In the present specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.
[0062] The units and method steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation is not considered to exceed the scope of the present invention.
[0063] Those of ordinary skill in the art can understand that all or part of the steps in the above methods can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium, such as a read-only memory, a magnetic disk, or an optical disc, etc. Optionally, all or part of the steps of the above embodiments can also be implemented using one or more integrated circuits. Correspondingly, the various modules / units in the above embodiments can be implemented in the form of hardware or in the form of software function modules. The present invention is not limited to any specific form of the combination of hardware and software.
[0064] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present invention, used to illustrate the technical solutions of the present invention, rather than to limit it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for modeling data transmission between GPU dies, comprising: Comprising: Configuring a GPU simulator for each GPU die so that each GPU die implements die functions in an independent GPU simulator respectively, and the simulator executes the GPU die operation logic by using GPGPU-Sim; Based on the SST structure simulation toolkit, each GPU simulator abstracts each GPU die simulator into a corresponding SST component, connects each GPU die simulator through SST Link, uses SST Link to send SST events and receives SST events through the event processor in the peer SST component to simulate the data transmission between GPU dies; Using SST Link to send SST events to simulate the data transmission between GPU dies includes: Taking Ariel as the master device connected to multiple GPU simulators to detect the specified function call in the program by using the Ariel PIN instrumentation tool, and the specified function is the cudaMemcpyPeer function; When the specified function call is detected, converting the specified function call into an SST event and sending it to the corresponding SST component of the GPU simulator; After the corresponding component of the GPU simulator receives the SST event, using the cudaMemcpy function in the Device To Host direction to extract the SST event data and store the data in the SST component; The SST component then uses the cudaMemcpy function in the Host To Device direction to directly send the data to the SST component of the peer GPU simulator without passing through the CPU.
2. The GPU die - to - die data transfer modeling method according to claim 1, wherein, Abstracting each GPU die simulator into a corresponding SST component includes: Instantiating the GPU die simulator by using the SST component code, creating a GPU type that inherits from the SST component, and creating a clock function and an external Link interface of the SST component, where the clock function is used to manage the clocks of all SST components of the SST structure simulation toolkit; Registering an event processing function for the Link of each SST component corresponding to the GPU die simulator, and the event processing function is used to process the SST events passed to the corresponding component of the GPU die simulator through SST Link.
3. A GPU die - to - die data transfer modeling system, characterized in that, For implementing the method for modeling data transmission between GPU dies as described in claim 1 or 2, the system includes: a die simulation configuration module and a component instantiation module, wherein, The die simulation configuration module is used to configure a GPU simulator for each GPU die so that each GPU die implements die functions in an independent GPU simulator respectively, and the simulator executes the GPU die operation logic by using GPGPU-Sim; The component instantiation module is used to, based on the SST structure simulation toolkit, abstract each GPU die simulator into a corresponding SST component by each GPU simulator, connect each GPU die simulator through SST Link, use SST Link to send SST events and receive SST events through the event processor in the peer SST component to simulate the data transmission between GPU dies.
4. An electronic device, characterized in that, Including: At least one processor, and a memory coupled to the at least one processor; Wherein, the memory stores a computer program, and the computer program can be executed by the at least one processor to implement the method according to claim 1 or 2.
5. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed, it can implement the method according to claim 1 or 2.
Citation Information
Patent Citations
Multi-core particle interconnection simulation method and device, storage medium and electronic equipment
CN117236263A