Simulation method, device, equipment, medium and program product
By creating a virtual node collection and a virtual communication network in a discrete event simulation system, the problem of difficulty in simulating the graphics processing state of large-scale GPU clusters is solved in the prior art, and a high-precision simulation effect is achieved.
Patent Information
- Application Number
- CN202411930303.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to simulate the graphics processing state of large-scale GPU clusters, and it is impossible to effectively simulate the graphics processing state in the GPU cluster.
By creating a virtual node collection and a virtual communication network in a discrete event simulation system, the data processing status of physical nodes in the GPU cluster is simulated, and the matching degree between virtual nodes and physical nodes and the consistency of network topology is improved.
Large-scale simulation of the graphics processing state of physical nodes in the GPU cluster is realized, which improves the accuracy and stability of the simulation, and meets the requirements for GPU node simulation in the GPU cluster.
Smart Images

Figure CN119938470A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of cloud computing technology, and in particular to a simulation method, device, equipment, medium and program product. Background Art
[0002] In practical applications, simulation of a graphics processing unit (GPU) is usually implemented based on a simplified mathematical model. The above method can simulate the graphics processing state of a few GPUs, but cannot simulate the graphics processing state of a large-scale GPU cluster. Summary of the invention
[0003] Based on the above technical problems, the embodiments of the present application provide a simulation method, device, equipment, medium and program product.
[0004] The technical solution provided by the embodiment of the present application is as follows:
[0005] The present application embodiment first provides a simulation method, the method comprising:
[0006] Acquire a first set; wherein the first set includes a set of at least some physical nodes in a GPU cluster; the physical nodes include GPU nodes in the GPU cluster, and node devices based on which communication between any GPU nodes in the GPU cluster is based;
[0007] Creating a second set corresponding to the first set in a discrete event simulation system; wherein the second set includes a set of virtual nodes corresponding to the at least part of the physical nodes;
[0008] constructing a virtual communication network corresponding to the second set in the discrete event simulation system;
[0009] Based on at least one virtual operating state of the virtual nodes in the virtual communication network, the data processing state of at least part of the physical nodes is simulated.
[0010] The embodiment of the present application also provides a simulation device, the simulation device comprising:
[0011] An acquisition module, configured to acquire a first set; wherein the first set includes a set of at least some physical nodes in a GPU cluster; the physical nodes include GPU nodes in the GPU cluster, and node devices based on which communication between any GPU nodes in the GPU cluster is based;
[0012] A processing module, configured to create a second set corresponding to the first set in a discrete event simulation system; and construct a virtual communication network corresponding to the second set in the discrete event simulation system; wherein the second set includes a set of virtual nodes corresponding to the at least part of the physical nodes;
[0013] The simulation module is used to simulate the data processing state of at least part of the physical nodes based on at least one virtual operation state of the virtual nodes in the virtual communication network.
[0014] An embodiment of the present application also provides an electronic device, comprising a processor and a memory; wherein a computer program is stored in the memory; when the computer program is executed by the processor, it can implement any of the simulation methods described above.
[0015] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored; when the computer program is executed by a processor of an electronic device, the simulation method as described above can be implemented.
[0016] An embodiment of the present application also provides a computer program product, which includes a computer program; when the computer program is executed by a processor of an electronic device, it can implement any of the simulation methods described above.
[0017] The simulation method provided in the embodiment of the present application obtains a first set including at least some physical nodes in a GPU cluster, and the physical nodes include GPU nodes in the GPU cluster, and node devices based on which communication between any GPU nodes in the GPU cluster is based, and creates a second set corresponding to the first set in a discrete event simulation system, and the second set includes a set of virtual nodes corresponding to at least some physical nodes, so that the functions of the physical nodes in the first set can be fully and accurately represented through the virtual nodes in the second set; and, by means of the advantages of the discrete event simulation system in the dimensions of dynamic simulation and system modeling, the matching degree of the virtual nodes in the second set with the physical nodes in the first set in terms of graphics processing functions can be improved; at the same time, A virtual communication network corresponding to the second set is constructed in the discrete event simulation system. With the help of the network protocol that can be supported by the discrete event simulation, the stability of the virtual communication network can be improved, and the consistency between the network topology of the virtual communication network and the network topology corresponding to at least some physical nodes can also be improved; on this basis, based on at least one virtual operating state of the virtual nodes in the virtual communication network, the data processing state of at least some physical nodes is simulated, which can improve the consistency between the virtual operating state and the data processing state, thereby improving the accuracy of the simulation of the data processing state of at least some physical nodes, and then realizing large-scale simulation of the graphics processing state of the physical nodes in the GPU cluster, which can meet the simulation requirements of the GPU nodes in the GPU cluster. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 A schematic diagram of a simulation method according to an embodiment of the present invention;
[0019] Figure 2 Another schematic diagram of a simulation method according to an embodiment of the present invention;
[0020] Figure 3 A schematic diagram of the structure of a simulation device provided in an embodiment of the present application;
[0021] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0022] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0023] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0024] GPU collective communication is a technology that implements data exchange between multiple GPU nodes in a distributed system. Through GPU collective communication, the efficiency of parallel computing of GPU nodes can be improved, and the graphics processing performance of GPU nodes can be improved. In practical applications, common application areas of GPU collective communication include machine learning, scientific computing, and graphics rendering.
[0025] The basic operations of GPU collective communication include broadcast, reduce, allreduce, scatter, gather, allgather, and all-to-all. Broadcast is used to copy data on one GPU to all other GPUs. Reduce is used to aggregate data on all GPUs to one GPU according to rules including summing or finding the maximum value. Allreduce is used to aggregate data on all GPUs according to certain rules and copy the results to all GPUs. Scatter is used to divide data on one GPU into several blocks and distribute them to all other GPUs. Gather is used to collect data on all GPUs to one GPU. Allgather is used to collect data on all GPUs and copy the results to all GPUs. All-to-all is used to divide data on each GPU into several blocks and send and receive data blocks on other GPUs according to certain rules so that each GPU has a complete data set.
[0026] With the widespread application of GPUs in machine learning and high-performance computing, efficient data exchange and collaboration between GPU devices are becoming increasingly important; at the same time, machine learning and high-performance computing have also become a major hot topic in academic research and corporate technology research. However, the high cost of GPU devices has restricted the progress of GPU research in machine learning and high-performance computing. In order to solve the above problems, related technologies provide a technical solution for simulating the graphics processing process of GPU devices to alleviate the negative impact of the high cost of GPU resources on its machine learning and high-performance computing. At the same time, it can also verify the logic of the algorithm in advance through simulation before the actual GPU system test.
[0027] The GPU simulation methods provided in the related art mainly include the following:
[0028] Mathematical analysis based on simplified models mathematically analyzes performance by establishing a communication model. However, the simplified model cannot fully reflect the data processing state of the GPU device or node in the actual graphics processing process.
[0029] The simulation scenario is constructed by manually coding a predefined event sequence to simulate the GPU data processing process through discrete event simulation, but this solution can only simulate simple static scenarios; the cluster environment is built using the actual system and the algorithm is debugged on a real physical platform, but this method is costly and difficult to expand.
[0030] In order to solve the above technical problems, the embodiments of the present application provide a simulation method, device, equipment, medium and program product.
[0031] Figure 1 A flow chart of the simulation method provided in the embodiment of the present application is shown in FIG. Figure 1 As shown, the method may include the following steps:
[0032] Step 101: Get the first set.
[0033] The first set includes a set of at least part of the physical nodes in the GPU cluster; the physical nodes include the GPU nodes in the GPU cluster, and a set of node devices based on which communication between any GPU nodes in the GPU cluster is based.
[0034] In one implementation, the GPU cluster may include a GPU cluster that has not yet been constructed or has been constructed and includes multiple GPU nodes.
[0035] In one implementation, a GPU cluster can implement distributed graphics processing.
[0036] In one implementation, the node device on which communication between any GPU nodes in the GPU cluster is based may include a data transfer device, a network status detection device, or a data storage device.
[0037] In one implementation, the network topology of the GPU cluster may be predetermined.
[0038] Step 102: Create a second set corresponding to the first set in a discrete event simulation system.
[0039] The second set includes a set of virtual nodes corresponding to at least some of the physical nodes.
[0040] In one embodiment, the discrete event simulation system can implement process simulation and model simulation, and can simulate multiple data communication protocols, multiple network topologies, and multiple data processing algorithms in a computer simulation environment; illustratively, the discrete event simulation system may include an NS3 system.
[0041] Among them, the NS3 system is a discrete event simulator that can realize NS3 simulation; among them, NS3 simulation is a process of simulating a computer network using NS3 software, which can realize simulation in exploring the performance research of network protocols, topology, and traffic.
[0042] The NS3 system has the following features:
[0043] First of all, the NS3 system can support multiple network technologies. The NS3 system can simulate multiple networks including wired networks, wireless networks, mobile networks, satellite networks, and sensor networks. It can also simulate the interconnection status and interoperation process between node devices in the above-mentioned networks.
[0044] The NS3 system supports multiple network protocols, including Transmission Control Protocol / Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Internet Control Message Protocol (ICMP), Address Resolution Protocol (ARP), Domain Name System (DNS) and Hyper Text Transfer Protocol (HTTP), etc. It can also simulate the data interaction behavior between the above network protocols.
[0045] The NS3 system supports a variety of network applications. The NS3 system can simulate a variety of network applications including File Transfer Protocol (FTP), Telnet, Voice over Internet Protocol (VoIP), video streaming, and Point to Point (P2P), and can simulate their impact and requirements on the network.
[0046] Supporting multiple simulation modes, the NS3 system can support multiple simulation modes including offline simulation, real-time simulation and distributed simulation to meet different simulation needs and scenarios.
[0047] Supporting multiple simulation outputs, the NS3 system can output various simulation results including log files, trace files, statistical data, and graphical interfaces, and can also intuitively display the above simulation results.
[0048] In one implementation, at least part of the physical nodes may include a set of GPU nodes in a GPU cluster that need to perform simulation and node devices on which their communication depends.
[0049] In one implementation, there may be a one-to-one correspondence between the virtual nodes in the second set and the physical nodes in the first set.
[0050] In one implementation, the second set may be created in the following manner:
[0051] The data computing operations, data transmission operations and state switching that can be performed by the physical nodes in the first set are analyzed to obtain a physical parameter set, and then code is created based on the physical parameter set, and virtual nodes corresponding to the physical nodes in the first set are constructed through the code to obtain a second set.
[0052] Exemplarily, the first function set possessed by the virtual nodes in the second set in the NS3 system can match the second function set possessed by at least some physical nodes; wherein the first function set and the second function set may include functions such as graphics processing, data storage, data transmission, and status update.
[0053] Step 103: construct a virtual communication network corresponding to the second set in the discrete event simulation system.
[0054] In one embodiment, the virtual communication network may include a virtual network in the NS3 system for connecting the virtual nodes in the second set.
[0055] In one embodiment, the connection relationship between virtual nodes in the virtual communication network may be the same as the connection relationship between at least some physical nodes.
[0056] In one embodiment, a virtual communication network may be constructed in the following manner:
[0057] Determine a one-to-one mapping relationship between a first identification set of at least some physical nodes and a second identification set of virtual nodes in a second set, determine a physical network topology corresponding to at least some physical nodes, and then connect the virtual nodes in the second set based on the physical network topology and the above mapping relationship to obtain a virtual communication network; wherein the first identification set includes a set of identifications of at least some physical nodes; and the second identification set includes a set of identifications of virtual nodes in the second set.
[0058] Step 104: Based on at least one virtual operating state of the virtual nodes in the virtual communication network, simulate the data processing state of at least part of the physical nodes.
[0059] In one embodiment, the virtual operating state may include the state in which data processing resources of a virtual node in a virtual communication network are occupied or released during graphics processing, data transmission, data storage, and status update; exemplarily, the data processing resources may include at least one of memory capacity, disk or hard disk storage capacity, bandwidth, threads, and processes.
[0060] In one implementation, the virtual operation state may also include a change state of a graphic processing effect of the virtual node during the process in which the data processing resources of the virtual node are occupied or released.
[0061] In one embodiment, at least one virtual operating state may be obtained by:
[0062] Obtain a physical trigger operation for at least some physical nodes to switch to at least one data processing state, then construct a virtual trigger operation corresponding to the physical trigger operation in the NS3 system, and execute the above virtual trigger operation for the virtual communication network, and when at least some virtual nodes in the second set start to run, collect the running status of at least some virtual nodes, and integrate the above running status to obtain at least one virtual running state; wherein the physical trigger operation may include a set of operations for triggering at least some physical nodes to perform data processing, and the virtual trigger operation may include a set of operations for triggering at least some virtual nodes to perform data processing.
[0063] In one implementation, the data processing state may include a state in which data processing resources of at least some physical nodes are occupied or released during the data processing process.
[0064] In one implementation, the data processing state may also include a change state of a graphic processing effect of a physical node during a process in which the data processing resources of the physical node are occupied or released.
[0065] In one implementation, the data processing status of at least some of the physical nodes may be simulated by any of the following methods:
[0066] Statistics are collected on at least one virtual operating state in units of time to obtain a first statistical result, and then the first statistical result is used as a simulation result of the data processing state of at least part of the physical nodes in the time dimension.
[0067] Statistics are collected on at least one virtual operating state in units of nodes to obtain a second statistical result, and then the second statistical result is used as a simulation result of the data processing state of at least part of the physical nodes in the node dimension.
[0068] From the above, it can be seen that the simulation method provided in the embodiment of the present application obtains a first set including at least part of the physical nodes in the GPU cluster, and the physical nodes include the GPU nodes in the GPU cluster, and the node devices based on which communication between any GPU nodes in the GPU cluster is based, and creates a second set corresponding to the first set in the discrete event simulation system, and the second set includes a set of virtual nodes corresponding to at least part of the physical nodes. In this way, the functions of the physical nodes in the first set can be fully and accurately represented through the virtual nodes in the second set; and, by means of the advantages of the discrete event simulation system in the dynamic simulation and system modeling dimensions, the matching degree of the virtual nodes in the second set with the physical nodes in the first set in terms of graphics processing functions can be improved; at the same time When constructing a virtual communication network corresponding to the second set in the discrete event simulation system, with the help of the network protocol that can be supported by the discrete event simulation, the stability of the virtual communication network can be improved, and the consistency between the network topology of the virtual communication network and the network topology corresponding to at least some physical nodes can also be improved; on this basis, based on at least one virtual operating state of the virtual nodes in the virtual communication network, the data processing state of at least some physical nodes is simulated, which can improve the consistency between the virtual operating state and the data processing state, thereby improving the accuracy of the simulation of the data processing state of at least some physical nodes, thereby realizing large-scale simulation of the graphics processing state of the physical nodes in the GPU cluster, which can meet the simulation requirements of the GPU nodes in the GPU cluster.
[0069] Based on the foregoing embodiment, in the simulation method provided in the embodiment of the present application, the discrete event simulation system includes the NS3 system; accordingly, constructing a virtual communication network corresponding to the second set in the discrete event simulation system can be achieved in the following manner:
[0070] Determine a target topology structure; in the NS3 system, associate the virtual nodes in the second set based on the target topology structure to obtain a virtual communication network.
[0071] In one implementation, the target topology structure may include a network topology structure for associating at least some physical nodes, and may also be any type of pre-set network topology structure; illustratively, the target topology structure may include a Ring structure and a Tree structure, etc.
[0072] In one implementation, the target topology may include any type of network topology that needs to be simulated for a GPU cluster.
[0073] In one implementation, the target topology structure may include connection relationships that at least some physical nodes should have.
[0074] In one embodiment, the virtual communication network can be obtained by:
[0075] The connection relationship that at least some physical nodes should have is obtained from the target topology structure, and then the virtual nodes in the second set are connected by corresponding configuration in the NS3 system according to the above connection relationship to obtain a virtual communication network.
[0076] As can be seen from the above, in the simulation method provided by the embodiment of the present application, the discrete event simulation system includes the NS system, and after determining the target topology, in the NS3 system, the virtual nodes in the second set are associated based on the target topology to obtain a virtual communication network. In this way, through the above operation, it is possible to achieve targeted association of the virtual nodes in the second set, thereby improving the flexibility of constructing the virtual communication network; and, with the advantage of the NS3 system supporting a variety of network topologies and network protocols, when the target topology changes, it is also possible to construct a variety of virtual communication networks corresponding to the diverse target topologies.
[0077] Based on the foregoing embodiment, in the simulation method provided in the embodiment of the present application, the virtual nodes in the second set are associated with the target topology structure to obtain a virtual communication network, which can be achieved by the following steps:
[0078] Step A1: Obtain a first configuration file associated with a target topology structure.
[0079] The first configuration file at least includes the association relationship between the first part of virtual nodes in the second set.
[0080] In one implementation, the first configuration file may be pre-configured according to the target topology after the target topology is determined.
[0081] In one implementation, the first configuration file may be presented in the form of an *.xml file; illustratively, the first configuration code corresponding to the first configuration file may be as follows:
[0082] <!-- Define connections between nodes -->
[0083] <connections>
[0084] <!-- Node 0 connected to 3 -->
[0085] <connection>
[0086] <gpu1> 0< / gpu1>
[0087] <gpu2> 3< / gpu2>
[0088] <bw> 25Gbps< / bw>
[0089] <delay> 10us< / delay>
[0090] < / connection>
[0091] <!-- Node 1 connected to 0 -->
[0092] <connection>
[0093] <gpu1> 1< / gpu1>
[0094] <gpu2> 0< / gpu2>
[0095] <bw> 25Gbps< / bw>
[0096] <delay> 10us< / delay>
[0097] < / connection>
[0098] <!-- Other Links -->
[0099] <connection> ...< / connection>
[0100] < / connections>
[0101] Among them, gpu1 to gpu3 can be used to represent three virtual GPU nodes respectively. In the first configuration code, <connections>Tags define the connection relationship between virtual GPU nodes. <connection>The label defines a connection between virtual GPU nodes and can specify the virtual GPU nodes at both ends of the connection. Of course, you can also specify the connection from GPU to NIC, or from NIC to GPU. <bw>The label defines the bandwidth of the link corresponding to the above connection. <delay>The label defines the delay parameters of the link. In practical applications, different network topologies can be constructed by defining any number of connections.
[0102] Exemplarily, the second configuration code represented by the first configuration file corresponding to the Ring type network topology structure may be as follows:
[0103] <topology>
[0104] <name> Ring< / name>
[0105] <!-- Define 4 GPU nodes -->
[0106] <gpus>
[0107] <gpu id="0"gpu_mem="16GB"gpu_bw="100Gbps" / >
[0108] <gpu id="1"gpu_mem="16GB"gpu_bw="100Gbps" / >
[0109] <gpu id="2"gpu_mem="16GB"gpu_bw="100Gbps" / >
[0110] <gpu id="3"gpu_mem="16GB"gpu_bw="100Gbps" / >
[0111] < / gpus>
[0112] <!-- Define Connection -->
[0113] <connections>
[0114] <!-- Node 0 connected to 1 -->
[0115] <connection>
[0116] <gpu1> 0< / gpu1>
[0117] <gpu2> 1< / gpu2>
[0118] <bw> 25Gbps< / bw>
[0119] <delay> 10us< / delay>
[0120] < / connection>
[0121] <!-- Node 1 connects to 2 -->
[0122] <connection>
[0123] <gpu1> 1< / gpu1>
[0124] <gpu2> 2< / gpu2>
[0125] <bw> 25Gbps< / bw>
[0126] <delay> 10us< / delay>
[0127] < / connection>
[0128] <!-- Node 2 connects to 3 -->
[0129] <connection>
[0130] <gpu1> 2< / gpu1>
[0131] <gpu2> 3< / gpu2>
[0132] <bw> 25Gbps< / bw>
[0133] <delay> 10us< / delay>
[0134] < / connection>
[0135] <!-- Node 3 connected to 0 -->
[0136] <connection>
[0137] <gpu1> 3< / gpu1>
[0138] <gpu2> 0< / gpu2>
[0139] <bw> 25Gbps< / bw>
[0140] <delay> 10us< / delay>
[0141] < / connection>
[0142] < / connections>
[0143] < / topology>
[0144] Through the second configuration code, a virtual GPU node gpu0->gpu1->gpu2->gpu3->
[0145] The Ring type topology of gpu0. By parsing the second configuration code, a corresponding virtual communication network can be constructed.
[0146] Exemplarily, the third configuration code represented by the first configuration file corresponding to the Tree type network topology structure may be as follows:
[0147] <topology>
[0148] <name> Tree< / name>
[0149] <!-- Define 4 GPU nodes -->
[0150] <gpus>
[0151] <gpu id="0"gpu_mem="16GB"gpu_bw="100Gbps" / >
[0152] <gpu id="1"gpu_mem="16GB"gpu_bw="100Gbps" / >
[0153] <gpu id="2"gpu_mem="16GB"gpu_bw="100Gbps" / >
[0154] <gpu id="3"gpu_mem="16GB"gpu_bw="100Gbps" / ><gpu id="4"gpu_mem="16GB"gpu_bw="100Gbps" / ><gpu id="5"gpu_mem="16GB"gpu_bw="100Gbps" / ><gpu id="6"gpu_mem="16GB"gpu_bw="100Gbps" / >< / gpus>
[0155] <!-- Define Connection -->
[0156] <connections>
[0157] <!-- Node 0 connected to 1 -->
[0158] <connection>
[0159] <gpu1> 0< / gpu1>
[0160] <gpu2> 1< / gpu2>
[0161] <bw> 25Gbps< / bw>
[0162] <delay> 10us< / delay>
[0163] < / connection>
[0164] <!-- Node 0 connects to 2 -->
[0165] <connection>
[0166] <gpu1> 0< / gpu1>
[0167] <gpu2> 2< / gpu2>
[0168] <bw> 25Gbps< / bw>
[0169] <delay> 10us< / delay>
[0170] < / connection>
[0171] <!-- Node 1 connected to 3 -->
[0172] <connection>
[0173] <gpu1> 1< / gpu1>
[0174] <gpu2> 3< / gpu2>
[0175] <bw> 25Gbps< / bw>
[0176] <delay> 10us< / delay>
[0177] < / connection>
[0178] <!-- Node 1 connected to 4 -->
[0179] <connection>
[0180] <gpu1> 1< / gpu1>
[0181] <gpu2> 4< / gpu2>
[0182] <bw> 25Gbps< / bw>
[0183] <delay> 10us< / delay>
[0184] < / connection>
[0185] <!-- Node 2 connects to 5 -->
[0186] <connection>
[0187] <gpu1> 2< / gpu1>
[0188] <gpu2> 5< / gpu2>
[0189] <bw> 25Gbps< / bw>
[0190] <delay> 10us< / delay>
[0191] < / connection>
[0192] <!-- Node 2 connected to 6 -->
[0193] <connection>
[0194] <gpu1> 2< / gpu1>
[0195] <gpu2> 6< / gpu2>
[0196] <bw> 25Gbps< / bw>
[0197] <delay> 10us< / delay>
[0198] < / connection>
[0199] < / connections>
[0200] < / topology>
[0201] Exemplarily, through the third configuration code, a simple Tree-type topology structure is implemented between the virtual GPU nodes gpu 0->gpu 1, gpu 2, gpu 1->3, gpu 4, and gpu 2->gpu 5, gpu 6. After parsing the third configuration code, a corresponding virtual communication network can be constructed.
[0202] In one implementation, the first portion of virtual nodes may include virtual GPU nodes in the aforementioned configuration code.
[0203] Step A2: Based on the association relationship, associate the first part of virtual nodes to obtain a virtual communication network.
[0204] In one embodiment, the virtual communication network can be obtained by:
[0205] Based on the connection relationship between the virtual nodes represented by the association relationship, the connection status between the virtual nodes in the first part of the virtual nodes is configured in an association manner, thereby obtaining a virtual communication network.
[0206] As can be seen from the above, the simulation method provided in the embodiment of the present application, after obtaining the first configuration file associated with the target topology structure, associates the first part of the virtual nodes based on the association relationship between the first part of the virtual nodes in the second set contained in the first configuration file to obtain a virtual communication network. In this way, through the above method, the efficiency and controllability of associating the first part of the virtual nodes are improved, thereby improving the flexibility of constructing the virtual communication network.
[0207] Based on the foregoing embodiment, in the simulation method provided in the embodiment of the present application, the virtual nodes in the second set are associated with the target topology structure to obtain a virtual communication network, which can also be achieved by the following steps:
[0208] Step B1: Obtain a second configuration file associated with the target topology structure.
[0209] The second configuration file at least includes capability parameters of a second part of virtual nodes in the second set.
[0210] In one implementation, the second configuration file may be determined after the target topology structure is determined, and may also be flexibly adjusted or configured according to actual simulation requirements.
[0211] In one implementation, the configuration code corresponding to the second configuration file may be included in the configuration code corresponding to the first configuration file, or may exist independently of the first configuration file.
[0212] In one embodiment, the second part of virtual nodes may be the same as the first part of virtual nodes, and may also include other virtual nodes except the first part of virtual nodes; illustratively, the second part of virtual nodes may include virtual GPU nodes, and virtual network card nodes for supporting data transmission of virtual GPU nodes, and may also include virtual channel nodes for realizing data transmission between different virtual network card nodes.
[0213] In one embodiment, the capability parameters may include the performance parameters of the virtual node in terms of computing resources and network resources; the performance parameters may include parameters such as the memory capacity, video memory size, network bandwidth, and data delay of the virtual node; illustratively, the above code <bw>Tags and <delay>Tags can be used to configure performance parameters such as network bandwidth and data delay.
[0214] Exemplarily, the fourth configuration code for configuring the virtual GPU node in the second configuration file may be as follows:
[0215] <!-- Define 4 GPU nodes -->
[0216] <gpus>
[0217] <gpu id="0"gpu_mem="16GB"gpu_bw="100Gbps" / >
[0218] <gpu id="1"gpu_mem="16GB"gpu_bw="100Gbps" / >
[0219] <gpu id="2"gpu_mem="16GB"gpu_bw="100Gbps" / >
[0220] <gpu id="3"gpu_mem="16GB"gpu_bw="100Gbps" / >
[0221] < / gpus>
[0222] In the above code <gpus>Four virtual GPU nodes are defined under the tag. Each virtual GPU node has an id for unique identification. In addition, each virtual GPU node can be configured with the following parameters:
[0223] gpu_mem: used to represent the memory size of the virtual GPU node;
[0224] gpu_bw: used to indicate the network bandwidth of the virtual GPU node.
[0225] Exemplarily, the fifth configuration code for configuring the virtual network card node in the second configuration file may be as follows:
[0226] <!--Define 4 NIC nodes-->
[0227] <nics>
[0228] <nic id="0"nic_delay="10us"nic_bw="100Gbps" / >
[0229] <nic id="1"nic_delay="10us"nic_bw="100Gbps" / >
[0230] <nic id="2"nic_delay="10us"nic_bw="100Gbps" / >
[0231] <nic id="3"nic_delay="10us"nic_bw="100Gbps" / >
[0232] < / nics>
[0233] The configuration parameters contained in the above configuration code will be used to set the performance of virtual GPU nodes and virtual network card nodes, i.e., NIC nodes, in the subsequent simulation; and, by pre-defining the performance parameters of virtual GPU nodes and NICs, the above performance parameters can be directly referenced when building virtual network connections, thereby improving the efficiency of virtual network connections; and, configuring the above performance parameters through *.xml files also provides conditions for adjusting the performance parameters.
[0234] It should be noted that before configuring the capability parameters of the virtual nodes through the second configuration file, the virtual nodes can be created in advance through the following pseudo code:
[0235]
[0236] In the above code, a virtual GPU node class is created, which contains the configuration of parameters such as the graphics card computing power and video memory size of the virtual GPU node. At the same time, a virtual network node class is also created, which contains the configuration of parameters such as the interface bandwidth, delay, and queue buffer for the virtual network node.
[0237] Step B2: configuring the second part of virtual nodes based on the capability parameters to obtain a virtual communication network.
[0238] In one embodiment, the virtual communication network can be obtained by:
[0239] While building the connection relationship between the first part of virtual nodes based on the first configuration file, the second part of virtual nodes are configured based on the capability parameters, thereby obtaining a virtual communication network.
[0240] Specifically, when building a virtual communication network, an instance is created for each virtual GPU node and virtual network card node, and data representing the connection relationship and performance parameters between different virtual nodes are read from the corresponding *.xml file to initialize each virtual node.
[0241] Wherein, gpu_bw is used to represent the computing throughput of the virtual GPU node, which can be determined by converting the computing rate of the GPU into a bandwidth indicator; illustratively, the conversion formula between the computing capacity of the virtual GPU node and gpu_bw can be: gpu_bw = gpu_cores*gpu_freq*
[0242] operations_per_cycle; where gpu_cores is the number of cores (CUDA cores) of the virtual GPU node, gpu_freq is the frequency of the virtual GPU node, and operations_per_cycle is the number of operations per clock cycle, which can be specifically assumed to be 1. For example, if the virtual GPU node includes 3584 CUDA cores, its frequency is 1.5GHz, and there is 1 operation per cycle, then gpu_bw=3584*1.5*1=5376Giga ops / s, and the above data is converted into bandwidth as gpu_bw=5376 / 8=672GB / s, thus realizing the conversion between computing throughput and bandwidth indicators.
[0243] As can be seen from the above, in the simulation method provided by the embodiment of the present application, after obtaining the second configuration file associated with the target topology structure, the second part of the virtual nodes are configured based on the capability parameters of the second part of the virtual nodes in the second set contained in the second configuration file to obtain a virtual communication network. In this way, through the above operation, the flexible and controllable configuration of the capability parameters of the second part of the virtual nodes is achieved, so that the data processing capability of the virtual communication network can match the capability parameters in the second configuration file.
[0244] Based on the foregoing embodiment, in the simulation method provided in the embodiment of the present application, the virtual nodes in the second set are associated with the target topology structure to obtain a virtual communication network, which can also be achieved by the following steps:
[0245] Step C1: Obtain a virtual GPU node and a virtual network card device from the second set.
[0246] Step C2: Establish a point-to-point channel in the NS3 system.
[0247] In one embodiment, a point-to-point channel is used to implement data transmission between different virtual GPU nodes; illustratively, two ends of the point-to-point channel can be connected to two different virtual network card devices respectively, and different virtual network card devices can be associated with corresponding virtual GPU nodes respectively.
[0248] In one implementation, the point-to-point channel may be PointToPointChannel; and when creating the point-to-point channel, parameters such as link data rate, maximum transmission unit (MTU), and queue mode may also be set.
[0249] For example, a point-to-point channel can be created and:
[0250]
[0251] Step C3: Based on the network topology represented by the target topology structure, the virtual GPU nodes and the virtual network card devices are connected through point-to-point channels to obtain a virtual communication network.
[0252] In one embodiment, the virtual communication network can be obtained by:
[0253] The virtual nodes that need to be connected are determined according to the target topology structure, and then the virtual nodes that need to be connected are respectively connected through point-to-point channels to obtain a virtual communication network; the virtual nodes may include virtual GPU nodes and virtual network card nodes.
[0254] Exemplarily, connecting virtual nodes via a point-to-point channel can be implemented by the following code:
[0255]
[0256] for conn in topology.connections:
[0257] node1=topology.nodes[conn.node1]
[0258] node2=topology.nodes[conn.node2]
[0259] #Create a channel
[0260] channel=create_p2p_channel(conn.bw,conn.delay)
[0261] #Connect nodes
[0262] connect_nodes(node1,node2,channel)
[0263] Exemplarily, the above code adopts an object-oriented modular approach to abstract virtual nodes, connections between virtual nodes, and network topology structures between virtual nodes into classes and functions; specifically, the above code first defines two node classes, GPUNode and NetworkNode; wherein GPUNode contains relevant parameters of the virtual GPU node, and NetworkNode contains parameters of the virtual network node, and both of the above nodes will create NS3 network card devices; the create_p2p_channel function is used to create NS3's point-to-point channel and set the bandwidth and delay parameters of the channel, and the connect_nodes function installs the network card devices of the two nodes on the nodes and connects them to the same channel to achieve connection between virtual nodes; the connect_topology function is used to read the topology configuration, obtain virtual node instances, connection information between virtual nodes, and create channels to connect virtual nodes, thereby completing the construction of a virtual communication network.
[0264] Exemplarily, the topology configuration may be obtained through the *.xml file in the aforementioned embodiment.
[0265] As can be seen from the above, the simulation method provided in the embodiment of the present application, after obtaining the virtual GPU node and the virtual network card device from the second set, constructs a point-to-point channel in the NS3 system, and connects the virtual GPU node and the virtual network card device through the point-to-point channel based on the network topology represented by the target topology structure to obtain a virtual communication network. In this way, through the above operations, targeted and flexible connection of the virtual GPU node and the virtual network card device is achieved.
[0266] Based on the foregoing embodiments, in the simulation method provided in the embodiments of the present application, based on at least one virtual operation state of a virtual node in a virtual communication network, simulating the data processing state of at least part of the physical nodes can be achieved by the following steps:
[0267] Step D1: determine the data processing flow included in the target algorithm.
[0268] In one embodiment, the target algorithm may include a distributed data processing algorithm or a communication algorithm based on which distributed graphics processing is implemented between GPU nodes in a GPU cluster; illustratively, the distributed data processing algorithm may include steps and links on how each GPU node performs graphics processing operations, how data is transmitted between each GPU node, and how status is synchronized; the communication algorithm may implement the GPU cluster communication function.
[0269] In one implementation, the target algorithm may be predetermined or flexibly adjusted.
[0270] In one implementation, the data processing flow may include steps and links of how each GPU node performs graphics processing operations, how each GPU node performs data transmission and state synchronization.
[0271] In one embodiment, the data processing flow may further include the coordination state of each GPU node during the graphics processing process; illustratively, the coordination state may include parallel, serial, and parallel and serial cross-execution states.
[0272] Step D2: simulate and control at least one virtual operating state of the virtual node based on the data processing flow to obtain a control result.
[0273] In one embodiment, the virtual operation state may include the state of virtual nodes including virtual GPU nodes and virtual network nodes during the data processing process.
[0274] In one embodiment, the control result may include at least one result of changes in input data, output data, storage data, transmission data, and data storage status of the virtual GPU node and / or virtual network node during the data processing process.
[0275] In one implementation, the control result can be obtained by:
[0276] During the continuous data processing process, the control result is obtained by collecting or capturing the virtual operating status of the virtual node and counting the virtual operating status in the time dimension.
[0277] Step D3: Determine the data processing status based on the control result.
[0278] In one embodiment, the data processing status may be determined by:
[0279] Based on the advancement status of the data processing flow, the data in the control result is counted to obtain a third statistical result, and the third statistical result is determined as the data processing status; illustratively, the third statistical result can characterize the data processing capabilities, data processing results and data processing status changes of virtual nodes including virtual GPU nodes during the advancement of the data processing flow.
[0280] From the above, it can be seen that the simulation method provided in the embodiment of the present application, after determining the data processing flow included in the target algorithm, simulates and controls at least one virtual operating state of the virtual node based on the data processing flow, obtains a control result, and determines the data processing state based on the control result. In this way, through the above process, the correlation between the data processing state and the data processing flow included in the target algorithm can be improved, thereby improving the accuracy of the data processing state; and, in the case of changes in the target algorithm, the data processing flow and control results obtained by the above method can change accordingly, thereby enabling flexible simulation of a variety of target algorithms.
[0281] Based on the foregoing embodiments, in the simulation method provided in the embodiments of the present application, at least one virtual operating state of a virtual node is simulated and controlled based on a data processing flow to obtain a control result, which can be achieved by the following steps:
[0282] Step E1, obtain an interface set.
[0283] The interfaces in the interface set are used to trigger the virtual GPU nodes in the second set to perform at least one target operation; the target operation includes an operation that can be performed by the GPU nodes in the GPU cluster.
[0284] In one embodiment, the interface set may include a set of target operations that can trigger the execution of GPU collective communication; illustratively, the target operations may include Broadcast, Reduce, Allreduce, Scatter, Gather, Allgather, and All-to-all, etc.
[0285] In one implementation, interfaces in the interface set may be predefined in the NS3 system; specifically, the pseudo code corresponding to the interface set may be as follows:
[0286]
[0287] def allreduce(self, tensor):
[0288] pass
[0289] def reducescatter(self,tensor):
[0290] pass
[0291] def allgather(self, tensor):
[0292] pass
[0293] def gather(self,tensor,root):
[0294] pass
[0295] def reduce(self,tensor,root):
[0296] pass
[0297] def barrier(self):
[0298] pass
[0299] Step E2: Screen the interfaces in the interface set based on the data processing flow to obtain the target interface.
[0300] In one implementation, the target interface may be obtained by:
[0301] Based on the operation set included in the data processing flow, the target operation that can be triggered by the interface set is screened to obtain the operation screening result, and then the interface set corresponding to the operation screening result is determined as the target interface.
[0302] Step E3: In the process of calling the target interface to control at least one virtual state, node state data of the virtual nodes in the virtual communication network are collected.
[0303] In one implementation, the node status data may include at least one of the status of virtual node input and output data, the status of storage data, and the occupation and release status of data processing resources.
[0304] In one embodiment, the process of calling the target interface to control at least one virtual state can be implemented by instantiating the extended target interface after the target algorithm expands, rewrites or customizes the target interface to obtain the extended target interface; exemplarily, for Ring and Tree type network topologies, when the target interface is AllReduce and AllGather, the extended target interfaces corresponding to the above two virtual network topologies can include RingAllReduce and RingAllGather, as well as TreeAllReduce and TreeAllGather, respectively.
[0305] Specifically, the extended operation on the target interface in the communication algorithm can be as follows: class RingAllReduce(Communicator):
[0306] def allreduce(self, tensor):
[0307] #1. Divide the tensor into multiple chunks
[0308] chunks = split_tensor(tensor)
[0309] #2. Schedule chunks in parallel by pipeline
[0310] for iin range(self.pipeline_depth):
[0311] #2.1 Get the current and next chunk
[0312] chunk_1 = chunks[i]
[0313] chunk_2=chunks[(i+1)%depth]
[0314] #2.2 Calculate the target node under the ring topology
[0315] dst_1 = (src_1 + 1) % num_gpus
[0316] dst_2 = (src_2 + 1) % num_gpus
[0317] #2.3 Schedule the sending of the current chunk
[0318] schedule_send(chunk_1,src_1,dst_1)
[0319] #2.4 Schedule the sending of the next chunk
[0320] schedule_send(chunk_2,src_2,dst_2)
[0321] #2.5 Wait for the current chunk to complete and schedule reception
[0322] wait_chunk_1()
[0323] schedule_recv(chunk_1,src_1,dst_1)
[0324] #2.6 Wait for the next chunk to complete and schedule reception
[0325] wait_chunk_2()
[0326] schedule_recv(chunk_2,src_2,dst_2)
[0327] #3. Merge all chunk results
[0328] merged = merge_chunks(chunks)
[0329] #4. Return results
[0330] return merged
[0331] The above pseudo code first divides the tensor tensor into multiple chunks for pipeline processing; then in the main loop, the chunks are scheduled in a pipeline parallel manner, the target node under the ring topology is calculated for each chunk, the send event of chunk 1 is registered, and then the send event of chunk 2 is registered to achieve parallel flow of the two chunks, and when waiting for chunk 1 to complete, its receive event is registered at the same time, and when waiting for chunk 2 to complete, its receive event is registered at the same time, so as to optimize the efficiency of the pipeline through asynchronous overlapping sending and receiving; finally, the data of all chunks are merged. Through the above algorithm, the characteristics of the Rng topology are utilized to realize pipeline parallel and overlapping transmission, which can improve data throughput efficiency.
[0332] In one implementation, the node status data may be collected according to a preset collection period, or may be obtained through continuous collection.
[0333] Step E4: Determine the control result based on the node status data.
[0334] In one embodiment, the control result can be determined by:
[0335] The node status data is integrated in the time dimension to obtain the control result.
[0336] The node status data is integrated in the virtual node dimension to obtain the control result.
[0337] From the above, it can be seen that the simulation method provided in the embodiment of the present application, after obtaining the interface set, screens the interfaces in the interface set based on the data processing flow to obtain the target interface, and in the process of calling the target interface to control at least one virtual operating state, collects the node status data of the virtual node in the virtual communication network, and determines the control result based on the node status data, and the interface in the interface set is used to trigger the virtual GPU node to perform at least one target operation. In this way, through the above operation, not only the accurate screening of the interfaces in the interface set is achieved, but also the accuracy and real-time performance of the node status data can be improved, thereby improving the accuracy of the control result.
[0338] Based on the foregoing embodiment, in the simulation method provided in the embodiment of the present application, based on the data processing flow, at least one virtual operating state of the virtual node is simulated and controlled to obtain a control result, which can also be achieved by the following steps:
[0339] Step F1: Determine an event processing mechanism associated with the data processing flow.
[0340] Among them, the event processing mechanism includes a scheduling mechanism and a registration mechanism.
[0341] In one embodiment, the scheduling mechanism may include a process or method in which the NS3 system allocates processor usage rights to threads; illustratively, the scheduling mechanism may include cooperative thread scheduling and preemptive thread scheduling.
[0342] In one embodiment, the scheduling mechanism can be implemented by a scheduler in NS3; wherein the scheduler can be used to simulate an asynchronous event-driven execution process, the scheduler allows multiple events to be scheduled simultaneously, and callbacks are triggered in sequence according to the time sequence to simulate the asynchronous concurrent execution process; for the scheduler, each event can accurately set the execution time, so that the scheduler determines how to trigger the callback according to these times to accurately simulate the progress of the passage of time.
[0343] In one embodiment, the registration mechanism can dynamically add or modify the behavior of the program when the code is running. It stores a set of functions or objects through a registry, which can be called or instantiated at runtime. In practical applications, the registration mechanism provides a flexible way to manage and decouple code, making it more modular and easier to manage.
[0344] In one implementation, the event handling mechanism may be determined by:
[0345] Determine the scheduling mechanism based on the asynchronous state or concurrent state required by the data processing flow, and then determine the registration mechanism based on the state switching requirements between the asynchronous state or concurrent state.
[0346] Step F2: Determine an objective function associated with at least one virtual operating state.
[0347] In one embodiment, the objective function may include at least one function for processing data related to the virtual operating state and constructing conditions related to the virtual operating state.
[0348] In one embodiment, the objective function can be determined by:
[0349] An objective function is constructed according to data or conditions based on which at least one virtual operating state is generated and a method for processing at least one virtual operating state.
[0350] Step F3: register the target function based on the registration mechanism, and monitor the scheduling result of the target function based on the callback mechanism.
[0351] In one implementation, the scheduling result may include whether the target function is scheduled, the data processing operation performed after the target function is scheduled, the data on which the above data processing operation is based, the processing result corresponding to the data processing operation, etc.
[0352] In one implementation, after the target function is registered based on the registration mechanism, when the callback mechanism schedules the target function, at least one kind of data of the target function during the execution of the data processing operation can be collected to obtain the scheduling result.
[0353] Specifically, by registering the send / receive events and the callback functions contained in the corresponding target functions, the data transmission process including slice transmission and pipeline parallel transmission in the communication algorithm can be modeled, and the scheduling results corresponding to the above data transmission process can be obtained.
[0354] The pseudo code implementation of the event scheduling mechanism is as follows:
[0355] #Send schedule
[0356] def schedule_send(chunk,src,dst):
[0357] #Calculate the sending time
[0358] time = chunk.size / bandwidth
[0359] #Create a send event
[0360] event = SendEvent(
[0361] time=time,
[0362] src=src,
[0363] dst=dst,
[0364] chunk=chunk )
[0366] #Register to the scheduler
[0367] scheduler.schedule(event)
[0368] #Receive schedule
[0369] def schedule_recv(chunk,src,dst):
[0370] time = chunk.size / bandwidth
[0371] event = RecvEvent(
[0372] time=time,
[0373] src=src,
[0374] dst=dst,
[0375] chunk=chunk )
[0377] scheduler.schedule(event)
[0378] #Send callback
[0379] def handle_send(event):
[0380] #Send logic
[0381] send(event.chunk)
[0382] #Receive callback
[0383] def handle_recv(event):
[0384] #Receiving logic
[0385] data = receive(event.src)
[0386] store(event.chunk,data)
[0387] #Scheduler
[0388] scheduler = Simulator()
[0389] #Register callback
[0390] scheduler.register_handler(SendEvent,handle_send)
[0391] scheduler.register_handler(RecvEvent,handle_recv)
[0392] #run
[0393] scheduler.run()
[0394] In the above pseudo code, schedule_send and schedule_recv are used to register send and receive events. Event objects are created and put into the scheduler; the event objects may contain metadata including source target and chunk data, and may also include execution time; handle_send and handle_recv may be callback functions, which will be called when the event object is triggered, and the callback functions implement the actual sending and receiving logic.
[0395] In the above code, the Simulator object acts as an event scheduler, which can schedule registered events according to the time sequence. It can also register the event type and the corresponding callback function through register_handler, and call simulator.run() to start the event processing loop. In this way, the asynchronous sending / receiving process can be simulated through the event scheduling mechanism, and new events and callbacks can be flexibly added to expand the simulation logic.
[0396] The event-driven approach described above can precisely control the time schedule, thereby enabling a more accurate simulation of the execution process of the communication algorithm.
[0397] During the execution of the simulated communication algorithm, the communication algorithm will be divided into pipeline processing of multiple data blocks (chunks). When the current chunk is processed, the communication request for the next chunk can be submitted immediately without waiting for the current one to be completely finished, thereby hiding the communication delay and improving the utilization of the pipeline; for example, when calculating on a virtual GPU node, the communication of the next chunk is overlapped. In this way, through the event scheduling mechanism, the next event is registered while waiting for the callback, and the request overlap can be flexibly scheduled, and new requests can continue to be submitted in the callback. Specifically, the pseudo code for implementing overlap optimization through the event scheduling mechanism can be as follows:
[0398] #Chunk processing function
[0399] def process_chunk(chunk):
[0400] #Processing chunks
[0401] pass
[0402] for chunk in chunks:
[0403] #Send chunk
[0404] event1 = send_chunk(chunk)
[0405] #Register sending completion event
[0406] event1.on_finished(schedule_next_chunk)
[0407] #Wait for sending to complete
[0408] wait_for(event1)
[0409] #Receive chunk
[0410] event2 = recv_chunk(chunk)
[0411] #Register to receive completion events
[0412] event2.on_finished(process_chunk)
[0413] #Wait for receiving to complete
[0414] wait_for(event2)
[0415] #Send completion callback
[0416] def schedule_next_chunk(event):
[0417] #Send the next chunk immediately
[0418] next_chunk = get_next_chunk()
[0419] send_event=send_chunk(next_chunk)
[0420] #Recursive registration sending completion event
[0421] send_event.on_finished(schedule_next_chunk)
[0422] In the above pseudo code, for each chunk, the callback functions for sending and receiving completion are registered respectively. In the sending completion callback, the next chunk is sent immediately, thus forming a recursive schedule. In the receiving completion callback, the data of the chunk is processed, and the wait_for function is used to wait for each event to complete. In this way, when the sending of the current chunk is completed, the sending of the next chunk can be initiated immediately, and when the receiving of the current chunk is completed, the processing function is triggered again. In this way, with the help of the event scheduling mechanism, overlapping transmission between chunks is realized, and the waiting time between different chunks is reduced, thereby realizing the overlapping optimization of the pipeline and hiding the communication delay.
[0423] Step F4: Determine a control result based at least on the scheduling result.
[0424] In one implementation, the control result can be obtained by:
[0425] The processing result included in the scheduling result is determined as the control result.
[0426] Integrate the scheduling results associated with each virtual node in the virtual communication network in the time dimension to obtain the control result for the virtual communication network; illustratively, in order to integrate the scheduling results in the time dimension, the time information can be obtained by designing the top-level interface. For example, taking AllReduce as an example, the pseudo code of the AllReduce top-level interface design can be as follows:
[0427] stats=allreduce(input_tensor)
[0428] print(stats.total_time)
[0429] print(stats.startup_time)
[0430] print(stats.bandwidth)
[0431] In the above pseudo code, tensors of various data types are supported as input and output data. The interface function automatically generates and appends a unique tensor name for execution tracking, and returns a statistical object containing time data after the communication algorithm is executed. The statistical data fields may include the following fields:
[0432] total_time: used to indicate the overall execution time
[0433] startup_time: used to indicate the initialization time
[0434] bandwidth: used to indicate the average bandwidth
[0435] Through the above parameters, detailed time statistics of the algorithm execution process can be obtained conveniently and accurately.
[0436] As can be seen from the above, in the simulation method provided by the embodiment of the present application, after determining the event processing mechanism associated with the data processing flow and the target function associated with at least one virtual operating state, the target function is registered based on the registration mechanism in the event processing mechanism, and the scheduling result for the target function is monitored based on the callback mechanism in the event processing mechanism, and then the control result is determined at least based on the scheduling result. In this way, in the above process, the registration mechanism and callback mechanism in NS3 are fully utilized, so that the data processing process of the GPU node in the GPU communication network can be simulated more accurately, thereby improving the accuracy of the scheduling result and the control result.
[0437] Based on the foregoing embodiment, in the simulation method provided in the embodiment of the present application, based on the data processing flow, at least one virtual operating state of the virtual node is simulated and controlled to obtain a control result, which can also be achieved by the following steps:
[0438] Step G1, constructing simulation control logic; in the process of simulating control of at least one virtual operating state based on the data processing flow, statistically analyzing state data of the simulation control based on the simulation control logic to obtain statistical results.
[0439] In one implementation, the simulation control logic may include the number of simulations, the simulated communication algorithm, and the network topology structure targeted by the simulation, and may also include a method for tracking and counting data during the simulation process.
[0440] Exemplarily, the pseudo code corresponding to the simulation control logic may be as follows:
[0441] #Execution control class
[0442] class ExecController:
[0443] def__init__(self,topo,algorithm,dataset):
[0444] self.topo=topo
[0445] self.algorithm = algorithm
[0446] self.dataset = dataset
[0447] def run_experiments(self,num_trials):
[0448]
[0449] In the above pseudo code, the ExecController class encapsulates the configuration of the topology, algorithm, and dataset. The run_experiments method is used to execute multiple rounds of simulation. Each round generates new simulation results, and each simulation result can be independent of each other.
[0450] Through the above pseudo code, it is possible to create a simulation environment, input network topology and communication algorithm, execute the simulation process, and obtain simulation results.
[0451] In one embodiment, the statistical results may include the above-mentioned simulation results, which may be saved to a file through the record_result method, and each round of simulation operation and its corresponding simulation result may have a unique identifier to facilitate subsequent tracking of the simulation results.
[0452] In one embodiment, the network topology structure, communication algorithm and data set based on which the simulation needs to be simulated can be flexibly configured to automatically perform multiple simulations and automatically collect simulation results through simulation control logic. In this way, through the above operations, richer performance data of the virtual GPU node can be obtained, providing a basis for analyzing the performance of the algorithm in different scenarios.
[0453] Step G2: Determine the control result based on the statistical result.
[0454] In one embodiment, the control result can be determined by:
[0455] According to the identification of the simulation process corresponding to the statistical result, the statistical results corresponding to each simulation process are summarized and integrated, and the summarized and integrated result is determined as the control result.
[0456] Exemplarily, both the statistical results and the control results can be realized by the statistical function in the NS3 framework; wherein the pseudo code corresponding to the statistical function can be as follows:
[0457] from ns.stats import TimeSeriesStats
[0458] stats = TimeSeriesStats()
[0459] #Define custom statistical indicators
[0460] stats.register_uint('chunks_sent')
[0461] stats.register_double('gpu_util')
[0462] for chunk in chunks:
[0463] #Run the communication algorithm
[0464] stats = allreduce(chunk)
[0465] # Update statistics
[0466] stats.update('bytes_sent',len(chunk))
[0467] stats.update('gpu_util',get_gpu_util())
[0468] # Output results
[0469] print(stats.get_total('chunks_sent'))
[0470] print(stats.get_mean('gpu_util'))
[0471] print(get_avg_bandwidth())
[0472] #Write the statistical results to a file
[0473] stats.write_time_series('stats.csv')
[0474] #Draw a time series graph
[0475] plot_time_series(stats.get_time_series('throughput'))
[0476] def get_avg_bandwidth():
[0477] total_time=stats.get_total_time()
[0478] total_bytes=stats.get_total('bytes_sent')
[0479] return total_bytes / total_time
[0480] In the above code, we first create the ns-3 TimeSeriesStats statistical object stats. Then we register two custom statistical indicators:
[0481] chunks_sent: used to record the number of chunks sent
[0482] gpu_util: used to record the utilization of virtual GPU nodes
[0483] While the allreduce algorithm is running, the following two statistics can be updated:
[0484] chunks_sent: used to indicate the number of increased chunks
[0485] gpu_util: used to record the current GPU utilization
[0486] Then, the results of each statistical indicator can be obtained through the methods provided by stats:
[0487] get_total(): used to indicate the total amount of indicators obtained
[0488] get_mean(): used to get the average value
[0489] You can also define the get_avg_bandwidth() method to calculate the average bandwidth, that is, calculate the average bandwidth as the number of bytes sent / total time.
[0490] Finally, the statistical results can be written to a comma-separated values (CSV) file in a time series format, and a throughput time series curve can be drawn to visualize the statistical results.
[0491] Exemplarily, during the communication algorithm simulation process, the statistical results can be updated in real time; during the above process, the statistical framework of the NS3 system can be fully utilized, combined with custom statistical indicators, so as to flexibly simplify the above statistical process.
[0492] Figure 2 Another schematic diagram of the simulation method provided in the embodiment of the present application is as follows: Figure 2 As shown, the process may include the following steps:
[0493] Step 201, start.
[0494] Step 202: Configure the XML file.
[0495] Exemplarily, the user may configure the xml file according to the scale of the GPU cluster.
[0496] Exemplarily, the xml file may include the first configuration file and the second configuration file in the aforementioned embodiment.
[0497] Exemplarily, a second set including, for example, 100 virtual GPU nodes may be set through an XML file, and two GPU cards may be configured for each virtual GPU node in the second set, and GPU model parameters may be set through the XML file.
[0498] Exemplarily, the number of virtual network nodes, the number of ports, and network parameters may also be set through the XML file; wherein the network parameters may include bandwidth, delay, and the like.
[0499] Exemplarily, the XML file can also be used to specify the connection method of the network topology structure, whether it is fully connected or hierarchically connected, etc., and can also select the communication algorithm to be simulated. It can also set the data set corresponding to the communication algorithm, set the number of simulation rounds, and the operating control parameters such as the amount of data to be transmitted during each round of simulation.
[0500] Step 203: Determine whether the format is correct.
[0501] Exemplarily, by loading an XML file, it is possible to verify whether the XML file complies with the syntax and format requirements of the XML in the NS3 system. For example, it is possible to verify whether the XML file contains a root node, and it is also possible to verify one by one whether the parameter names and formats contained in the virtual GPU nodes are correct. It is also possible to determine whether the parameter values of the virtual network nodes are within a reasonable range, and it is also possible to verify whether the communication algorithm name is in the supported algorithm set.
[0502] Exemplarily, if the format is judged to be incorrect, a prompt message indicating the incorrect format may be output; the prompt message may include that the number of virtual GPU nodes is incorrect, the network delay exceeds the possible range, and step 202 is re-executed to prompt the user to modify the xml file.
[0503] Exemplarily, if the format is determined to be correct, step 204 may be executed.
[0504] Step 204: Load the XML file.
[0505] Exemplarily, after loading the XML file, the XML file can be parsed to extract each parameter name and parameter value therefrom, and each parameter name and parameter value can be stored in a dictionary object. It can also be parsed into a custom class object, and an instance object can be established based on the XML file, so that subsequent program code can directly apply the above parameters.
[0506] Exemplarily, if the number of the above parameters is large, incremental loading may be selected, that is, only the parameters required for use in the current simulation process are loaded.
[0507] Among them, a dictionary object is a data structure for storing key-value pairs. The key in the data structure is unique and can be associated with at least one value.
[0508] Step 205: Build a topology.
[0509] Exemplarily, the virtual network node in the NS3 system can be instantiated according to the parameters in the xml file; and a point-to-point communication channel connecting the virtual GPU node and the virtual network node can be created according to the network topology in the xml file, and parameters such as bandwidth and delay can be set at the same time.
[0510] Step 206: Event scheduling.
[0511] For example, according to the actual communication algorithm to be simulated, different event classes such as sending events and receiving events can be constructed, and the constructor of the event class can be implemented to encapsulate event-related parameters and write callback processing functions after event triggering. During the communication algorithm simulation process, event objects can be created and scheduling operations can be performed by calling the NS-3 scheduler.
[0512] Step 207: Execute control.
[0513] Exemplarily, the xml file may also include control data for simulation rounds. In this case, the control data may be obtained from the xml file to create an execution control module and encapsulate the control logic to manage the overall simulation process.
[0514] Exemplarily, according to the number of rounds corresponding to the current simulation process, the process from running the simulation to the statistical analysis is controlled, and the simulation results of each round of simulation are recorded. It can also be determined whether the total number of simulation rounds has been reached.
[0515] Step 208: Run the corresponding algorithm to perform simulation according to the configuration.
[0516] Exemplarily, the constructed network topology can be loaded here, and an instance of the selected communication algorithm can be created according to the configuration to generate a data set that this round of simulation depends on; then the instance of the communication algorithm is called to perform a simulation based on the network topology and the data set; in the above process, the simulation process of the communication algorithm can rely on the event scheduler to drive.
[0517] Step 209: Update statistics.
[0518] Exemplarily, during the communication algorithm simulation process, a statistical function may be called to update corresponding statistical variables according to pre-configured statistical indicators to obtain statistical results.
[0519] Exemplarily, the statistical results may be stored in a common statistical management module, or may be stored in a result object corresponding to the current round of simulation, and the statistical results may also support data aggregation of multiple rounds of simulation.
[0520] Step 210: Determine whether there are more simulations.
[0521] Exemplarily, the execution control module can determine whether the current round number reaches the total round number. If it does not reach the total round number, step 208 can be executed to trigger the next round of simulation; if it has reached the total round number, step 211 can be executed; if the total round number changes, error handling is required.
[0522] Step 211: Update statistics.
[0523] For example, the statistical results corresponding to each round of simulation can be summarized, and according to the analysis requirements, the statistical results can be integrated to generate an experimental report, and the experimental report can also be visualized to show the comprehensive simulation results.
[0524] Step 212: Analyze the results.
[0525] Exemplarily, the simulation results corresponding to each round of simulation process can be summarized, and according to the needs of data analysis, the simulation results can be statistically analyzed, a test report can be generated, and a visual analysis can be performed to obtain a global simulation result.
[0526] Step 213, end.
[0527] Through the above process, a controllable, diversified, flexible and automated full-process simulation of GPU cluster communication is achieved in the NS3 system.
[0528] From the above, it can be seen that the simulation method provided in the embodiment of the present application, after constructing the simulation control logic, in the process of simulating and controlling at least one virtual operating state based on the data processing process, statistically simulates the state data of the control based on the simulation control logic, obtains the statistical results, and then determines the control results based on the statistical results. In this way, through the above process, it is possible to achieve accurate and stable control of the process of obtaining the statistical results, thereby improving the accuracy of the control results.
[0529] Through the simulation method provided in the embodiment of the present application, a specific network topology structure and communication model including virtual GPU nodes can be constructed in the NS3 system to simulate the actual communication status of the simulation GPU cluster; and by adopting discrete event drive, large-scale scenario simulation can be achieved, thereby improving the comparative efficiency of different simulation processes; during the simulation process, by adjusting the parameter configuration, a diversified large-scale GPU cluster model can be constructed, thereby providing data support for simulating and testing the extreme performance of the GPU cluster.
[0530] At the same time, since the NS3 system is an open source simulation platform, its models are reusable and the open source code is extensible. Therefore, by simulating the GPU cluster in the NS3 system, the simulation cost can be reduced. On the other hand, the simulation method provided in the embodiment of the present application can also support customized application requirements, thereby meeting the simulation needs of manufacturers, universities, and research institutes that conduct GPU cluster communication research.
[0531] It should be noted that the simulation method provided in the embodiment of the present application can use simulation environments such as OMNeT++ for secondary development to meet simulation requirements, and can also combine the advantages of other simulation platforms to achieve a hybrid simulation solution.
[0532] Through the above-mentioned simulation method, at least the performance bottleneck of GPU cluster communication can be analyzed, the performance of different communication algorithms can be compared, the optimal parameters of the communication algorithm can be sought, larger-scale scenarios can be explored, and the cost of actual debugging can be reduced; among them, by analyzing the performance bottleneck, the execution process of the communication algorithm under different network topologies can be observed in detail, so as to determine the performance bottleneck of different communication algorithms; by comparing the performance of different communication algorithms, different communication algorithms can be screened to obtain the optimal communication algorithm; by finding the optimal parameters of the communication algorithm, the influence of different parameters on the communication algorithm can be evaluated; by exploring larger-scale scenarios, the extreme performance of the GPU cluster can be tested; through pre-executed simulation operations, the cost of modifying and debugging the actual GPU cluster can be reduced, so as to promote the research and development of distributed GPU cluster communication algorithms and improve the system design efficiency of GPU clusters.
[0533] Based on the above embodiments, the present application also provides a simulation device. Figure 3 A schematic diagram of the structure of the simulation device provided in the embodiment of the present application is shown in FIG. Figure 3 As shown, the simulation device 3 may include:
[0534] The acquisition module 301 is used to acquire a first set; wherein the first set includes a set of at least part of the physical nodes in the GPU cluster; the physical nodes include the GPU nodes in the GPU cluster, and the node devices based on which any GPU nodes in the GPU cluster communicate with each other;
[0535] The processing module 302 is used to create a second set corresponding to the first set in the discrete event simulation system; construct a virtual communication network corresponding to the second set in the discrete event simulation system; wherein the second set includes a set of virtual nodes corresponding to at least part of the physical nodes;
[0536] The simulation module 303 is used to simulate the data processing state of at least part of the physical nodes based on at least one virtual operation state of the virtual nodes in the virtual communication network.
[0537] In some embodiments, the discrete event simulation includes an NS3 system; a processing module 302, configured to determine a target topology; and in the NS3 system, associating virtual nodes in the second set based on the target topology to obtain a virtual communication network.
[0538] In some embodiments, the acquisition module 301 is used to acquire a first configuration file associated with the target topology structure; wherein the first configuration file at least includes the association relationship between the first part of virtual nodes in the second set;
[0539] The processing module 302 is used to associate the first part of virtual nodes based on the association relationship to obtain a virtual communication network.
[0540] In some embodiments, the acquisition module 301 is used to acquire a second configuration file associated with the target topology structure; wherein the second configuration file at least includes capability parameters of a second part of virtual nodes in the second set;
[0541] The processing module 302 is used to configure the second part of virtual nodes based on the capability parameters to obtain a virtual communication network.
[0542] In some embodiments, the acquisition module 301 is used to acquire a virtual GPU node and a virtual network card device from the second set;
[0543] The processing module 302 is used to build a point-to-point channel in the NS3 system; based on the network topology represented by the target topology structure, the virtual GPU node and the virtual network card device are connected through the point-to-point channel to obtain a virtual communication network.
[0544] In some embodiments, the processing module 302 is used to determine the data processing flow included in the target algorithm; simulate and control at least one virtual operating state of the virtual node based on the data processing flow to obtain a control result; and determine the data processing state based on the control result.
[0545] In some embodiments, the acquisition module 301 is used to acquire an interface set; wherein the interface in the interface set is used to trigger the virtual GPU node to perform at least one target operation; the target operation includes an operation that can be performed by the GPU node in the GPU cluster;
[0546] The processing module 302 is used to screen the interfaces in the interface set based on the data processing flow to obtain the target interface; in the process of calling the target interface to control at least one virtual operating state, collect the node status data of the virtual node in the virtual communication network; based on the node status data, determine the control result.
[0547] In some embodiments, the processing module 302 is used to determine an event processing mechanism associated with the data processing process; wherein the event processing mechanism includes a scheduling mechanism and a registration mechanism;
[0548] The processing module 302 is used to determine a target function associated with at least one virtual operating state; register the target function based on a registration mechanism, and monitor a scheduling result for the target function based on a callback mechanism; and determine a control result based at least on the scheduling result.
[0549] In some embodiments, the processing module 302 is used to construct a simulation control logic; in the process of simulating the control of at least one virtual operating state based on the data processing flow, the state data of the simulation control is statistically simulated based on the simulation control logic to obtain statistical results; and the control result is determined based on the statistical results.
[0550] Based on the above embodiments, the present application also provides an electronic device, Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application, such as Figure 4 As shown, the electronic device 4 includes a processor 401 and a memory 402; wherein the memory 402 stores a computer program; when the computer program is executed by the processor 401, it can implement any of the above-mentioned simulation methods.
[0551] Based on the foregoing embodiments, the embodiments of the present application further provide a computer-readable storage medium, in which a computer program is stored; when the computer program is executed by a processor of an electronic device, it can implement any of the simulation methods described above.
[0552] Based on the foregoing embodiments, an embodiment of the present application further provides a computer program product, the program product comprising a computer program; when the computer program is executed by a processor of an electronic device, it can implement any of the simulation methods described above.
[0553] The above description of various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other, and for the sake of brevity, they will not be repeated herein.
[0554] The methods disclosed in the various method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0555] The features disclosed in the various product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0556] The features disclosed in the various method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0557] It should be noted that the computer-readable storage medium may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory (Flash Memory), or a hard disk.
[0558] Memory), magnetic surface storage, optical disk, or Compact Disc Read-Only
[0559] Memory, CD-ROM) and other memories; it can also be various electronic devices including one or any combination of the above memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0560] It should be noted that, in this article, the term "includes", "comprising" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0561] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0562] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus necessary general hardware nodes, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0563] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0564] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0565] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0566] The above are only preferred embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.< / gpus> < / delay> < / bw> < / delay> < / bw> < / connection> < / connections>
Claims
1. A simulation method, characterized in that: The method comprises: Acquire a first set; wherein the first set includes a set of at least some physical nodes in a GPU cluster; the physical nodes include GPU nodes in the GPU cluster, and node devices based on which communication between any GPU nodes in the GPU cluster is based; Creating a second set corresponding to the first set in a discrete event simulation system; wherein the second set includes a set of virtual nodes corresponding to the at least part of the physical nodes; constructing a virtual communication network corresponding to the second set in the discrete event simulation system; Based on at least one virtual operating state of the virtual nodes in the virtual communication network, the data processing state of at least part of the physical nodes is simulated.
2. The method according to claim 1, characterized in that The discrete event simulation system includes an NS3 system; and constructing a virtual communication network corresponding to the second set in the discrete event simulation system includes: Determine the target topology; In the NS3 system, the virtual nodes in the second set are associated based on the target topology structure to obtain the virtual communication network.
3. The method according to claim 2, characterized in that The step of associating the virtual nodes in the second set based on the target topology structure to obtain the virtual communication network includes: Acquire a first configuration file associated with the target topology structure; wherein the first configuration file at least includes association relationships between a first part of virtual nodes in the second set; Based on the association relationship, the first part of virtual nodes are associated to obtain the virtual communication network.
4. The method according to claim 2, characterized in that: The step of associating the virtual nodes in the second set based on the target topology structure to obtain the virtual communication network includes: Acquire a second configuration file associated with the target topology structure; wherein the second configuration file at least includes capability parameters of a second part of virtual nodes in the second set; The second part of virtual nodes is configured based on the capability parameters to obtain the virtual communication network.
5. The method according to claim 2, characterized in that: The step of associating the virtual nodes in the second set based on the target topology structure to obtain the virtual communication network includes: Acquire a virtual GPU node and a virtual network card device from the second set; Establishing a point-to-point channel in the NS3 system; Based on the network topology represented by the target topology structure, the virtual GPU node and the virtual network card device are connected through the point-to-point channel to obtain the virtual communication network.
6. The method according to claim 1, characterized in that The simulating the data processing state of at least part of the physical nodes based on at least one virtual operation state of the virtual nodes in the virtual communication network comprises: Determine the data processing flow included in the target algorithm; Performing simulation control on at least one virtual operating state of the virtual node based on the data processing flow to obtain a control result; Based on the control result, the data processing state is determined.
7. The method according to claim 6, characterized in that The simulating control of at least one virtual operating state of the virtual node based on the data processing flow to obtain a control result includes: Acquire an interface set; wherein an interface in the interface set is used to trigger a virtual GPU node in the second set to perform at least one target operation; the target operation includes an operation that can be performed by a GPU node in the GPU cluster; Filtering the interfaces in the interface set based on the data processing flow to obtain a target interface; In the process of calling the target interface to control the at least one virtual operating state, collecting node state data of virtual nodes in the virtual communication network; Based on the node status data, the control result is determined.
8. The method according to claim 6, characterized in that The simulating control of at least one virtual operating state of the virtual node based on the data processing flow to obtain a control result includes: Determining an event processing mechanism associated with the data processing flow; wherein the event processing mechanism includes a scheduling mechanism and a registration mechanism; determining an objective function associated with the at least one virtual operating state; Registering the target function based on the registration mechanism, and monitoring the scheduling result for the target function based on the callback mechanism; The control result is determined based at least on the scheduling result.
9. The method according to claim 6, characterized in that The simulating control of at least one virtual operating state of the virtual node based on the data processing flow to obtain a control result includes: Build simulation control logic; In the process of simulating the control of the at least one virtual operating state based on the data processing flow, counting the state data of the simulated control based on the simulation control logic to obtain a statistical result; The control result is determined based on the statistical result.
10. A simulation device, characterized in that: The simulation device comprises: An acquisition module, configured to acquire a first set; wherein the first set includes a set of at least some physical nodes in a GPU cluster; the physical nodes include GPU nodes in the GPU cluster, and node devices based on which communication between any GPU nodes in the GPU cluster is based; A processing module, configured to create a second set corresponding to the first set in a discrete event simulation system; and construct a virtual communication network corresponding to the second set in the discrete event simulation system; wherein the second set includes a set of virtual nodes corresponding to the at least part of the physical nodes; The simulation module is used to simulate the data processing state of at least part of the physical nodes based on at least one virtual operation state of the virtual nodes in the virtual communication network.
11. An electronic device, characterized in that: The electronic device comprises a processor and a memory; wherein the memory stores a computer program; when the computer program is executed by the processor, the simulation method as claimed in any one of claims 1 to 9 can be implemented.
12. A computer-readable storage medium, characterized in that: The storage medium stores a computer program; when the computer program is executed by a processor of an electronic device, it can implement the simulation method as described in any one of claims 1 to 9.
13. A computer program product, characterized in that The program product comprises a computer program; When the computer program is executed by a processor of an electronic device, the simulation method according to any one of claims 1 to 9 can be implemented.
Citation Information
Cited By
GPU cluster pre-silicon verification method, simulation platform, storage medium and program product
CN121441867A
Gpu cluster silicon pre- validation method, simulation platform, storage medium and program product
CN121441867B