Simulation methods, devices, simulators, media and equipment for machine learning models
By generating a model computation graph and using a simulation engine to simulate the operators of the machine learning model, the problems of high simulation complexity and high cost are solved, and an efficient and low-cost simulation process is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BYTEDANCE TECHNOLOGY CO LTD
- Filing Date
- 2026-04-21
- Publication Date
- 2026-06-02
AI Technical Summary
Simulation of machine learning models suffers from high computational complexity, high cost, and low efficiency, especially when it is difficult to achieve optimal configuration in training and inference performance deployment.
By receiving the model to be simulated and the system configuration information, a model calculation diagram is generated, the operator is simulated using the simulation engine, and configuration optimization operations are performed to obtain the optimal simulation results.
This improves the efficiency and accuracy of simulation while reducing its cost.
Smart Images

Figure CN122133513A_ABST
Abstract
Description
Technical Field
[0001] This article relates to the field of artificial intelligence technology, specifically to a simulation method, apparatus, simulator, medium, and device for a machine learning model. Background Technology
[0002] With the continuous development of artificial intelligence technology, machine learning models are being used more and more, especially simulations based on machine learning models, which have attracted widespread attention. However, simulations of machine learning models typically require enormous computational complexity and storage, leading to difficulties, high costs, and low efficiency in simulating the performance of machine learning models in training and inference. Summary of the Invention
[0003] This content section is provided to briefly introduce the concepts, which will be described in detail in the examples section later. This content section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0004] Firstly, a simulation method for a machine learning model is provided, the method comprising: Receive the simulation model and system configuration information; Obtain the model computation graph, which is generated based on the structural information of the model to be simulated and the system configuration information; The simulation operation is performed to obtain the simulation results, which are obtained by simulating the operators of the model computation graph using the simulation engine. Based on the simulation results, configuration optimization operations are performed to obtain the simulation results.
[0005] Secondly, a simulation device for a machine learning model is provided, the device comprising: The receiving module is configured to receive the model to be simulated and system configuration information; The acquisition module is configured to acquire the model computation graph, which is generated based on the structural information of the model to be simulated and the system configuration information. The simulation module is configured to perform simulation operations and obtain simulation results, which are obtained by simulating the operators of the model computation graph using the simulation engine. The exploration module is configured to perform configuration optimization operations based on the simulation results to obtain simulation results.
[0006] Thirdly, a simulator is provided, comprising: The front end is used to receive the model to be simulated and system configuration information; and to obtain the model calculation graph, which is generated based on the structural information of the model to be simulated and the system configuration information. The backend performs simulation operations and obtains simulation results, which are obtained by simulating the operators of the model computation graph using a simulation engine. An explorer is used to perform configuration optimization operations based on the simulation results to obtain the simulation results.
[0007] Fourthly, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect.
[0008] Fifthly, an electronic device is provided, comprising: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method described in the first aspect.
[0009] A sixth aspect provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in the first aspect.
[0010] The above technical solution, upon receiving the model to be simulated and system configuration information, generates a model computation graph based on the structural information of the model to be simulated and the system configuration information. Based on this, the simulation engine is used to simulate the operators of the obtained model computation graph to obtain simulation results. Configuration optimization operations can be performed on the simulation results to obtain the optimal simulation results. By performing simulation operations based on the model computation graph and exploring the optimal results obtained from the simulation, not only can the efficiency and accuracy of the simulation be improved, but the cost of the simulation can also be reduced.
[0011] Other features and advantages of this article will be described in detail in the following examples section. Attached Figure Description
[0012] The above and other features, advantages, and aspects of this document will become more apparent when viewed in conjunction with the accompanying drawings and the following examples. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings: Figure 1 This is a flowchart illustrating a simulation method for a machine learning model according to an exemplary embodiment.
[0013] Figure 2 This is an example diagram illustrating the architecture of a machine learning model in a simulation method for a machine learning model, according to an exemplary embodiment.
[0014] Figure 3This is an example diagram illustrating the computation and parallel strategies of a machine learning model in a simulation method according to an exemplary embodiment.
[0015] Figure 4 This is an overall framework diagram of the target simulator in a simulation method for a machine learning model, according to an exemplary embodiment.
[0016] Figure 5 This is an example diagram illustrating the acquisition of the model computation graph by the front end in a simulation method for a machine learning model according to an exemplary embodiment.
[0017] Figure 6 This is an example diagram of the backend architecture in a simulation method for a machine learning model, according to an exemplary embodiment.
[0018] Figure 7 This is an example diagram of link congestion in a simulation method of a machine learning model according to an exemplary embodiment.
[0019] Figure 8 This is a block diagram illustrating a simulation apparatus for a machine learning model according to an exemplary embodiment.
[0020] Figure 9 This is an example diagram illustrating an implementation of a simulator according to an exemplary embodiment.
[0021] Figure 10 This is a schematic diagram of the structure of an electronic device according to an exemplary embodiment. Detailed Implementation
[0022] The present invention will now be described in more detail with reference to the accompanying drawings. While certain scenarios are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the scenarios set forth herein; rather, these scenarios are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the accompanying drawings and the scenarios depicted are for illustrative purposes only and are not intended to limit the scope of this document.
[0023] It should be understood that the steps described in the method implementation may be performed in different orders and / or in parallel. Furthermore, the method implementation may include additional steps and / or omit the steps shown. The scope of this document is not limited in this respect.
[0024] The term "comprising" and its variations as used herein can be open-ended, meaning "including but not limited to". The term "based on" can mean "at least partially based on". The term "one case" means "at least one case"; the term "another case" means "at least one additional case"; the term "some cases" means "at least some cases". Definitions of other terms will be given in the following description.
[0025] It should be noted that the concepts of "first" and "second" mentioned here are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependencies.
[0026] It should be noted that the terms "one" and "more" used here are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0027] The names of messages or information exchanged between the multiple devices in the implementation are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0028] It is understandable that before using this article, users should be informed of the type, scope of use, and usage scenarios of the personal information involved in accordance with relevant laws and regulations, and their authorization should be obtained.
[0029] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media performing the operations described herein.
[0030] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0031] It is understood that the above notification and user authorization process is merely illustrative and does not limit the implementation method described in this article. Other methods that comply with relevant laws and regulations may also be applied to the implementation method described in this article.
[0032] At the same time, it is understood that the data involved in this article (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0033] In related technologies, the large design space complexity of parallel strategies, system optimization, and hardware configuration makes it extremely challenging to deploy and train optimal machine learning models for inference. Machine learning deployment typically requires determining a complete set of system-level parameters, which directly impacts the performance and cost of subsequent training and inference.
[0034] For example, system-level parameter combinations may include parallel strategy configuration, hardware resource configuration, scheduling and execution parameters, the number of tasks processed by each machine, and whether to buffer intermediate results. Therefore, how to accurately and quickly perform performance simulation of machine learning models is a technical problem that urgently needs to be solved.
[0035] To address the aforementioned issues, a simulation method, apparatus, simulator, medium, and equipment product for machine learning models are proposed. The simulation method generates a model computation graph based on the received structural and system configuration information of the model to be simulated. Simulation operations are then performed according to this computation graph. Since the simulation operations are performed on each operator within the computation graph, this paper enables precise simulation of the time of each operation at the operator level. Subsequent configuration optimization operations based on the simulation results can accurately and efficiently obtain higher-precision simulation results, such as precise simulation of operator-level delays and timeline trajectories.
[0036] Figure 1 This is a flowchart illustrating a simulation method for a machine learning model according to an exemplary embodiment. Figure 1 As shown, this paper provides a simulation method for a machine learning model. Specifically, this method can be executed using a machine learning model simulation device, which can be implemented in software and / or hardware. For example... Figure 1 As shown, the method may include the following steps.
[0037] In step S110, the simulation model and system configuration information are received.
[0038] This paper can be applied to target simulators, which can be unified, modular, and fine-grained simulators for training and inference of machine learning models. For example, machine learning models can include large language models (LLMs), visual models, audio models, and multimodal models.
[0039] Here, the machine learning model architecture can be a Transformer architecture based on a pure decoder, which can be as follows: Figure 2 As shown, based on Figure 2As can be seen, a machine learning model can include an embedding layer, stacked feature processing modules, and a language model head layer (LM Head Layer).
[0040] Here, the stacked feature processing modules can include multiple feature processing modules (Transformer modules), such as Transformer Block 1, Transformer Block 2, and Transformer Block n. Each feature processing module can include an attention mechanism and a feedforward network, where the attention mechanism can include Q (query), K (key), V (value) generation and multi-head attention. In other words, the machine learning model can be composed of stacked modules, each of which can consist of an attention layer and a feedforward network. To improve parameter efficiency, the feedforward network of the machine learning model can adopt a Mixture of Experts (MoE) structure.
[0041] Operations on machine learning models can include, for example Figure 3 The training task 301 and inference task 302 shown differ fundamentally in their execution flows for machine learning model training and inference. For example, the training process may include iterative forward propagation, back propagation, and optimizer updates, while the inference process can utilize key-value (KV) caching and accelerate forward propagation. In other words, the inference process may include pre-filling, decoding, and key-value caching. Therefore, in simulating machine learning models, different task types will result in different simulation flows. The specific steps for executing simulation operations based on task type will be described in detail later and will not be elaborated upon here.
[0042] Optionally, due to the scalability of machine learning models, simulations require both load-balanced parallel strategies and extensive performance tuning to optimize their execution efficiency on large-scale clusters. Therefore, parallel strategies are crucial for achieving efficient execution of machine learning models.
[0043] As an example, parallel strategies can include, for example, Figure 3 The examples shown are Tensor Parallelism (TP), Data Parallelism (DP), Expert Parallelism (EP), Pipeline Parallelism (PP), and Sequence Parallelism (SP).
[0044] Tensor parallelism 303 is used to partition a single weight tensor across multiple graphics processing units (GPUs), which can support larger hidden layer dimensions; data parallelism 304 can be used to partition optimizer states, gradients, and model parameters themselves across devices to optimize memory usage and communication volume of a single GPU; expert parallelism 305 can be applied to hybrid expert models, and different expert subnetworks can be distributed to GPUs based on expert parallelism strategies; pipeline parallelism 306 can be used to divide the Transformer module (feature processing module) into consecutive stages to assign each stage to a different GPU, during which micro-batch data can flow between stages to ensure the utilization of all GPUs; sequence parallelism 307 can split the input token sequence across devices in the normalization layer to accelerate normalization computation.
[0045] In deploying machine learning models, multiple parallelism strategies mentioned above are often used in combination. For example, tensor parallelism can be used for graphics processing units within a node, while pipelined parallelism can be used between nodes. Each combination of parallelism strategies places unique demands on computing resources and network connections within and between nodes.
[0046] As described above, a target simulator can be a simulator for training and inference of machine learning models. In other words, a target simulator can be a unified, modular, and fine-grained simulator used to accurately predict the performance of machine learning models. Through this target simulator, the optimal configuration scheme of the model to be simulated can be obtained, thus effectively improving system throughput. The framework of a target simulator can be as follows: Figure 4 As shown, based on Figure 4 It can be seen that the target simulator (simulation platform) may include a front end 410, a back end 420, and an explorer 430.
[0047] Among them, the front end 410 can be a graph-based front end, which can be used to parse the model graph and perform a series of compiler-like processing operations, such as tracing operators, injecting parallel strategies, scheduling execution, and obtaining analysis results; the back end 420 can be a multi-engine back end, which can execute operator-level simulation through interchangeable worker nodes. For example, the back end 420 can support multiple simulation modes, and it can also include an overlapping processor, which can be used to capture overlapping phenomena and estimate the performance degradation caused by resource contention (overlap); the explorer 430 can be a design explorer (design space explorer / parameter searcher), which can be used to explore the design space through suboptimal configuration pruning, that is, to perform configuration optimization and determine the target solution, which balances cost optimization and performance optimization.
[0048] In summary, graph-based front-ends can be used to construct forward and backward computation graphs and perform optimization and analysis; multi-engine back-ends can be used for computational and communication operators in accurate simulation models; and design explorers can be used to search for optimal configurations through guided pruning.
[0049] In one scenario, the input data for the target simulator can be the model to be simulated and system configuration information. For example, the target simulator can receive input data such as the model to be simulated and system configuration information through a native model interface; that is, the target simulator can be compatible with multiple native models. Therefore, the front end of the target simulator can receive the model to be simulated and system configuration information.
[0050] Furthermore, the model to be simulated can be a native model of the complete Transformer architecture without operator fusion, quantization compression, or structural rewriting. For example, the model to be simulated can be a native model without performance optimizations such as model distillation, parameter pruning, or precision conversion. For instance, the model to be simulated can be a high-speed inference model (Virtual Large Language Model) or a custom deep learning model.
[0051] It should be noted that the front-end receiving the model to be simulated can be receiving the structural information of the model to be simulated. This structural information can be an inherent attribute of the model to be simulated, which can be used by the front-end to construct the original computational graph. For example, the model to be simulated can include at least one of the following: model meta-information, network layers and module composition, core operators and computational logic, and tensor shape and dimension information.
[0052] The model metadata can include the model type, size, and source; the network layers and module composition can include the number of Transformer layers, the structure of the encoder and decoder, the attention module, the feedforward network, and the normalization method; the core operators and operational logic can include computation operators, the topological relationships of residual connections / jump connections, and forward / backward computation logic; the tensor shape and dimension information can include the hidden layer dimension, the head dimension, the intermediate dimension or vocabulary size of the feedforward network (FFN), and the tensor data type.
[0053] Optionally, the system configuration information may be the external environment and strategy parameters for simulation operation, which can be used for parallel injection and optimization at the front end, accurate performance simulation at the back end, and configuration optimization in the design space. For example, the system configuration information may include at least one of the following: hardware cluster specifications, multi-dimensional parallel strategy, runtime hyperparameters, computation graph optimization rules, communication link bandwidth, and design space exploration constraints.
[0054] Here, hardware cluster configuration information may include graphics processor model, number of graphics processors per node or number of nodes, and graphics processor memory size or memory bandwidth, etc.; parallel strategy configuration information may include tensor parallelism partitioning dimension and partitioning number, pipeline parallelism stage number / layer partitioning method, data parallelism grouping size, expert parallelism expert allocation method, sequence parallelism partitioning length, and hybrid parallelism combination strategy, etc.; simulation running hyperparameters include batch size, micro-batch size, sequence length, inference / training mode, and maximum cache length, etc.; computation graph optimization configuration includes operator fusion switch, operator reordering strategy, deceleration model selection, and overlapping execution strategy, etc.; communication link and network configuration includes intra-node communication, inter-node communication, communication domain partitioning, and communication algorithm selection; configuration optimization operations (design space exploration constraint configuration) may include adjustable parameter search range, performance constraints, pruning rules, and optimization objectives, etc., where optimization objectives may include minimum extension, maximum throughput, and maximum GPU utilization, etc.
[0055] In step S120, the model calculation diagram is obtained. The model calculation diagram is generated based on the structural information and system configuration information of the model to be simulated.
[0056] As an optional approach, upon receiving the model to be simulated and system configuration information, the front-end can generate a model computation graph based on the structural information of the model and the system configuration information. Therefore, the front-end can adopt a computation graph-oriented design, meaning it can serve as the interface between user input and the operator-level back-end, such as... Figure 5 As shown, the front end can be used to convert input data into simulation graphs and obtain optimization results through multiple graph operations.
[0057] In one scenario, the front-end can generate a first computational graph based on the structural and system configuration information of the model to be simulated. Based on this, optimization operations are performed on the first computational graph to obtain a second computational graph. These optimization operations can be used to perform structural and semantic optimization on the first computational graph. Additionally, the front-end can acquire a parallel strategy and insert communication operators into the second computational graph according to this strategy to obtain the model computational graph.
[0058] The model to be simulated can come from mainstream frameworks. After obtaining input data such as the model to be simulated and system configuration information, the front end can use a graph tracer (graph tracing tool) to generate a first computational graph, which automatically converts the native model (the model to be simulated) into an intermediate computational graph. For example, the graph tracer can first perform symbolic tracing on the model to be simulated to capture its operator-level structure. Based on this, the traced modules are optimized and downgraded to generate a unified forward computational graph (model computational graph) suitable for simulation.
[0059] Since the architecture of a machine learning model can consist of multiple repeating feature processing blocks (Transformer blocks), the front end can extract and simulate a single Transformer decoder module. This not only improves the simulation speed of the machine learning model, but also maintains the architectural fidelity and numerical accuracy.
[0060] It's important to note that before performing single-module simulation, the model's architecture type can be obtained first. Then, module simulations can be performed specifically based on this architecture type. For example, single-module simulation can be executed for symmetric models or loads that do not require pipelined parallelism. Conversely, for asymmetric models or loads requiring pipelined parallelism, different layers can be traced as independent computation graphs according to the model definition, and explicitly scheduled by device rank to ensure accurate modeling of the device ranks and inter-stage dependencies for all pipelined parallelism. In other words, a single computation graph can be generated for symmetric models, while independent computation graphs can be generated for asymmetric models.
[0061] Furthermore, to support decoupled services, the front end can trace pre-filling and decoding operations as independent computation graphs, thereby simulating heterogeneous execution scenarios where different stages are mapped to different hardware clusters. In other words, when generating the first computation graph, the front end can obtain the task type of the simulation task. If the simulation task is an inference task, it can generate the computation graph corresponding to pre-filling and the computation graph corresponding to the decoding operation, etc. That is, different task types will result in different generated computation graphs.
[0062] In another scenario, during the process of generating the first computational graph based on the structural information of the model to be simulated and the system configuration information, the task type of the simulation task can be obtained. This task type can belong to the system configuration information; that is, the system configuration information can include task types. Based on this, the first computational graph corresponding to the model to be simulated is generated according to the task type.
[0063] As an example, in response to the task type being training, a forward graph and a backward graph can be generated based on the structural information of the model to be simulated and the system configuration information. That is, when the simulation task is a training task, the generated first computation graph can include a forward graph and a backward graph.
[0064] As can be seen, for training tasks requiring backpropagation of the computation graph, the graph tracker can automatically generate forward and backward graphs. For example, a virtual loss function and a virtual input tensor can be added to the forward graph to support dynamic output shapes. After tracking is complete, multiple exploitation phases process the joint graph to optimize its structure and semantics. This processing may include adding tensor metadata, renaming inputs, renaming weight nodes, and removing redundant operations. Redundant operations may include view transformations, view splitting, and other operations that do not change the tensor's shape or data type.
[0065] like Figure 5 As shown, after obtaining the joint graph including both forward and reverse information, a partitioner (using the default partitioning rule) can divide the joint graph into independent forward and reverse graphs. Furthermore, the reverse graph can be further optimized by performing post-processing operations. For example, these post-processing operations may include decomposing automatic functionalization operations, removing self-cloning operations, and eliminating dead code. This operation graph tracker efficiently and automatically generates a concise and optimized reverse graph for a given forward computation.
[0066] As another example, in response to the task type being inference, a forward graph can be generated based on the structural information and system configuration information of the model to be simulated. That is, when the simulation task is an inference task, the generated first computational graph can include the forward graph. In other words, after the original forward graph is generated by the graph tracker for the inference task, the original forward graph can be directly used as the first computational graph.
[0067] Machine learning models typically combine various optimization methods and parallel strategies during training and inference processes. To support these methods and strategies and prepare for future adaptation to new technologies, this paper proposes a compiler-like design. In other words, optimization methods and parallel strategies are abstracted into graph operations (optimization and parallelization) that directly act on the computation graph. Based on these graph operations, a stage can be added or removed, enabling or disabling the corresponding optimization method in the simulation task. Different stages can be freely combined to achieve joint optimization.
[0068] Understandably, after obtaining the first computational graph, the front end can perform optimization and parallelization processes on it to obtain the model computational graph, i.e., the front end can enter... Figure 5The diagram illustrates optimization stage 501 and parallel processing stage 502. Optimization processing (optimization operations) can include operator rewriting, operator fusion, recalculation, or reordering. The optimization operations described in this paper can be operator-level optimizations, such as operator rewriting or operator fusion, based on the operator level. Here, optimization operations can be implemented using a flexible graph operation framework of matching and replacement.
[0069] For example, during the optimization phase, the front-end can traverse the first computation graph to identify whether it contains a target node pattern that matches predefined matching conditions. These predefined matching conditions can be user-inputted conditions, which may include implementation conditions for high-performance operators or fusion operators. Once a target node pattern matching the predefined conditions is found, the back-end can modify the node type and attributes to achieve the desired optimization. For instance, if a difference is detected between the original training graph and the actual computation graph, the node pattern of the computation graph can be identified, and then it can be determined whether the node pattern is the target node pattern. If so, multiple operators can be merged into one, or inefficient operator forms can be replaced.
[0070] It's important to note that the optimization rules for optimization operations can be preset or continuously updated as needed. New optimization rules can be easily added by defining custom matching patterns and corresponding transformation operations. Furthermore, for quantization optimization, the front-end can modify node precision by adding a quantization stage, or it can directly track the quantized model. This further improves the simulator's usability and coverage. In other words, upon receiving input data, the front-end can determine whether quantization is required for the current simulation. If quantization is required, the front-end can modify node precision by adding a quantization stage during optimization operations, or it can track the quantized model.
[0071] After optimizing the first computation graph to obtain the second computation graph, the front end can sequentially apply the corresponding processing stages to the second computation graph to implement a multi-parallel strategy, thereby providing a flexible combination of hybrid parallel strategies. As described above, machine learning models deployed in multi-GPU systems typically employ a combination of various parallel methods. These parallel strategies introduce aggregate communication operations to coordinate computation and data transmission between devices. Therefore, when obtaining the second computation graph, the parallel strategy can be acquired, and communication operators can be inserted into the second computation graph based on this strategy to obtain the model computation graph.
[0072] For example, the front end can insert communication operators into the optimized computation graph using a series of dedicated parallel processing strategies. Beforehand, the parallel strategies can be categorized, and the corresponding communication operator insertion operation can be performed based on the type of parallel strategy.
[0073] For the first type of parallel strategy, the tensor and its associated computational operators can be partitioned across multiple graphics processing units (GPUs). This paper can trace the computation graph, adjust the tensor shape of the partitioned operators, and insert corresponding communication operators before and after each partitioned operation according to the selected strategy. Here, the first type of parallel strategy can be a partition-based parallel strategy, which can include tensor parallelism, sequence parallelism, and expert parallelism. Furthermore, communication operators can include full reduction, full collection, and reduced scattering, etc.
[0074] For the second type of parallel strategy, the training or inference process can be divided into multiple stages, with explicit dependencies between devices. This process can be achieved through a front-end scheduling pattern generator. Based on the second type of parallel strategy, dependencies between stages can be constructed, and send or receive communication operators can be inserted to model data transmission between stages. Furthermore, the simulator can simultaneously support 1F1B scheduling (One Forward, One BackwardScheduling) and DualPipe scheduling (Dual Pipeline Scheduling) that overlaps computation and communication. This ensures that the generated dependency graph and communication events can be recorded for subsequent timeline analysis and visualization. Here, the second type of parallel strategy can be a pipelined parallel strategy.
[0075] For the third type of parallel strategy, the training dataset can be partitioned across various graphics processors, allowing each device to independently compute gradients. The simulator can support the simulation of parallel frameworks for distributed data of multiple native models. For example, for the first native model, gradients can be synchronized through ensemble communication operations, meaning the front end can explicitly model this process by inserting communication operators. For the second native model, the front end can support multiple schemes through additional parameter configuration and / or optimizer state slicing, optimizer state synchronization, and prefetch analysis. In other words, when inserting / injecting communication operators into the second computation graph, the insertion strategies adopted for different types of native models may also differ.
[0076] This simulator can output simulation results at multiple dimensions and granularities, including coarse-grained system-level metrics and fine-grained operator-level metrics. Coarse-grained system-level metrics can include model floating-point operation utilization, parallel communication overhead, and memory usage; fine-grained operator-level metrics can include operator latency, efficiency, and analyzer-style tracking data. The analyzer allows for customized analysis workflows, enabling flexible and convenient addition of new metrics. The analyzer can be used to perform tasks such as... Figure 5The analysis passes are shown below. Multi-granularity analysis can generate both system-level metrics and output fine-grained, style-specific tracing information, thus providing rich information for performance debugging.
[0077] Understandably, after obtaining the model computation graph, the front end can send it to the back end to instruct the back end to perform simulation on the model computation graph and obtain simulation results. Subsequently, the front end can receive the simulation results sent by the back end, and based on this, the front end can enter the analysis phase, that is, it can perform analysis operations on the simulation results sent by the back end to obtain analysis results.
[0078] For example, simulation results can include multiple performance metrics. During analysis, the front end can obtain the category of each performance metric in the simulation results and determine the corresponding analysis strategy. Based on this, the simulation results are analyzed according to the analysis strategy to obtain the analysis results. In other words, the front end can support the calculation of both dependency-independent metrics (performance metrics of the first category) and dependency-aware metrics (performance metrics of the second category) through a unified graph-based processing method.
[0079] As an example, in response to a performance metric categorized as "Category 1," a first analysis strategy corresponding to Category 1 is obtained. Here, the performance metric in Category 1 can be one that has no dependency on other operators, and the first analysis strategy can be used to obtain system-level results based on graph traversal. It is evident that for metrics such as model floating-point operation utilization that do not require operator dependencies (Category 1 performance metrics), the front-end analyzer can directly perform operations on the computation graph to calculate the metric for each node through graph traversal and aggregate these metrics to obtain system-level results.
[0080] As another example, in response to a performance metric categorized as "second category," a second analysis strategy corresponding to the second category is obtained. The performance metric for the second category can be based on simulation results, and the second analysis strategy can be used to obtain operator-level results based on the dependencies between operators in the model computation graph. It is evident that for dependency-aware metrics such as detailed timeline trajectories (performance metrics of the second category), the front-end analyzer can interact with the back-end, receiving simulation results from the back-end to obtain the start and end times of each operator. Then, by combining the dependencies between operators, the timeline is corrected, thus generating an accurate execution timeline. Based on this corrected timeline, analyzer-style tracking data and latency-related metrics can be further derived, such as module latency, end-to-end latency, and floating-point operation utilization.
[0081] It should be noted that since both optimization and analysis operations are applied to the computation graph, the simulator can support the alternating execution of optimization and analysis operations within the same simulation process, which can improve the efficiency, consistency, and flexibility of the simulation to a certain extent.
[0082] Furthermore, accurate peak memory estimation is crucial for reliable simulation of machine learning models. Underestimating peak memory can lead to memory overflow errors, while overestimating it can result in suboptimal configurations. Peak memory consumption in machine learning models primarily originates from optimizer states, model weights and gradients, activation tensors, and temporary buffer allocation. This paper's graph-based front-end can analyze the lifecycle of each tensor during backpropagation, enabling accurate modeling of activation and temporary memory, thereby improving the fidelity of memory simulation. In other words, when the simulation task is a training task, the front-end can analyze the lifecycle of each tensor during backpropagation during its analysis operations. Additionally, the front-end can traverse the backpropagation graph to determine the allocation, usage, and release timing of each intermediate tensor during backpropagation iterations.
[0083] In summary, the front-end, through graph generation, optimization, parallel processing, and analysis stages, can reproduce the real memory behavior of graphics processing units (GPUs). For example, it can reproduce the timing of temporary tensor reuse and release, which directly affects peak memory usage. Furthermore, fine-grained graph analysis allows for the evaluation of memory-sensitive training configurations, and operator-level lifecycle modeling enables high-fidelity GPU memory simulation to a certain extent. For instance, the front-end output data may include operator-level computation graphs (model computation graphs), initial execution trajectories, operator dependencies, tensor shapes, and memory information.
[0084] This paper proposes a modular processing stage that enables plug-and-play analysis and optimization stages, meaning that the core architecture does not need to be reconstructed, and various parallel strategies and new optimization methods can be modeled to a certain extent.
[0085] In step S130, a simulation operation is performed to obtain simulation results. The simulation results are obtained by simulating the operators of the model computation graph using a simulation engine.
[0086] As described above, after obtaining the model computation graph, the front-end can send it to the back-end of the target simulator. The back-end of the target simulator is responsible for simulating individual computational or communication operators. In other words, the back-end can receive the model computation graph, perform simulation operations on it, and obtain simulation results. These simulation results can be obtained by simulating the operators in the model computation graph using a simulation engine.
[0087] As an alternative approach, during the simulation operation, the backend can determine the simulation engine corresponding to the operator. Based on this, the simulation operation is performed according to the determined simulation engine; that is, different operators may correspond to different simulation engines. For example, the backend architecture can be as follows: Figure 6 As shown, based on Figure 6 It is known that the backend simulation engine may include at least one of the following: a performance analysis engine 601 (Profiling Engine), a prediction engine 602 (Prediction Engine), and a performance modeling engine 603 (Analytical Engine).
[0088] Understandably, the model computation graph can include multiple operators, each of which can be simulated using a performance analysis engine, a prediction engine, or a performance modeling engine. Furthermore, the backend fusion engine 604 supports collaborative execution of multiple engines, meaning different operators can be simulated by different engines.
[0089] The performance analysis engine performs cache matching based on operator type and tensor shape. Its input data includes operator type and tensor shape. The engine executes operators on the target hardware and simulates each operator by analyzing its runtime, thus providing accurate latency data. For graphics processor loads, the target simulator automatically generates analysis tasks for each operator, then distributes these tasks to the internal graphics processor cluster for execution, recording latency upon task completion.
[0090] In addition, based on Figure 6 It is known that the backend of the target simulator can integrate an analysis database, which can be used to cache the measured results of common operators and input shapes. Upon receiving a simulation request, the performance analysis engine can first query this database. If it is determined that a matching operator and shape combination exists in the analysis database, the cached delayed data can be directly reused, thus significantly improving simulation efficiency.
[0091] In one scenario, in response to the operator being a first-category operator, the performance analysis engine can be used as a simulation engine. The first-category operator can be a regular operator or a simple operator, and the backend analysis database stores measured results matching the first-category operator. Therefore, when it is determined that an operator is a first-category operator, the measured results matching the first-category operator can be retrieved from the analysis database, and these measured results can be used as the first simulation result.
[0092] Optionally, the prediction engine is used to predict operator results based on operator type and tensor shape. That is, the prediction engine may include a lightweight machine learning model that can directly estimate operator latency based on operator type and tensor shape. Here, the lightweight machine learning model can be a machine learning model trained based on historical measured data. It is evident that the prediction engine can achieve fast latency estimation, and is particularly suitable for analyzing unknown input shapes not covered by the database. In other words, the prediction engine can achieve good generalization and estimation for complex unknown loads without relying on purely analytical approximations.
[0093] In the target simulator, each type of operator can correspond to a customized random forest predictor (lightweight machine learning model) trained based on the analysis database data. This avoids real-time hardware execution and can maintain high accuracy under different operator shapes, thereby significantly improving the speed of large-scale simulation workloads.
[0094] In one scenario, in response to an operator being classified as a second category operator, the prediction engine is used as a simulation engine. That is, when it is determined that no measured results matching the operator are stored in the analysis database, the operator can be classified as a second category operator. In other words, the second category operator can be an operator for which no corresponding measured results are stored in the analysis database, and the type of the operator is an operator type that can be covered by the lightweight machine learning model.
[0095] Optionally, the performance modeling engine is used to obtain operator results based on a mathematical analytical model. This performance modeling engine comprehensively analyzes the execution time of the operator by combining the mathematical model with the operator's computation, memory requirements, and hardware capabilities. In other words, it can calculate the operator's performance based on hardware theory and mathematical models through formula derivation. Here, the mathematical analytical model may include a roofline model or a link center model, etc.
[0096] Furthermore, when performing simulation operations based on the performance modeling engine, the types of operators can be further identified, including computational operators and communication operators. For computational operators, a roofline model can be used to calculate the computation time and memory access time on the target hardware, and the larger of the two is taken as the final computation time for that operator. In this process, the floating-point operation capability and memory bandwidth of the hardware can be pre-configured according to the simulation hardware, while the computational floating-point operation quantity / memory access quantity of the operator can be calculated in real time based on its input shape.
[0097] Optionally, a hierarchical link-centric model can be used for communication operators to ensure cross-platform portability and accuracy. The link-centric model can model the cluster topology using calibrated per-hop latency and effective bandwidth obtained from performance analysis. Here, the target simulator supports the application of both ring and tree-based aggregated communication algorithms in various topologies. The performance modeling engine can decompose high-level aggregated communication operations (communication operators) into physical link-level data transmission. For each link, the total link latency can be obtained by aggregating calibrated handshake latency and transmission latency calculated based on data volume and effective bandwidth. This fine-grained approach enables the target simulator to accurately assess network congestion based on bandwidth sharing and topology constraints.
[0098] In one scenario, in response to an operator being of the third category, the performance modeling engine is used as the simulation engine. That is, the third category of operators are operators other than those of the first and second categories. In other words, when the analysis database does not store the actual test results corresponding to the operator, and it is determined that the operator cannot be processed by the lightweight machine learning model, the performance modeling engine can be used to simulate the operator.
[0099] It should be noted that the backend can include at least one of the aforementioned simulation engines; that is, the backend can be configured with one or more simulation engines. When only one simulation engine is configured, the corresponding simulation operation can be executed based on that engine. During this process, if it is determined that the operator and the simulation engine are incompatible, the operator can be tested to obtain actual simulation data. For example, when a simulation request is received, if it is determined through detection that the analysis database does not store measured data matching the operator to be simulated, then the operator can be tested to obtain simulation data.
[0100] Furthermore, if multiple simulation engines are configured in the backend, a simulation engine matching the current operator to be simulated can be selected through engine fusion. For example, if measured performance data matching the type, tensor size, and hardware configuration of the current operator to be simulated is found in the analysis database, the performance analysis engine can be used for simulation; if it is determined that no measured data matching the operator to be simulated exists in the analysis database, and the type and feature dimensions of the current operator to be simulated are within the coverage of the prediction engine's built-in model, the prediction engine can be used for simulation; when there is no matching measured data, the operator is not covered by a lightweight machine learning model (prediction model), or pure theoretical analytical calculations are required, the performance modeling engine can be used as a fallback simulation solution, that is, the current operator to be simulated is simulated based on the performance modeling engine.
[0101] The aforementioned fusion engine can be used to integrate multiple simulation backends within a single execution flow. This means the fusion engine can achieve an adaptive balance between simulation speed and accuracy, while maintaining compatibility with new models and operators. For example, for operators lacking analysis and prediction data, simulation can be performed using the performance modeling engine; for other operators, simulation can be performed using either the performance analysis engine or the prediction engine.
[0102] This paper proposes a priority fallback mechanism to enable engine selection. Each engine can maintain a registry of supported operators. The fusion engine can dynamically select the highest priority backend available for each operator and fall back to a lower priority engine (performance modeling engine) when necessary. This not only ensures comprehensive coverage of heterogeneous loads but also does not affect the overall simulation fidelity and scalability.
[0103] In summary, the target simulator's backend can integrate multiple backend engines, and may include a fusion engine. Based on this fusion engine, various different simulation engines (backend engines) can be combined and used, thus achieving an optimal balance between simulation speed and accuracy. In other words, this paper achieves an optimal balance between simulation speed and accuracy by fusing three backend engine modes: performance analysis engine, prediction engine, and performance modeling engine.
[0104] As an alternative approach, after determining the simulation engine corresponding to the operator to be simulated, simulation operations can be performed based on that simulation engine to obtain a first simulation result. Based on this, the first simulation result can be corrected to obtain a second simulation result. This correction can be performed on operators with overlapping relationships, and the correction process can be an overlap processing procedure, which can be... Figure 6 The illustrated computational overlap processor 605 is primarily used to perform communication analysis operations to obtain communication latency and link utilization. Its input data may include communication type, size, and cluster information. Furthermore, the computational overlap processor 605 may include multiple processors, which may be graphics processors.
[0105] During the training or inference process of machine learning models, communication operators are often executed in overlap with computation operators or other communication operators to improve overall performance. Therefore, in the simulation process, this paper can perform fine processing on this phenomenon to generate fine-grained and accurate trajectories.
[0106] For example, the target simulator can be configured with a coarse-grained first reduction model for each overlapping operator. This first reduction model can be a proportional deceleration model, based on which a deceleration factor can be applied to the overlapping portion of two operators. The reduction factor can be trained using analytical data from the hardware cluster. Thus, the first reduction model is used to scale the execution time of computation and communication by a preset fixed ratio when computation and communication overlap, thereby simply and quickly correcting the performance degradation caused by resource contention.
[0107] Optionally, a fine-grained second reduction model can be configured for scenarios where communication overlaps. This second reduction model can be a bandwidth-aware deceleration model. It is evident that the second reduction model can be used to dynamically calculate the actual available bandwidth and corresponding deceleration factor based on factors such as peak link bandwidth, number of concurrent communication tasks, data volume, and hardware topology, thereby achieving accurate performance correction based on congestion awareness.
[0108] Understandably, during the correction of the first simulation results, the overlap type of the operators can be obtained. Different overlap types correspond to different reduction models. For example, when the overlap type of the operators is the first overlap type, a reduction factor can be applied to the overlapping portion of the two operators based on the first reduction model. Here, the first overlap type can be an overlap between computation and communication. Similarly, when the overlap type of the operators is the second overlap type, a reduction factor can be applied to the overlapping portion of the two operators based on the second reduction model. Here, the second overlap type can be an overlap between communication operations.
[0109] For example, when computation operators and communication operators overlap, independent deceleration factors can be assigned to each computation operator and communication operator; when communication operators overlap, a single deceleration factor can be assigned to the two communication operators, meaning that the two communication operators can share the same deceleration factor, and the deceleration factor can be applied to the overlapping portion between the two communication operators.
[0110] It should be noted that when using a performance modeling engine to simulate overlapping communication operations, i.e., when simulating overlapping communication operators, a fine-grained bandwidth-aware deceleration model (second reduction model) can be enabled. For example... Figure 7 As shown, the deceleration degree of each operator can be determined by both the effective bandwidth of the cluster and the link congestion situation. For each part of the overlapping operator, the backend can check the link congestion situation of each interconnection layer and calculate the deceleration degree based on the contention ratio of the effective bandwidth, thereby simulating the underlying network packet-level congestion control. Figure 7In the diagram, the first timeline 701 is the timeline of the original communication operation, and the second timeline 702 is the timeline after considering congestion-aware simulation. By comparing the two timelines, it can be seen that using congestion-aware simulation can greatly ensure the accuracy of the simulation, that is, it can improve the overall performance of the simulation.
[0111] It should be noted that the reduction of operators with overlapping relationships can be performed after the simulation engine performs the simulation operation, or it can be performed during the simulation engine's simulation operation. For example, the performance modeling engine can reduce operators with overlapping relationships during the simulation to correct the simulation results. There is no explicit restriction on when to reduce overlapping operators, and it can be selected according to the actual situation.
[0112] In step S140, a configuration optimization operation is performed based on the simulation results to obtain the simulation results.
[0113] As an alternative approach, once simulation results are obtained, configuration optimization operations can be performed based on these results to obtain the final simulation results. These configuration optimization operations can be design space exploration operations, which can be performed by... Figure 4 The Explorer 430 implementation is shown. The design space can refer to the set of parameters constituting all feasible deployment configuration schemes when the simulation model (machine learning model) is deployed for distributed training or inference on the target hardware cluster. Here, feasible deployment configuration schemes include, but are not limited to, adjustable deployment parameters such as: parallel strategy combinations, number of computing devices, tensor parallelism, pipeline parallelism, data parallelism, batch size, sequence length, operator optimization strategies, and memory allocation strategies.
[0114] Optionally, configuration optimization (design space exploration) can be a process within the design space that, based on multi-granularity performance indicators from front-end feedback, automatically iterates, evaluates, prunes inferior solutions, and selects the best among various deployment configuration schemes according to preset constraints and optimization goals.
[0115] Furthermore, to reduce the engineering costs of analyzing simulation results during configuration optimization, the target simulator can incorporate a native design space search function, i.e., an integrated explorer (design explorer). This explorer can start from the target model and task. The target simulator can receive a complete design space input by the user, which can include different selection strategies for the number of graphics processors and parallelism scale. To maximize search efficiency, the explorer can support rule-based pruning of the search space. Users can predefine known inefficient scenarios in the simulator; based on this, the explorer can directly skip the corresponding subspace simulation, thus achieving pruning.
[0116] For example, the input data of the explorer can be the computation graph structure (model computation graph) obtained by the front end, the simulation capability of the back end, user constraint information, and adjustable parameter range, etc., and its output data can be the performance results of all valid configurations, the set of optimal configurations, performance bottleneck analysis, and the best deployment configuration, etc.
[0117] The simulation results presented in this paper can include high-level summary information and fine-grained tracking information. High-level summary information can include floating-point operation counts, model floating-point operation rate, memory usage, and estimated power consumption / thermal design power. Fine-grained tracking information can include single-layer tracking information and 3D tracking information. Single-layer tracking information can include a single-layer timeline, while 3D tracking information can include full-scale 3D timeline tracking information corresponding to the full range of 3D multi-GPU configurations. Additionally, the simulation results may include optimal configuration information from design exploration.
[0118] In summary, the input data for the target simulator can include the model to be simulated (the native model) and system configuration information describing the GPU (Graphics Processing Unit) / NPU (Neural Processing Unit) specifications, interconnect topology, and target parallel scheme. Based on this input data, the target simulator can output high-level summary metrics, which may include at least one piece of information such as floating-point operation count, model floating-point operation utilization, memory usage, power consumption, or thermal design power estimation.
[0119] It should be noted that when the fine-grained mode is detected to be enabled, the output data of the target simulator can also include the model's execution trajectory (execution record). This execution trajectory can be determined according to the style of the model to be simulated; that is, different styles of the model to be simulated will correspond to different output execution trajectories. For example, the execution trajectory can include a single-layer timeline and the full 3D multi-graphics processor trajectory.
[0120] The aforementioned target simulator is used to evaluate the performance of machine learning model training and inference on large-scale systems. It is versatile and scalable, meaning that the target simulator can achieve accurate end-to-end simulation on different graphics processor models. In addition, the target simulator can scale to ultra-large-scale configurations of nearly ten thousand graphics processors under different cluster sizes for training tasks, and supports optimal combinations of all parallel strategies such as data parallelism, pipelined parallelism, expert parallelism, sequence parallelism, and tensor parallelism.
[0121] This paper, upon receiving the model to be simulated and system configuration information, generates a model computation graph based on the structural information of the model and the system configuration information. Based on this, a simulation engine is used to simulate the operators of the obtained model computation graph to obtain simulation results. Configuration optimization operations can be performed on these simulation results to obtain the optimal simulation results. By performing simulation operations based on the model computation graph and exploring the optimality of the simulation results, not only can the efficiency and accuracy of the simulation be improved, but the cost of the simulation can also be reduced.
[0122] Figure 8 This is a schematic diagram of the structure of a simulation device for a machine learning model according to an exemplary embodiment. Figure 8 As shown, this paper provides a simulation device 800 for a machine learning model, which may include a receiving module 810, an acquisition module 820, a simulation module 830, and an exploration module 840.
[0123] The receiving module 810 is configured to receive the model to be simulated and system configuration information; The acquisition module 820 is configured to acquire a model calculation graph, which is generated based on the structural information of the model to be simulated and the system configuration information. The simulation module 830 is configured to perform simulation operations and obtain simulation results, which are obtained by simulating the operators of the model computation graph using a simulation engine. The exploration module 840 is configured to perform configuration optimization operations based on the simulation results to obtain simulation results.
[0124] In some implementations, the acquisition module 810 includes: The generation submodule is configured to generate a first computational graph based on the structural information of the model to be simulated and the system configuration information. An optimization submodule is configured to perform optimization operations on the first computation graph to obtain a second computation graph. The optimization operations are used to perform structural and semantic optimization on the first computation graph. The insertion submodule is configured to acquire a parallel strategy and insert communication operators into the second computation graph according to the parallel strategy to obtain the model computation graph.
[0125] In some implementations, the system configuration information includes a task type, and the generation submodule is further configured to generate a forward graph and a backward graph based on the structural information of the model to be simulated and the system configuration information in response to the task type being a training type.
[0126] In some implementations, the system configuration information includes a task type, and the generation submodule is further configured to generate a forward graph based on the structural information of the model to be simulated and the system configuration information in response to the task type being a reasoning type.
[0127] In some implementations, the simulation results include multiple performance metrics, and the exploration module 840 is further configured to acquire the categories of the performance metrics, determine the analysis strategies corresponding to the categories, analyze the simulation results according to the analysis strategies, obtain analysis results, and perform configuration optimization operations based on the analysis results.
[0128] In some implementations, the exploration module 840 is further configured to obtain a first analysis strategy corresponding to the first category in response to the category of the performance metric being a first category, wherein the performance metric of the first category is a performance metric that has no dependency on other operators, and the first analysis strategy is used to obtain system-level results based on graph traversal.
[0129] In some implementations, the exploration module 840 is further configured to, in response to the performance metric being classified as a second category, obtain a second analysis strategy corresponding to the second category, wherein the performance metric for the second category is a performance metric obtained based on the simulation results, and the second analysis strategy is used to obtain operator-level results based on the dependencies between operators in the model computation graph.
[0130] In some implementations, the simulation module 830 is further configured to determine the simulation engine corresponding to the operator and perform simulation operations based on the simulation engine.
[0131] In some implementations, the simulation engine includes at least one of the following: A performance analysis engine is used to perform cache matching based on operator type and tensor shape; A prediction engine is used to predict operator results based on operator type and tensor shape; A performance modeling engine, which is used to obtain operator results based on a mathematical analytical model.
[0132] In some embodiments, the simulation module 830 is further configured to perform simulation operations according to the simulation engine to obtain a first simulation result; and to correct the first simulation result to obtain a second simulation result, wherein the correction is performed on operators with overlapping relationships.
[0133] This paper, upon receiving the model to be simulated and system configuration information, generates a model computation graph based on the structural information of the model and the system configuration information. Based on this, a simulation engine is used to simulate the operators of the obtained model computation graph to obtain simulation results. Configuration optimization operations can be performed on these simulation results to obtain the optimal simulation results. By performing simulation operations based on the model computation graph and exploring the optimality of the simulation results, not only can the efficiency and accuracy of the simulation be improved, but the cost of the simulation can also be reduced.
[0134] This paper provides a simulator that includes a front-end, a back-end, and an explorer. The front-end receives the model to be simulated and system configuration information, and acquires the model computation graph, which is generated based on the structural information of the model to be simulated and the system configuration information. The back-end executes simulation operations to obtain simulation results, which are obtained by simulating the operators of the model computation graph using a simulation engine. The explorer performs configuration optimization operations based on the simulation results to obtain simulation results.
[0135] The simulator can be the aforementioned target simulator (simulation platform), and the implementation process of the target simulator can be as follows: Figure 9 As shown, based on Figure 9 As can be seen, the target simulator can perform simulations at the operator level based on the computation graph, and it can support adaptation to computing cards and communication modules that are aware of cluster topology. The input data of the target simulator can include the native model (the model to be simulated) and system configuration information. It can simulate not only training tasks but also inference tasks. Furthermore, the output data of the target simulator can include optimal configuration information, expected performance, and execution timelines.
[0136] The front end of the target simulator can be used to parse the model diagram and perform a series of compiler-like processing operations; the back end can perform operator-level simulation through interchangeable working nodes; the explorer can be used to explore the design space through suboptimal configuration pruning, that is, to achieve configuration optimization and determine the target solution, which balances cost optimization and performance optimization.
[0137] The target simulator proposed in this paper achieves optimal end-to-end time simulation accuracy. It supports both training and inference tasks, allowing for simultaneous training and inference. It enables direct use of native models, provides an operator-level analysis perspective, and flexibly integrates hardware backends and parallel strategies, thereby achieving precise configuration optimization and performance analysis in large-scale scenarios. For example, in training tasks, the target simulator can use an operation graph-based simulation approach, employing a precise backend model to accurately simulate the time of each operation at the operator level, while also adapting to the operational logic of different frameworks. In inference tasks, because analysis-based simulation can be used in both computation and communication processes, it can accurately predict the first token output time, thus obtaining higher-precision prediction results.
[0138] The target simulator achieves high simulation accuracy not only at the end-to-end level but also at the operator-level granularity. Furthermore, it accurately simulates operator-level latency and timeline trajectories, closely matching actual hardware execution behavior. Simultaneously, the target simulator reliably captures the actual memory allocation during the training of complex hybrid expert models, faithfully reproducing the memory behavior of the graphics processing unit.
[0139] The aforementioned simulator views machine learning model simulation as a compiler-like transformation process. At each stage, it progressively optimizes the representation of the model, schedule, and system to balance simulation speed, accuracy, and scalability, achieving more comprehensive coverage and higher accuracy. Furthermore, the simulator proposed in this paper not only significantly reduces the cost of finding optimal configurations for machine learning model training and service deployment but also lowers the domain knowledge requirements for engineers, thus reducing simulation costs.
[0140] The following is for reference. Figure 10 The diagram illustrates a structural schematic of an electronic device 1000 suitable for implementing the above-described technical solution. Terminal devices may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Personal Computers), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs (Televisions), desktop computers, etc. Figure 10 The electronic device shown is merely an example and should not be construed as limiting its functionality or scope of use.
[0141] like Figure 10As shown, the electronic device 1000 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1008 into a random access memory (RAM) 1003. The RAM 1003 also stores various programs and data required for the operation of the electronic device 1000. The processing unit 1001, the ROM 1002, and the RAM 1003 are interconnected via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0142] Typically, the following devices can be connected to the input / output interface 1005: input devices 1006 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 1007 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1008 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows electronic device 1000 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 10 An electronic device 1000 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0143] In particular, depending on certain circumstances, the processes described in the flowchart above can be implemented as computer software programs. For example, a computer program product is provided, comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. This computer program can be downloaded and installed from a network via communication device 1009, or installed from storage device 1008, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the above-described methods.
[0144] It should be noted that the aforementioned computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM, or flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In one case, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In another case, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (Radio Frequency), etc., or any suitable combination thereof.
[0145] In some implementations, clients can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (LANs), wide area networks (WANs), the internet (e.g., the Internet), and peer-to-peer networks (e.g., ad-hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0146] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0147] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: receive a model to be simulated and system configuration information; acquire a model computation graph, the model computation graph being generated based on the structural information of the model to be simulated and the system configuration information; perform a simulation operation to obtain a simulation result, the simulation result being obtained by simulating the operators of the model computation graph using a simulation engine; and perform a configuration optimization operation based on the simulation result to obtain a simulation result.
[0148] Computer program code for performing the above operations can be written in one or more programming languages or a combination thereof. These programming languages include, but are not limited to, object-oriented programming languages, as well as conventional procedural programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0149] The flowcharts and block diagrams in the accompanying figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products under various scenarios. In this respect, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the figures. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0150] The modules mentioned above can be implemented in software or hardware. In some cases, the name of a module does not necessarily limit the module itself; for example, the acquisition module can also be described as "acquiring a model computation graph, which is generated based on the structural information of the model to be simulated and the system configuration information."
[0151] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field-Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application-Specific Standard Parts (ASSPs), Systems on Chips (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0152] In this context, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0153] The above description is merely illustrative and explains the technical principles employed. Those skilled in the art should understand that the scope of this document is not limited to technical solutions formed by specific combinations of the above-described technical features, but also includes other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features provided herein that have similar functions.
[0154] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. Multitasking and parallel processing may be advantageous in certain contexts. Similarly, while some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of this paper. Certain features described in the context of a single example can also be implemented in combination in a single example. Conversely, various features described in the context of a single example can also be implemented individually or in any suitable sub-combination in multiple examples.
[0155] Although this document has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims. Regarding the aforementioned apparatus, the specific manner in which the various modules perform their operations has already been described in detail in the section concerning the method, and will not be elaborated upon here.
Claims
1. A simulation method for a machine learning model, comprising: Receive the model to be simulated and system configuration information; Obtain the model computation graph, which is generated based on the structural information of the model to be simulated and the system configuration information; The simulation operation is performed to obtain the simulation results, which are obtained by simulating the operators of the model computation graph using the simulation engine. Based on the simulation results, configuration optimization operations are performed to obtain the simulation results.
2. The method according to claim 1, wherein obtaining the model computation graph includes: A first computational graph is generated based on the structural information of the model to be simulated and the system configuration information. An optimization operation is performed on the first computation graph to obtain a second computation graph. The optimization operation is used to perform structural and semantic optimization on the first computation graph. Obtain the parallel strategy, and insert the communication operator into the second computation graph according to the parallel strategy to obtain the model computation graph.
3. The method according to claim 2, wherein the system configuration information includes task type, and the step of generating a first computational graph based on the structural information of the model to be simulated and the system configuration information includes: In response to the task type being training type, a forward graph and a backward graph are generated based on the structural information of the model to be simulated and the system configuration information.
4. The method according to claim 2, wherein the system configuration information includes task type, and the step of generating a first computational graph based on the structural information of the model to be simulated and the system configuration information includes: In response to the task type being inference type, a forward graph is generated based on the structural information of the model to be simulated and the system configuration information.
5. The method according to claim 1, wherein the simulation results include multiple performance metrics, and the step of performing configuration optimization based on the simulation results includes: Obtain the category of the performance metric and determine the analysis strategy corresponding to the category; The simulation results are analyzed according to the analysis strategy to obtain the analysis results; Based on the analysis results, configuration optimization operations will be performed.
6. The method according to claim 5, wherein determining the analysis strategy corresponding to the category includes: In response to the performance metric being classified as a first category, a first analysis strategy corresponding to the first category is obtained. The performance metric of the first category is a performance metric that has no dependency on other operators. The first analysis strategy is used to obtain system-level results based on graph traversal.
7. The method according to claim 5, wherein determining the analysis strategy corresponding to the category includes: In response to the performance metric being classified as a second category, a second analysis strategy corresponding to the second category is obtained. The performance metric for the second category is a performance metric obtained based on the simulation results. The second analysis strategy is used to obtain operator-level results based on the dependencies between operators in the model computation graph.
8. The method according to claim 1, wherein performing the simulation operation comprises: Determine the simulation engine corresponding to the operator; The simulation operation is performed according to the simulation engine.
9. The method according to claim 8, wherein the simulation engine comprises at least one of the following: A performance analysis engine is used to perform cache matching based on operator type and tensor shape; A prediction engine is used to predict operator results based on operator type and tensor shape; A performance modeling engine, which is used to obtain operator results based on a mathematical analytical model.
10. The method according to claim 8, wherein performing the simulation operation according to the simulation engine comprises: The simulation operation is performed according to the simulation engine to obtain the first simulation result; The first simulation result is corrected to obtain the second simulation result, and the correction is performed on operators with overlapping relationships.
11. A simulation device for a machine learning model, comprising: The receiving module is configured to receive the model to be simulated and system configuration information; The acquisition module is configured to acquire the model computation graph, which is generated based on the structural information of the model to be simulated and the system configuration information. The simulation module is configured to perform simulation operations and obtain simulation results, which are obtained by simulating the operators of the model computation graph using the simulation engine. The exploration module is configured to perform configuration optimization operations based on the simulation results to obtain simulation results.
12. An emulator, comprising: The front end is used to receive the simulation model and system configuration information; Obtain the model computation graph, which is generated based on the structural information of the model to be simulated and the system configuration information; The backend is used to perform simulation operations and obtain simulation results, which are obtained by simulating the operators of the model computation graph using the simulation engine. An explorer is used to perform configuration optimization operations based on the simulation results to obtain the simulation results.
13. A computer-readable medium having a computer program stored thereon, wherein, When executed by a processing device, the computer program performs the steps of the method described in any one of claims 1-10.
14. An electronic device, comprising: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-10.
15. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-10.