A switching decision determination method for a multi-level simulation model and a related device
By acquiring real-time metrics of the module to be simulated and combining them with an operator feature knowledge base to identify the scenario type, the target simulation model is determined. This solves the problem of balancing simulation speed and accuracy in simulation verification, and enables flexible switching and a balance between efficiency and accuracy during the simulation process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING AIJIE KEXIN TECHNOLOGY CO LTD
- Filing Date
- 2026-04-01
- Publication Date
- 2026-07-24
AI Technical Summary
In existing technologies, it is impossible to flexibly switch between different simulation models during the simulation verification process, making it difficult to maximize simulation speed while ensuring simulation accuracy, and thus difficult to achieve an effective balance between simulation efficiency and accuracy.
By acquiring real-time metrics of the module to be simulated, and combining them with the operator feature knowledge base to identify the scene type, a target simulation model that matches the scene type result in terms of simulation accuracy is determined, and simulation models with different simulation accuracies are flexibly switched during the simulation process.
It achieves a balance between simulation speed and simulation accuracy during the simulation process, thereby improving simulation efficiency and accuracy.
Smart Images

Figure CN122452298A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of simulation, and more specifically, to a method and apparatus for determining switching decisions for multi-level simulation models. Background Technology
[0002] With the rapid development of artificial intelligence (AI) technology, the design complexity of AI chips is increasing daily, posing a severe challenge to system-level simulation verification technology. In the AI chip design process, simulation verification is a crucial step in ensuring the correctness of chip functionality and evaluating performance and power consumption. Simulation verification requires simulation using simulation models; different simulation models have varying levels of accuracy and speed, with higher accuracy models typically having lower simulation speeds.
[0003] In existing technologies, a single simulation accuracy model is typically used before simulation begins, rather than switching between different simulation models. During simulation, the system cannot intelligently switch between different simulation models based on the characteristics of the workload at different execution stages. This makes it impossible to maximize simulation speed while ensuring simulation accuracy, and consequently, it is difficult to achieve an effective balance between simulation efficiency and accuracy.
[0004] Therefore, how to provide a technical solution that can flexibly switch between simulation models of different simulation accuracies during the simulation process in order to balance simulation speed and simulation accuracy has become a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] This application provides at least one method and related apparatus for determining switching decisions for multi-level simulation models. It acquires real-time indicators for the module to be simulated, and combines them with an operator feature knowledge base to identify the scene type. It then determines the target simulation model that matches the scene type result in terms of simulation accuracy, thereby enabling the determination of simulation decisions that allow for flexible switching between simulation models with different simulation accuracies during the simulation process, so as to balance simulation speed and simulation accuracy.
[0006] Firstly, this application provides a method for determining switching decisions in a multi-level simulation model, wherein the simulation accuracy of the multi-level simulation model increases progressively. The method includes: In the current simulation scenario, the workload monitor is used to obtain real-time metrics for the module to be simulated corresponding to the original simulation model. The intelligent switching decision-maker determines the scenario type result corresponding to the current simulation scenario based on real-time indicators and through a pre-set operator feature knowledge base; the operator feature knowledge base includes the correspondence between theoretical indicators and scenario types; The intelligent switching decision-maker determines the target simulation model from the multi-level simulation models that matches the simulation accuracy of the scenario type results; If the original simulation model is different from the target simulation model, the intelligent switching decision-maker determines the switching decision corresponding to the current simulation scenario based on the target simulation model so as to switch the original simulation model to the target simulation model by referring to the switching decision.
[0007] Secondly, this application also provides a switching decision determination device for a multi-level simulation model, wherein the simulation accuracy of the multi-level simulation model increases progressively. The device includes: The workload monitor is used to obtain real-time metrics for the module to be simulated corresponding to the original simulation model in the current simulation scenario. The intelligent switching decision-maker is used to determine the scenario type result corresponding to the current simulation scenario based on real-time indicators and a pre-set operator feature knowledge base; the operator feature knowledge base includes the correspondence between theoretical indicators and scenario types. The intelligent switching decision-maker is also used to determine the target simulation model from the multi-level simulation models that matches the simulation accuracy of the scenario type results; The intelligent switching decision-maker is also used to determine the switching decision corresponding to the current simulation scenario based on the target simulation model if the original simulation model is different from the target simulation model, so as to switch the original simulation model to the target simulation model by referring to the switching decision.
[0008] Thirdly, this application also provides an electronic device, including: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory via the bus, and when the machine-readable instructions are executed by the processor, a switching decision determination method for a multi-level simulation model provided in this application is executed.
[0009] Fourthly, this application also provides a computer-readable storage medium storing a computer program, which is executed by a processor to perform a switching decision determination method for a multi-level simulation model provided in this application.
[0010] Fifthly, this application also provides a computer program product, including a computer program that is executed by a processor to perform a switching decision determination method for a multi-level simulation model provided in this application.
[0011] In summary, this application provides a method and related apparatus for determining switching decisions for multi-level simulation models, where the simulation accuracy of the multi-level simulation models increases progressively. The method includes: acquiring real-time metrics for the module to be simulated corresponding to the original simulation model through a workload monitor in the current simulation scenario; an intelligent switching decision-maker determining the scenario type result corresponding to the current simulation scenario based on the real-time metrics and a pre-set operator feature knowledge base; the operator feature knowledge base including the correspondence between theoretical metrics and scenario types; the intelligent switching decision-maker determining a target simulation model from the multi-level simulation models that matches the scenario type result in terms of simulation accuracy; if the original simulation model differs from the target simulation model, the intelligent switching decision-maker determines the switching decision corresponding to the current simulation scenario based on the target simulation model so as to switch the original simulation model to the target simulation model by referring to the switching decision. Through the above method, real-time metrics for the module to be simulated are obtained, and scenario type identification is performed in conjunction with the operator feature knowledge base to determine the target simulation model that matches the scenario type result in terms of simulation accuracy. This enables the determination of simulation decisions for flexibly switching simulation models with different simulation accuracies during the simulation process, thus balancing simulation speed and simulation accuracy.
[0012] Other advantages of this application will be explained in more detail in conjunction with the following description and figures.
[0013] It should be understood that the above description is merely an overview of the technical solution of this application, so as to provide a general understanding of the technical means of this application and to implement it in accordance with the contents of the specification. In order to make the above and other objects, features and advantages of this application more apparent and understandable, specific embodiments of this application are illustrated below. Attached Figure Description
[0014] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly described below. The accompanying drawings are incorporated in and constitute a part of this specification. These drawings illustrate embodiments conforming to this application and are used together with the specification to explain the technical solutions of this application. It should be understood that the drawings only illustrate certain embodiments of this application and should not be considered as a limitation on the scope of protection. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort. Furthermore, the same reference numerals denote the same components throughout the drawings. In the drawings: Figure 1 A flowchart illustrating a method for determining switching decisions for a multi-level simulation model, provided in an embodiment of this application; Figure 2 A schematic diagram of a switching decision determination device for a multi-level simulation model provided in an embodiment of this application; Figure 3This is a general architecture diagram provided for an embodiment of this application. Detailed Implementation
[0015] Exemplary embodiments of this application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of this application and to fully convey the scope of this application to those skilled in the art.
[0016] In the description of embodiments of this application, it should be understood that terms such as “comprising” or “having” are intended to indicate the presence of the disclosed features, figures, steps, behaviors, components, portions or combinations thereof in this specification, and do not exclude the possibility of the presence of one or more other features, figures, steps, behaviors, components, portions or combinations thereof.
[0017] Unless otherwise stated, " / " means "or". For example, A / B can mean A or B. In this article, "and / or" is merely a way of describing the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can mean: A alone, A and B at the same time, and B alone.
[0018] The terms "first," "second," etc., are used only for ease of description to distinguish identical or similar technical features and should not be construed as indicating or implying the relative importance or number of these technical features. Therefore, a feature defined by "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of embodiments of this application, unless otherwise stated, the term "multiple" means two or more.
[0019] With the rapid development of artificial intelligence technology, the design complexity of AI chips is increasing daily, posing a severe challenge to system-level simulation verification technology. In the AI chip design process, simulation verification is a crucial step in ensuring the correctness of chip functionality and evaluating performance and power consumption. Simulation verification requires simulation using simulation models; different simulation models have varying levels of accuracy and speed, with higher accuracy models typically having lower simulation speeds.
[0020] Currently, mainstream simulation methods can be mainly divided into the following four categories of simulation models with progressively increasing accuracy, based on different levels of precision: The first type is the functional model, which has the fastest simulation speed and can quickly handle large-scale workloads, making it suitable for early functional verification and software stack development. However, this type of model ignores pipeline details, hardware resource contention, and specific signal timing, resulting in insufficient accuracy in characterizing the microarchitectural behavior of key modules and making it difficult to deeply analyze performance bottlenecks.
[0021] The second category is architecture-level models, which can describe data paths, topological connections, and approximate transmission delays, achieving a preliminary balance between simulation accuracy and speed. However, due to their lack of precise characterization of internal pipeline states and instruction-level behavior, they struggle to provide accurate assessments of the efficiency of complex parallel computing.
[0022] The third type is the microarchitecture-level model, which can reflect the behavior of pipeline registers, state machines, and control logic. It has high simulation accuracy and is suitable for local optimization of key modules. However, for complex scenarios such as large-scale matrix operations, its simulation overhead is still too high, making it difficult to support the overall evaluation of long-term or large-scale networks.
[0023] The fourth type is the periodic accurate model, which can faithfully simulate the signal flip of each clock cycle. It has the highest simulation accuracy, but the simulation speed is extremely slow and the resource consumption is huge. It is almost impossible to use it for training or inference performance evaluation of a complete AI network.
[0024] In existing technologies, a single simulation accuracy model is typically used before simulation begins, rather than switching between different simulation models. During simulation, the system cannot intelligently switch between different simulation models based on the characteristics of the workload at different execution stages. Consequently, it cannot maximize simulation speed while ensuring simulation accuracy, making it difficult to achieve an effective balance between simulation efficiency and accuracy.
[0025] In view of this, this application provides a method and related apparatus for determining switching decisions for multi-level simulation models. It obtains real-time indicators for the module to be simulated, and combines them with an operator feature knowledge base to identify the scene type. It then determines the target simulation model that matches the scene type result in terms of simulation accuracy, thereby enabling the determination of simulation decisions that allow for flexible switching between simulation models with different simulation accuracies during the simulation process, so as to balance simulation speed and simulation accuracy.
[0026] The switching decision determination method for multi-level simulation models provided in this application can be implemented using a computer device, which can be a terminal device or a server. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. Terminal devices include, but are not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle terminals, and aircraft. The terminal device and the server can be directly or indirectly connected via wired or wireless communication, and this application does not impose any limitations on this connection.
[0027] The following describes a method for determining the switching decision of a multi-level simulation model provided in this application through method embodiments, such as... Figure 1 As shown, Figure 1 This application provides a flowchart of a method for determining switching decisions for a multi-level simulation model. The aforementioned computer device can be a server. The simulation accuracy of the multi-level simulation model increases progressively. In practical applications, the multi-level simulation model can include functional models, architecture-level models, micro-architecture-level models, and periodically accurate models with progressively increasing simulation accuracy. The method includes: S101. In the current simulation scenario, obtain the real-time metrics for the module to be simulated corresponding to the original simulation model through the workload monitor.
[0028] Specifically, the workload monitor is a dedicated module responsible for collecting real-time metrics of each simulation model.
[0029] The original simulation model refers to the simulation model currently in use and not yet switched for the module to be simulated in the current simulation scenario. In the current simulation scenario, the server can set the original simulation model based on the known workload type, or the original simulation model can be any of the multi-level simulation models. The module to be simulated refers to the hardware module of interest that needs to be simulated, including at least one of a computing module and a memory access module. In practical applications, the computing module includes a Graphics Processing Unit (GPU) computing core, internally containing Streaming Multiprocessors (SM), Tensor Cores, shared memory, tile buffers, double-buffered pipelines, etc. The memory access module includes a Network on Chip (NoC), multi-level cache (L1 / L2), and a Dynamic Random Access Memory (DRAM) controller.
[0030] Real-time metrics are quantitative data used to describe the real-time status and behavior of the module being simulated during operation, and they form the basis for subsequent workload characteristic analysis.
[0031] In the current simulation scenario, in order to identify the scenario type in subsequent steps, the real-time metrics corresponding to the current original simulation model are obtained through the workload monitor.
[0032] S102. The intelligent switching decision-maker determines the scenario type result corresponding to the current simulation scenario based on real-time indicators and through a pre-set operator feature knowledge base. The operator feature knowledge base includes the correspondence between theoretical indicators and scenario types.
[0033] Specifically, the operator feature knowledge base refers to a pre-set knowledge base that includes the correspondence between theoretical indicators and scenario types.
[0034] The server can compare and match the real-time metrics obtained from S101 with the theoretical metrics for various scene types pre-stored in the operator feature knowledge base to determine the scene type result corresponding to the current simulation scene. In practical applications, the real-time metrics can be compared with the theoretical metrics in the operator feature knowledge base to calculate the main metric fit, auxiliary multidimensional metric compatibility, and historical time series stability, thereby determining the scene type result and its confidence level. For example, when the real-time metrics include extremely high arithmetic strength and continuous regular memory access, it is determined that "the current stage may be the GEMM operator stage"; when the real-time metrics include rapid switching of computation modes and local mutation of memory access, it is determined that the stage is "attention computation".
[0035] S103, The intelligent switching decision-maker determines the target simulation model from the multi-level simulation models that matches the simulation accuracy of the scenario type results.
[0036] Specifically, after obtaining the current scene type result, the server can determine the target simulation model from the multi-level simulation models that matches the scene type result in terms of simulation accuracy, thereby determining the model that matches the current simulation scene.
[0037] S104. If the original simulation model is different from the target simulation model, the intelligent switching decision-maker determines the switching decision corresponding to the current simulation scenario based on the target simulation model so as to switch the original simulation model to the target simulation model by referring to the switching decision.
[0038] Specifically, after determining the target simulation model, the server needs to determine whether the target simulation model is consistent with the original simulation model. Only when the target simulation model is different from the original simulation model will the switching decision be triggered, thereby avoiding unnecessary frequent switching.
[0039] In practical applications, the main situations in which the target simulation model differs from the original simulation model include differences between the scenario type results corresponding to the current simulation scenario and the scenario type results corresponding to the previous simulation scenario, and a decrease in the confidence level of the target simulation model in the current simulation scenario.
[0040] After the target simulation model is determined in the aforementioned steps, an executable switching decision can be formed based on this information, so that the model dynamic switcher can switch the original simulation model to the target simulation model according to the switching decision.
[0041] By using the above method, real-time metrics for the module to be simulated are obtained, and scene type identification is performed by combining the operator feature knowledge base. The target simulation model that matches the scene type result in terms of simulation accuracy is determined. This enables simulation decisions that allow for flexible switching between simulation models with different simulation accuracies during the simulation process, so as to balance simulation speed and simulation accuracy.
[0042] In one possible implementation, the method further includes: The intelligent switching decision-maker determines the confidence level of the scenario type result based on the degree of matching between real-time indicators and theoretical indicators corresponding to the scenario type results. If the confidence level is less than the preset confidence level threshold, the intelligent switching decision-maker will determine the highest-precision simulation model among the multi-level simulation models as the target simulation model.
[0043] Specifically, in actual simulations, the matching between real-time metrics and theoretical metrics in the operator feature knowledge base is not always perfect. To ensure the reliability of switching decisions, the system needs to assess the confidence level of the current scene type judgment. When the confidence level is insufficient, a conservative strategy can be adopted, directly using the simulation model with the highest simulation accuracy for in-depth observation, in order to avoid loss of simulation accuracy due to incorrect judgments.
[0044] In one possible implementation, the intelligent switching decision-maker determines the confidence level of the scenario type result based on the degree of matching between real-time indicators and theoretical indicators corresponding to the scenario type results, including: The intelligent switching decision-maker determines the primary and secondary indicators corresponding to the scenario type results, as well as the theoretical ranges corresponding to the primary and secondary indicators respectively. The intelligent switching decision-maker determines the goodness of fit of the main indicator based on the real-time main indicator corresponding to the main indicator and the theoretical range of the main indicator. The intelligent switching decision-maker determines the compatibility of auxiliary indicators based on the real-time auxiliary indicators corresponding to the auxiliary indicators in the real-time indicators and the theoretical range of the auxiliary indicators. The intelligent switching decision-maker determines the historical time series stability based on the fluctuation of real-time indicators and the historical fluctuation of theoretical indicators corresponding to the scenario type results. The intelligent switching decision-maker determines the confidence level of the scenario type result based on the goodness of fit of the main indicator, the compatibility of the auxiliary indicator, and the stability of the historical time series.
[0045] Specifically, in this embodiment, the confidence level is not a simple judgment based on a single dimension, but rather a comprehensive evaluation of the credibility of the scene recognition result from multiple perspectives. That is, through multi-dimensional calculations, a comprehensive confidence level is finally synthesized to ensure the reliability of the judgment on the scene type.
[0046] First, determine the primary and secondary metrics corresponding to the scenario type results. Primary metrics refer to the core metrics for that scenario type, while secondary metrics refer to other metrics that can be used as a reference for that scenario type. In practical applications, as shown in Table 1, the primary and secondary metrics may differ for different scenario types. For example, when the scenario type result is determined to be "heavily computationally constrained," the primary metrics could be arithmetic strength and tensor core utilization, while secondary metrics could be memory access / computation ratio, vector unit activity, cache miss rate, DRAM bandwidth utilization, data dependency chain strength, memory access continuity, and shared memory conflict rate.
[0047] Table 1 Classification of Main and Auxiliary Indicators for Different Scenario Types
[0048] Furthermore, after determining the primary and secondary indicators corresponding to the scenario type results, the theoretical ranges of these indicators can be further defined for subsequent comparisons. In practical applications, Table 2 can be used as a reference to determine the theoretical ranges of the indicators for different scenario types.
[0049] Table 2. Theoretical range of theoretical indicators for different scenario types
[0050] Extract the real-time primary indicator corresponding to the primary indicator from the real-time indicators. Compare the value of the real-time primary indicator with the theoretical range of the primary indicator for the scene type and calculate the primary indicator fit. The fit can be calculated using distance metrics (such as Euclidean distance) or similarity metrics (such as cosine similarity). The higher the primary indicator fit value, the better the real-time primary indicator matches the theoretical range of the primary indicator.
[0051] Secondly, a further comparison can be made between real-time auxiliary indicators and their theoretical ranges. For example, if the scenario type result is "computationally constrained scenario," and the real-time auxiliary indicators simultaneously show both "cache miss rate exceeding 40%" and "DRAM bandwidth saturation," it indicates a potential deviation in the scenario type result. Based on the auxiliary indicators, the compatibility of the auxiliary indicators can be determined. Higher compatibility indicates fewer abnormal auxiliary indicators, while lower compatibility indicates more abnormal auxiliary indicators.
[0052] Based on the fluctuations of real-time indicators and the historical fluctuations of theoretical indicators corresponding to the scenario type results, the historical time series stability can be determined. The higher the historical time series stability, the more the real-time indicators match the historical performance of the scenario type results, and the higher the reliability of the judgment. The lower the historical time series stability, the less the real-time indicators match the historical performance of the scenario type results, and the lower the reliability of the judgment.
[0053] After obtaining the goodness of fit of the main indicator, the compatibility of the auxiliary indicators, and the historical time series stability, the confidence level can be determined by weighted summation. The weights of each indicator can be adjusted according to the specific application scenario.
[0054] In one possible implementation, the module to be simulated includes a computation module and a memory access module, and the multi-level simulation model includes a multi-level computation module simulation model for the computation module and a multi-level memory access module simulation model for the memory access module.
[0055] Specifically, in order to achieve refined simulation of AI chip systems, the modules to be simulated can include computing modules and memory access modules. Correspondingly, the multi-level simulation model can include a multi-level computing module simulation model for computing modules and a multi-level memory access module simulation model for memory access modules.
[0056] In one possible implementation, the original simulation module includes an original computational simulation model for the computation module and an original memory access simulation model for the memory access module. The intelligent switching decision-maker determines a target simulation model from the multi-level simulation models that matches the simulation accuracy of the scenario type results, including: The intelligent switching decision-maker determines the target computation simulation model and the target memory access simulation model from the multi-level computation module simulation model and the multi-level memory access module simulation model, which match the simulation accuracy of the scenario type results. If the scenario type results indicate that the current simulation scenario is a memory-constrained scenario, the accuracy of the target computation simulation model is lower than that of the target memory access simulation model. If the scenario type results indicate that the current simulation scenario is a computation-constrained scenario, the accuracy of the target computation simulation model is higher than that of the target memory access simulation model.
[0057] Specifically, since the module to be simulated consists of two different types of hardware—computation modules and memory access modules—and the performance bottleneck of the workload may fall on either type, model switching needs to be treated differently. In this embodiment, high-precision simulation resources are used for constrained modules, while low-precision simulation resources are used for unconstrained modules to improve efficiency.
[0058] When the scenario type result indicates that the current simulation scenario is a memory-constrained scenario, every detail of the memory access module may become the key to affecting the simulation accuracy. Therefore, the memory access module requires a simulation model with high simulation accuracy. On the other hand, since the computation module is often in a waiting state, its internal pipeline details have little impact on the overall simulation accuracy. The computation latency is usually hidden by the memory access latency, so a simulation model with low simulation accuracy can be used to advance quickly.
[0059] When the scenario type result indicates that the current simulation scenario is a computationally constrained scenario, the computation module requires a high-precision simulation model to accurately simulate it; while the details of the memory access module have a smaller impact on the overall process, and the memory access latency is usually hidden by the computation latency, so a low-precision simulation model can be used.
[0060] In one possible implementation, the multi-level simulation model includes a functional model with progressively increasing accuracy, an architecture-level model, a microarchitecture-level model, and a periodic precision model.
[0061] In one possible implementation, if the scenario type result indicates that the current simulation scenario is a heavily memory-constrained scenario, the target computation simulation model is a functional model, and the target memory access simulation model is a periodically accurate model; if the scenario type result indicates that the current simulation scenario is a lightly memory-constrained scenario, the target computation simulation model is an architecture-level model, and the target memory access simulation model is a microarchitecture-level model; if the scenario type result indicates that the current simulation scenario is a lightly computationally-constrained scenario, the target computation simulation model is a microarchitecture-level model, and the target memory access simulation model is an architecture-level model; if the scenario type result indicates that the current simulation scenario is a heavily computationally-constrained scenario, the target computation simulation model is a periodically accurate model, and the target memory access simulation model is a functional model.
[0062] Specifically, if the scenario type result indicates that the current simulation scenario is a heavily memory-restricted scenario, the target computation simulation model can be a functional model, omitting pipeline and execution unit details, and only performing computation according to ISA semantics and injecting the latency measured by the memory access model; the target memory access simulation model can be a periodic accurate model, accurately simulating the microarchitectural behavior of NoC routing arbitration, DRAM timing parameters, and cache coherence protocol with periodic precision.
[0063] If the scenario type result indicates that the current simulation scenario is a mildly memory-restricted scenario, the target computation simulation model can be an architecture-level model, tracking instruction-level pipeline statistics (such as issue width, execution unit utilization) but omitting microarchitectural hazard details; the target memory access simulation model is a microarchitectural-level model, simulating the periodic behavior of the tag array, consistency protocol state machine and cache line replacement strategy of the hierarchical cache (L1 / L2), and supporting fine-grained modeling of bank conflicts and replacement latency.
[0064] If the scenario type result indicates that the current simulation scenario is a lightly computationally constrained scenario, the target computation simulation model is a microarchitecture-level model, which accurately simulates the impact of pipeline stalls, instruction issue logic, execution unit contention, and branch prediction failures; the target memory access simulation model can be an architecture-level model, modeling average bandwidth and statistical latency.
[0065] If the scenario type result indicates that the current simulation scenario is a computationally intensive scenario, the target computation simulation model can be a periodic accurate model, which accurately simulates the systolic array microarchitecture of TensorCore, pipeline register states, and ultra-fine-grained hazards and delays inside the execution unit; the target memory access simulation model can be a functional model, which completes memory access operations with minimal simulation overhead and omits timing simulation.
[0066] In one possible implementation, if the original simulation model differs from the target simulation model, the method further includes: The intelligent handover decision-maker selects the handover boundary; In S104, the intelligent switching decision-maker, based on the target simulation model, determines the switching decision corresponding to the current simulation scenario so that the original simulation model can be switched to the target simulation model with reference to the switching decision. This includes: The intelligent switching decision-maker determines the switching decision corresponding to the current simulation scenario based on the switching boundary and the target simulation model, so as to switch the original simulation model to the target simulation model at the switching boundary with reference to the switching decision.
[0067] Specifically, when the target simulation model differs from the original simulation model, the server can select a switching boundary to determine when to switch. It then integrates and encapsulates the switching boundary and the target simulation model information so that the model dynamic switcher can switch the original simulation model to the target simulation model at the switching boundary based on the switching decision.
[0068] In one possible implementation, the intelligent handover decision-maker selects the handover boundary, including: The intelligent switching decision-maker selects the moment that satisfies the architectural visibility condition and the state closure condition as the switching boundary. The architectural visibility condition means that all visible side effects of the module to be simulated have been fully committed to the global visible view. The state closure condition means that there are no incomplete transactions that span time.
[0069] Specifically, the architecture visibility condition means that all visible side effects on the module to be simulated have been fully submitted to the global visibility view. The module to be simulated can include computation modules and memory access modules. The so-called "side effects" refer to the operation results that change the externally visible state, such as updating register values or writing memory data.
[0070] The state closure condition means that there are no incomplete transactions spanning time. Incomplete transactions include instructions being executed in the pipeline, memory requests that have not yet been written back, and bus transactions that have not converged.
[0071] Model switching cannot be performed at any time. In this embodiment, when both the architectural visibility condition and the state closure condition are met at a certain moment, it means that there is a complete state snapshot and no unfinished legacy transactions at that moment. The cost and risk of performing model switching are minimized, so this moment can be selected as the switching boundary.
[0072] In practical applications, the following types of switching boundaries can be included: Category 1: Trans-stage natural barriers.
[0073] 1) Prefill-Decode boundary (for autoregressive Transformer inference): The computational density and memory access patterns of the two stages are fundamentally different, and there is a natural control flow barrier between them.
[0074] 2) Gap between multiple kernels: In the CUDA programming model, the implicit synchronization point (kernel return) between successive kernel launches automatically ensures that all previous GPU activity is visible.
[0075] The second category: fine-grained synchronization points within sub-stages.
[0076] 1) Explicit ISA barrier instruction: syncthreads() forces all threads in CUDA to wait, at which point cross-thread data exchange in SharedMemory has been completed and the state is consistent.
[0077] 2) Fence / Flush instructions: Memory order guarantee instructions explicitly defined in the ISA layer (such as mfence in x86-64 and dmb in ARM) that ensure that all memory transactions preceding them are globally visible to subsequent instructions.
[0078] Category 3: Pipeline Drain.
[0079] The system reaches the ideal switching moment when the following conditions are met simultaneously: there are no instructions in the hardware pipeline, all FIFO buffers (such as TensorCore Feed, L2 Write-Back, DRAM command queue) are empty, and both compute and storage units are in a static state.
[0080] In one possible implementation, real-time metrics for the module to be simulated corresponding to the original simulation model are obtained through a workload monitor, including: The workload monitor obtains real-time metrics for the module to be simulated corresponding to the original simulation model through the performance monitoring interface. The performance monitoring interface includes a metric acquisition interface definition, which defines the acquired real-time metrics as a standardized metric vector of the structural standard.
[0081] Specifically, in this embodiment, in order to ensure that simulation models with different simulation accuracies can report comparable data to the same workload monitor, the system mandates a unified performance monitoring interface for all simulation models.
[0082] The performance monitoring interface refers to a standardized software interface for multi-level simulation models. It defines how multi-level simulation models with different simulation accuracies report performance data to the workload monitor, thereby standardizing the method, format, and timing semantics of each simulation model reporting performance indicators to the workload monitor.
[0083] The metric acquisition interface definition is used to define the acquired real-time metrics as standardized metric vectors with a structure standard. The typical operation is Get_Metrics(), which is used to return the real-time metrics in the current execution step. The standardized metric vector can include multiple metric entries, and each metric entry must contain at least: a unique metric identifier (MetricID), a metric value (Metric Value), and a metric validity flag or precision flag (Validity / Precision Flag).
[0084] In one possible implementation, the workload monitor obtains real-time metrics for the module to be simulated corresponding to the original simulation model through a performance monitoring interface, including: The workload monitor obtains real-time performance data for the module to be simulated corresponding to the original simulation model through the performance monitoring interface, and maps the real-time performance data to real-time indicators corresponding to a pre-set set of indicator identifiers; the set of indicator identifiers corresponds to the multi-level simulation model.
[0085] Specifically, real-time performance data generated by simulation models of different accuracies (functional models, architecture-level models, microarchitecture-level models, and periodically accurate models) have different formats, meanings, and granularities. To enable workload monitors to understand this data uniformly, various real-time performance data can be mapped to a unified set of indicator identifiers.
[0086] The set of indicator identifiers can adopt a predefined global indicator numbering system, including ID_001: Instructions Per Cycle (IPC), ID_002: Memory Access Bandwidth Utilization, ID_003: Arithmetic Strength, etc. In practical applications, indicators not supported by the current simulation model can be marked as invalid without affecting the overall parsing and subsequent processing of the indicator vector.
[0087] In one possible implementation, real-time metrics for the module to be simulated corresponding to the original simulation model are obtained through a workload monitor, including: The workload monitor obtains real-time metrics for the module to be simulated from the original simulation model at the execution boundary granularity, with the execution boundary corresponding to the switching boundary.
[0088] Specifically, the workload monitor collects metrics at the execution boundary granularity. The execution boundary and the switching boundary can be kept consistent, thus ensuring that the real-time metrics returned by each simulation model within the same execution boundary represent the statistical results within the same execution step, thereby ensuring the comparability of cross-model metrics in the time dimension.
[0089] In one possible implementation, the operator feature knowledge base includes at least one of dense matrix multiplication, attention computation, element-level operations, normalization layers, and data rearrangement.
[0090] Specifically, the pre-set operator feature knowledge base is built based on architectural theory analysis and a large amount of offline analysis data, which enables accurate identification of scene types.
[0091] In one possible implementation, the theoretical metrics in the operator feature knowledge base include at least two of the following: arithmetic strength, tensor core utilization, memory access / computation ratio, vector unit activity, cache miss rate, DRAM bandwidth utilization, data dependency chain strength, memory access continuity, and shared memory conflict rate.
[0092] Specifically, the pre-set operator feature knowledge base can include multiple theoretical indicators. In practical applications, the theoretical indicators corresponding to different scenario types are shown in Table 2 above.
[0093] Therefore, this application provides a method for determining switching decisions for multi-level simulation models, where the simulation accuracy of the multi-level simulation models increases progressively. The method includes: acquiring real-time metrics for the module to be simulated corresponding to the original simulation model through a workload monitor in the current simulation scenario; determining the scenario type result corresponding to the current simulation scenario based on the real-time metrics using a pre-set operator feature knowledge base; the operator feature knowledge base includes the correspondence between theoretical metrics and scenario types; determining a target simulation model from the multi-level simulation models that matches the scenario type result in terms of simulation accuracy; if the original simulation model differs from the target simulation model, determining a switching decision for the current simulation scenario based on the target simulation model so that the original simulation model can be switched to the target simulation model with reference to the switching decision. Through this method, real-time metrics for the module to be simulated are obtained, and scenario type identification is performed using the operator feature knowledge base to determine the target simulation model that matches the scenario type result in terms of simulation accuracy. This enables the determination of simulation decisions for flexibly switching simulation models with different simulation accuracies during the simulation process, thus balancing simulation speed and accuracy.
[0094] In the description of this specification, references to terms such as "some possible implementations," "some implementations," "example," "specific example," or "some examples" indicate that a specific feature, structure, material, or characteristic described in connection with that implementation or example is included in at least one implementation or example of this application, and the aforementioned terms do not necessarily refer to the same implementation or example. Furthermore, the described specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more implementations or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different implementations or examples described in this specification, as well as the features of different implementations or examples.
[0095] The method flowcharts for embodiments of this application describe certain operations as different steps performed in a certain order. Such flowcharts are illustrative and not restrictive. Some steps described herein may be grouped together and performed in a single operation, or some steps may be divided into multiple sub-steps, and some steps may be performed in an order different from that shown herein. The various steps shown in the flowcharts may be implemented in any way by any circuit structure and / or tangible mechanism (e.g., by software running on a computer device, hardware (e.g., logic functions implemented by a processor or chip), and / or any combination thereof).
[0096] Those skilled in the art will understand that in the methods described in the above specific embodiments, the order in which the steps are written does not imply a strict execution order, and the specific execution order of each step should be determined by its function and possible internal logic.
[0097] Based on the foregoing Figure 1 The following describes the switching decision determination device for multi-level simulation models provided in this application through a device embodiment. Figure 2 This application provides a schematic diagram of a switching decision determination device for a multi-level simulation model, as shown in the embodiments of the present application. Figure 2 As shown, the switching decision determination device 200 for multi-level simulation models includes: The workload monitor 201 is used to obtain real-time metrics for the module to be simulated corresponding to the original simulation model in the current simulation scenario. The intelligent switching decision-maker 203 is used to determine the scene type result corresponding to the current simulation scene based on real-time indicators and through a pre-set operator feature knowledge base 202; the operator feature knowledge base 202 includes the correspondence between theoretical indicators and scene types; The intelligent switching decision-maker 203 is also used to determine the target simulation model from the multi-level simulation models that matches the simulation accuracy of the scene type results; The intelligent switching decision 203 is also used to determine the switching decision corresponding to the current simulation scenario based on the target simulation model if the original simulation model is different from the target simulation model, so as to switch the original simulation model to the target simulation model at the switching boundary by referring to the switching decision.
[0098] In one possible implementation, the intelligent switching decision-maker 203 is also used for: The confidence level of the scenario type results is determined based on the degree of matching between real-time indicators and theoretical indicators corresponding to scenario type results. If the confidence level is less than the preset confidence level threshold, the highest-precision simulation model among the multi-level simulation models will be determined as the target simulation model.
[0099] In one possible implementation, the intelligent switching decision-maker 203 is used for: Determine the primary and secondary indicators corresponding to the results of each scenario type, as well as the theoretical ranges for each primary and secondary indicator; The fit of the main indicator is determined based on the real-time main indicator corresponding to the main indicator and the theoretical range of the main indicator. The compatibility of auxiliary indicators is determined based on the real-time auxiliary indicators corresponding to the auxiliary indicators in the real-time indicators and the theoretical range corresponding to the auxiliary indicators. Historical time series stability is determined based on the fluctuation of real-time indicators and the historical fluctuation of theoretical indicators corresponding to the scenario type results. The confidence level of the scenario type results is determined based on the goodness of fit of the main index, the compatibility of the auxiliary index, and the stability of the historical time series.
[0100] In one possible implementation, the module to be simulated includes a computation module and a memory access module, and the multi-level simulation model includes a multi-level computation module simulation model for the computation module and a multi-level memory access module simulation model for the memory access module.
[0101] In one possible implementation, the intelligent switching decision-maker 203 is used for: The original simulation module includes an original computation simulation model for the computation module and an original memory access simulation model for the memory access module. From the multi-level computation module simulation model and the multi-level memory access module simulation model, a target computation simulation model and a target memory access simulation model are determined that match the simulation accuracy of the scenario type results. If the scenario type results indicate that the current simulation scenario is a memory-constrained scenario, the accuracy of the target computation simulation model is lower than that of the target memory access simulation model; if the scenario type results indicate that the current simulation scenario is a computation-constrained scenario, the accuracy of the target computation simulation model is higher than that of the target memory access simulation model.
[0102] In one possible implementation, the multi-level simulation model includes a functional model with progressively increasing accuracy, an architecture-level model, a microarchitecture-level model, and a periodic precision model.
[0103] In one possible implementation, if the scenario type result indicates that the current simulation scenario is a heavily memory-constrained scenario, the target computation simulation model is a functional model, and the target memory access simulation model is a periodically accurate model; if the scenario type result indicates that the current simulation scenario is a lightly memory-constrained scenario, the target computation simulation model is an architecture-level model, and the target memory access simulation model is a microarchitecture-level model; if the scenario type result indicates that the current simulation scenario is a lightly computationally-constrained scenario, the target computation simulation model is a microarchitecture-level model, and the target memory access simulation model is an architecture-level model; if the scenario type result indicates that the current simulation scenario is a heavily computationally-constrained scenario, the target computation simulation model is a periodically accurate model, and the target memory access simulation model is a functional model.
[0104] In one possible implementation, the intelligent switching decision-maker 203 is also used for: If the original simulation model differs from the target simulation model, select the switching boundary; Intelligent switching decision maker 203, used for Based on the switching boundary and the target simulation model, the switching decision corresponding to the current simulation scenario is determined so that the original simulation model can be switched to the target simulation model at the switching boundary with reference to the switching decision.
[0105] In one possible implementation, the intelligent switching decision-maker 203 is used for: The switching boundary is selected when the architectural visibility condition and the state closure condition are met. The architectural visibility condition means that all visible side effects of the module to be simulated have been fully committed to the global visible view. The state closure condition means that there are no incomplete transactions that span time.
[0106] In one possible implementation, the workload monitor 201 is used for: The performance monitoring interface obtains real-time metrics for the module to be simulated corresponding to the original simulation model. The performance monitoring interface includes a metric acquisition interface definition, which defines the real-time metrics to be acquired as a standardized metric vector of the structural standard.
[0107] In one possible implementation, the workload monitor 201 is used for: The real-time performance data of the module to be simulated corresponding to the original simulation model is obtained through the performance monitoring interface, and the real-time performance data is mapped to real-time indicators corresponding to a pre-set set of indicator identifiers; the set of indicator identifiers corresponds to the multi-level simulation model.
[0108] In one possible implementation, the workload monitor 201 is used for: Real-time metrics for the module to be simulated corresponding to the original simulation model are obtained at the execution boundary granularity, with the execution boundary corresponding to the switching boundary.
[0109] In one possible implementation, the operator feature knowledge base 202 includes at least one of dense matrix multiplication, attention computation, element-level operations, normalization layer, and data rearrangement.
[0110] In one possible implementation, the theoretical metrics in the operator feature knowledge base 202 include at least two of the following: arithmetic strength, tensor core utilization, memory access / computation ratio, vector unit activity, cache miss rate, DRAM bandwidth utilization, data dependency chain strength, memory access continuity, and shared memory conflict rate.
[0111] It should be noted that the apparatus in the embodiments of this application can implement each process of the aforementioned method and achieve the same effect and function, which will not be elaborated here.
[0112] The following is through Figure 3 This application provides a detailed description of the switching method and switching decision determination method for multi-level simulation models: In the current simulation scenario, the workload monitor is used to obtain real-time metrics for the modules to be simulated, which include memory access modules and computation modules.
[0113] The intelligent switching decision-maker determines the scenario type result corresponding to the current simulation scenario based on real-time indicators and through a pre-set operator feature knowledge base. The operator feature knowledge base includes the correspondence between theoretical indicators and scenario types, which is used to provide decision support for the intelligent switching decision-maker.
[0114] The intelligent switching decision-maker determines the target simulation model from the multi-level simulation models that matches the simulation accuracy of the scenario type results.
[0115] If the original simulation model is different from the target simulation model, the intelligent switching decision-maker determines the switching decision corresponding to the current simulation scenario based on the target simulation model so as to switch the original simulation model to the target simulation model by referring to the switching decision.
[0116] The intelligent switching decision-maker sends a switching decision corresponding to the current simulation scenario to the model dynamic switcher. The switching decision includes a target simulation model that matches the simulation accuracy of the module to be simulated in the current simulation scenario, determined from the multi-level simulation models.
[0117] The model dynamic switcher acquires the first intermediate simulation data of the original simulation model for the current simulation scenario and saves the state.
[0118] The model dynamic switcher converts the first intermediate simulation data into second intermediate simulation data that matches the target simulation model, and sends the second intermediate simulation data to the target simulation model to perform state transition and state recovery.
[0119] The model dynamic switcher refers to the switching decision and switches the original simulation model to the target simulation model.
[0120] After the switch, the workload monitor obtains the performance feedback of the target simulation model for the module to be simulated in the current simulation scenario. The intelligent switching decision-maker determines the confidence level of the target simulation model based on the degree of matching between the performance feedback and the expected simulation data corresponding to the current simulation scenario. If the confidence level is lower than the preset back-switch confidence threshold, the model dynamic switcher switches the target simulation model back to the original simulation model; if the confidence level is higher than the preset switching confidence threshold, the model dynamic switcher switches the target simulation model to a simulation model with lower simulation accuracy compared to the target simulation model in the multi-level simulation model.
[0121] The following example illustrates this through a co-simulation implementation of dynamic switching between multi-precision models in the inference process of an autoregressive transformer (Transformer) model: This embodiment uses an AI chip designed for autoregressive Transformer inference tasks as the simulation object. The chip includes a computing module and a memory access module. The computing module may include one or more computing cores, and internally includes a matrix computing array, shared memory, on-chip buffers, and pipelined control structures; the memory access module includes an on-chip network, multi-level caches, and an external memory controller.
[0122] For the above modules, this embodiment constructs the following multi-level simulation model sets respectively:
[0123] All models implement a unified co-simulation interface specification, supporting state export, state import, operation control, and switching coordination.
[0124] 1. Setting up the original simulation model.
[0125] This embodiment selects the autoregressive Transformer inference task as a typical workload, and its execution flow includes: 1) Prefill phase: Input Prompt Token and perform large-scale matrix multiplication and attention calculation; 2) Decode stage: Token-by-to-token reasoning, frequent access to KV Cache, and significant memory access pressure.
[0126] In the initial stage of simulation, since sufficient real-time performance feedback has not yet been accumulated, the original simulation model combination can be set based on known workload types, operator prior characteristics, or default startup strategies. During subsequent operation, the intelligent switching decision-maker dynamically adjusts the model combination based on real-time metrics. Before entering the formal dynamic switching process, the autoregressive Transformer inference task uses the following original simulation model combination: Computing module: Architecture-level model; Storage module: Architecture-level model.
[0127] This original simulation model combination is used to rapidly advance the overall simulation progress of the Prefill phase, while ensuring that the performance characteristics at each stage are observable.
[0128] 2. Operation and feature determination in the Prefill stage.
[0129] During the Prefill phase, the workload monitor collects or derives real-time metrics for the current phase: GPU compute unit utilization, arithmetic intensity, storage tier traffic, average memory access latency, cache hit rate, or locality estimation metrics.
[0130] The intelligent switching decision-maker generates a confidence score based on the matching results between real-time indicators and the operator feature knowledge base. When the confidence score is higher than the preset high confidence score threshold, it is determined to be the corresponding high confidence scenario. When the confidence score is lower than the preset low confidence score threshold or the prediction error of multiple consecutive sampling windows exceeds the error threshold, it is determined to be a low confidence scenario or an unknown mode.
[0131] When real-time metrics indicate that the utilization rate of computing units remains at a high level, the arithmetic intensity is higher than the preset threshold, and the average memory access latency is low or hidden by the computation process, the current simulation scenario is determined to be a heavily computationally constrained scenario.
[0132] Based on the scenario type result being a heavily computationally constrained scenario, the calculation module uses a precise cycle model, and the memory access module uses a functional model.
[0133] 3. Model switching execution during the Prefill phase.
[0134] 1) Switch boundary selection.
[0135] The following stable execution boundaries are selected as switching boundaries: after a certain GEMM Kernel is completed in the Prefill phase; or after the GPU pipeline drain is completed and all in-transit instructions and memory access transactions have converged, at which point the architectural visibility condition and the state closure condition are satisfied.
[0136] 2) Switching actions.
[0137] The model dynamic switcher performs the following operations: switches the computing module from the architecture-level model to the periodic precise model; switches the memory access module from the architecture-level model to the functional model.
[0138] 3) Status handling method.
[0139] For internal states that affect the visible behavior of the subsequent architecture or significantly affect the initialization accuracy of the target simulation model, priority should be given to migration or reconstruction; for transient states that only serve the local execution process of the original simulation model source model and can be safely drained or discarded at the switching boundary, no cross-model migration should be performed.
[0140] On the computation module side: the register file, program counter, committed write-back data, and architecture-visible data in shared memory are retained; for Tile Buffer, matrix computation array FIFO, pipeline register, etc., which only serve the transient intermediate state of the micro-execution process, no cross-model migration is performed when the emptying condition is met.
[0141] On the memory access subsystem side: the memory content visible to the architecture is preserved; if the memory access transaction in transit has already converged at the switching boundary, the micro-queuing state and transient memory access transaction are not migrated.
[0142] After the switch is complete, the critical computational kernel of the Prefill phase continues to execute under the cycle-accurate model, which is used to accurately analyze pipeline conflicts, TensorCore utilization and other issues.
[0143] 4. The Decode phase runs and switches back.
[0144] After the Prefill phase ends, the system enters the Decode phase.
[0145] The workload monitor detected the following: decreased GPU compute unit utilization, a significant increase in KV cache access frequency, fluctuating cache hit rates or deterioration in localized metrics, with memory access latency becoming the dominant factor. These characteristics indicate that compute resources are idle while storage access has become the dominant bottleneck, thus classifying the current stage as a severely memory-constrained scenario.
[0146] Based on the scenario type result, which indicates a heavily memory-constrained scenario, the memory access module requires a higher-precision model, such as a periodically accurate model; the computation module can be switched to a functional model, or switched to an architecture-level model as needed.
[0147] 5. Model switching execution during the Decode phase.
[0148] 1) Switch boundary selection.
[0149] The following stable execution boundaries are selected as switching boundaries: after a single Token Decode Kernel is completed; at the synchronization point (barrier) of a thread or task group, the thread state is consistent.
[0150] 2) Switching actions.
[0151] The model dynamic switcher performs the following switches: switching the computation module from the periodic exact model to the functional model; and switching the memory access module from the functional model to the periodic exact model.
[0152] 3) Status handling method.
[0153] For internal states that affect the visible behavior of the subsequent architecture or significantly affect the initialization accuracy of the target simulation model, priority should be given to migration or reconstruction; for transient states that only serve the local execution process of the original simulation model source model and can be safely drained or discarded at the switching boundary, no cross-model migration should be performed.
[0154] On the computation module side: only the necessary architectural state is loaded, such as the generated token sequence, program control state, and register context related to subsequent execution.
[0155] On the memory access subsystem side: prioritize loading directly inheritable states, including architecture-visible states such as main memory content, as well as performance-related internal states such as cache-valid states and necessary directory information that are available in the source model.
[0156] For microstates not explicitly maintained by the functional model, including the memory access queue in transit, NoC arbitration state, and DRAM scheduling history state, recovery is achieved through default initialization, warm-up operation, or statistical reconstruction.
[0157] The Decode phase then runs under a precise memory access cycle model to analyze cache line conflicts, NoC congestion, and DRAM scheduling bottlenecks.
[0158] 6. Anomaly detection and rollback mechanism.
[0159] When the confidence level of the match between the performance feedback and the expected simulation data corresponding to the current simulation scenario is lower than the preset back-cut confidence threshold, or when the performance prediction error exceeds the threshold in multiple consecutive sampling windows, the system marks this stage as a low-confidence or unknown mode and automatically triggers the back-cut mechanism. The back-cut mechanism can switch at least one of the computing module and the memory access module to the microarchitecture-level model. In this embodiment, both can be switched to the microarchitecture-level model simultaneously to further locate the source of memory access conflicts and scheduling anomalies.
[0160] After the anomaly analysis is completed, the system switches to a suitable combination of high-abstract models with low simulation accuracy based on the new characteristics.
[0161] This application also provides an electronic device, including a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, the following processing is performed: In the current simulation scenario, the workload monitor is used to obtain real-time metrics for the module to be simulated corresponding to the original simulation model. Based on real-time metrics, the scenario type result corresponding to the current simulation scenario is determined through a pre-set operator feature knowledge base; the operator feature knowledge base includes the correspondence between theoretical metrics and scenario types. From the multi-level simulation models, identify the target simulation model that matches the simulation accuracy of the scene type results; If the original simulation model is different from the target simulation model, a switching decision is determined based on the target simulation model so that the original simulation model can be switched to the target simulation model with reference to the switching decision.
[0162] This application also provides a computer-readable storage medium storing a computer program. When a processor runs this computer program, it executes the steps of the switching decision determination method for a multi-level simulation model described in the above-described method embodiments. The storage medium can be a volatile or non-volatile computer-readable storage medium.
[0163] This application also provides a computer program product, including a computer program carrying program code. The instructions included in the program code can be used to execute the steps of the switching decision determination method for a multi-level simulation model described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.
[0164] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0165] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on its differences from other embodiments. In particular, the descriptions of the apparatus, device, and computer-readable storage medium embodiments are simplified because they are substantially similar to the method embodiments; relevant details can be found in the descriptions of the method embodiments.
[0166] The apparatus, device, and computer-readable storage medium provided in the embodiments of this application correspond one-to-one with the method. Therefore, the apparatus, device, and computer-readable storage medium also have similar beneficial technical effects as their corresponding methods. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the apparatus, device, and computer-readable storage medium will not be repeated here.
[0167] While the spirit and principles of this application have been described above with reference to several specific embodiments, it should be understood that this application is not limited to the disclosed specific embodiments, and the division of aspects does not imply that features in these aspects cannot be combined. This application is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.
Claims
1. A method for determining switching decisions in a multi-level simulation model, characterized in that, The simulation accuracy of the multi-level simulation model increases progressively, and the method includes: In the current simulation scenario, the workload monitor is used to obtain real-time metrics for the module to be simulated corresponding to the original simulation model. The intelligent switching decision-maker determines the scenario type result corresponding to the current simulation scenario based on the real-time indicators and through a pre-set operator feature knowledge base; the operator feature knowledge base includes the correspondence between theoretical indicators and scenario types. The intelligent switching decision-maker determines a target simulation model from the multi-level simulation models that matches the simulation accuracy of the scenario type result; If the original simulation model is different from the target simulation model, the intelligent switching decision-maker determines the switching decision corresponding to the current simulation scenario based on the target simulation model so as to switch the original simulation model to the target simulation model by referring to the switching decision.
2. The method according to claim 1, characterized in that, The method further includes: The intelligent switching decision-maker determines the confidence level of the scenario type result based on the degree of matching between the real-time indicators and the theoretical indicators corresponding to the scenario type results. If the confidence level is less than a preset confidence threshold, the intelligent switching decision-maker will determine the highest-precision simulation model among the multi-level simulation models as the target simulation model.
3. The method according to claim 2, characterized in that, The intelligent switching decision-maker determines the confidence level of the scenario type result based on the degree of matching between the real-time indicators and the theoretical indicators corresponding to the scenario type result, including: The intelligent switching decision-maker determines the primary and secondary indicators corresponding to the scenario type results, as well as the theoretical ranges corresponding to the primary and secondary indicators respectively. The intelligent switching decision-maker determines the fit of the main indicator based on the real-time main indicator corresponding to the main indicator and the theoretical range corresponding to the main indicator. The intelligent switching decision-maker determines the compatibility of the auxiliary indicators based on the real-time auxiliary indicators corresponding to the auxiliary indicators in the real-time indicators and the theoretical range corresponding to the auxiliary indicators. The intelligent switching decision-maker determines the historical time series stability based on the fluctuation of the real-time indicators and the historical fluctuation of the theoretical indicators corresponding to the scenario type results. The intelligent switching decision-maker determines the confidence level of the scenario type result based on the main indicator fit degree, the auxiliary indicator compatibility degree, and the historical time series stability.
4. The method according to claim 1, characterized in that, The module to be simulated includes a computing module and a memory access module. The multi-level simulation model includes a multi-level computing module simulation model for the computing module and a multi-level memory access module simulation model for the memory access module.
5. The method according to claim 4, characterized in that, The original simulation module includes an original computational simulation model for the computation module and an original memory access simulation model for the memory access module. The intelligent switching decision-maker determines a target simulation model from the multi-level simulation models that matches the simulation accuracy of the scenario type result, including: The intelligent switching decision-maker determines a target computing simulation model and a target memory access simulation model that match the simulation accuracy of the scenario type result from the multi-level computing module simulation model and the multi-level memory access module simulation model. If the scenario type result identifies the current simulation scenario as a memory-constrained scenario, the accuracy of the target computing simulation model is lower than that of the target memory access simulation model. If the scenario type result identifies the current simulation scenario as a computing-constrained scenario, the accuracy of the target computing simulation model is higher than that of the target memory access simulation model.
6. The method according to claim 5, characterized in that, The multi-level simulation model includes functional models with progressively increasing accuracy, architecture-level models, microarchitecture-level models, and periodic precision models.
7. The method according to claim 6, characterized in that, If the scenario type result identifies the current simulation scenario as a heavily memory-constrained scenario, the target computation simulation model is a functional model, and the target memory access simulation model is a periodic accurate model; if the scenario type result identifies the current simulation scenario as a lightly memory-constrained scenario, the target computation simulation model is an architecture-level model, and the target memory access simulation model is a microarchitecture-level model. If the scenario type result identifies the current simulation scenario as a lightly computationally constrained scenario, the target computation simulation model is a microarchitecture-level model, and the target memory access simulation model is an architecture-level model; If the scenario type result identifies the current simulation scenario as a heavily computationally constrained scenario, the target computation simulation model as a periodic accurate model, and the target memory access simulation model as a functional model.
8. The method according to claim 1, characterized in that, If the original simulation model is different from the target simulation model, the method further includes: The intelligent switching decision-maker selects the switching boundary; The intelligent switching decision-maker determines the switching decision corresponding to the current simulation scenario based on the target simulation model, so as to switch the original simulation model to the target simulation model by referring to the switching decision, including: The intelligent switching decision-maker determines the switching decision corresponding to the current simulation scenario based on the switching boundary and the target simulation model, so as to switch the original simulation model to the target simulation model at the switching boundary with reference to the switching decision.
9. The method according to claim 8, characterized in that, The intelligent handover decision-maker selects the handover boundary, including: The intelligent switching decision-maker selects the moment that satisfies the architectural visibility condition and the state closure condition as the switching boundary. The architectural visibility condition means that all visible side effects of the module to be simulated have been fully committed to the global visible view. The state closure condition means that there are no incomplete transactions that span time.
10. The method according to claim 1, characterized in that, The step of obtaining real-time metrics for the module to be simulated corresponding to the original simulation model through the workload monitor includes: The workload monitor obtains real-time metrics for the module to be simulated corresponding to the original simulation model through the performance monitoring interface; the performance monitoring interface includes a metric acquisition interface definition, which defines the acquired real-time metrics as a standardized metric vector of structural standards.
11. The method according to claim 10, characterized in that, The workload monitor obtains real-time metrics for the module to be simulated corresponding to the original simulation model through a performance monitoring interface, including: The workload monitor obtains real-time performance data for the module to be simulated corresponding to the original simulation model through the performance monitoring interface, and maps the real-time performance data to real-time indicators corresponding to a pre-set set of indicator identifiers; the set of indicator identifiers corresponds to the multi-level simulation model.
12. The method according to claim 1, characterized in that, The step of obtaining real-time metrics for the module to be simulated corresponding to the original simulation model through the workload monitor includes: The workload monitor obtains real-time metrics for the module to be simulated corresponding to the original simulation model at the execution boundary granularity, and the execution boundary corresponds to the switching boundary.
13. The method according to claim 1, characterized in that, The operator feature knowledge base includes at least one of dense matrix multiplication, attention computation, element-level operations, normalization layers, and data rearrangement.
14. The method according to claim 13, characterized in that, The theoretical metrics in the operator feature knowledge base include at least two of the following: arithmetic strength, tensor core utilization, memory access / computation ratio, vector unit activity, cache miss rate, DRAM bandwidth utilization, data dependency chain strength, memory access continuity, and shared memory conflict rate.
15. A switching decision-making device for a multi-level simulation model, characterized in that, The simulation accuracy of the multi-level simulation model increases progressively, and the device includes: The workload monitor is used to obtain real-time metrics for the module to be simulated corresponding to the original simulation model in the current simulation scenario. An intelligent switching decision-maker is used to determine the scene type result corresponding to the current simulation scene based on the real-time indicators and through a pre-set operator feature knowledge base; the operator feature knowledge base includes the correspondence between theoretical indicators and scene types; The intelligent switching decision-maker is also used to determine a target simulation model from the multi-level simulation model that matches the simulation accuracy of the scenario type result; The intelligent switching decision-maker is further configured to, if the original simulation model is different from the target simulation model, determine the switching decision corresponding to the current simulation scenario based on the target simulation model so as to switch the original simulation model to the target simulation model by referring to the switching decision.
16. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, a switching decision determination method for a multi-level simulation model as described in any one of claims 1 to 14 is performed.
17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs a switching decision determination method for a multi-level simulation model as described in any one of claims 1 to 14.
18. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, performs a switching decision determination method for a multi-level simulation model as described in any one of claims 1 to 14.