Method and system for evaluating computer hardware performance based on analogue simulation model

By building a hardware component model and a task load model, combined with event-driven scheduling simulation task execution, the problem of insufficient accuracy and single dimension of hardware performance evaluation in the existing technology is solved, and the accuracy and visualization of multi-index performance evaluation is achieved.

CN120353684AActive Publication Date: 2025-07-22BEIJING ZUNGUAN TECH

Patent Information

Application Number
CN202510845899.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

The prior art lacks a high-fidelity performance mapping mechanism for general computing tasks in computer hardware performance evaluation, the evaluation dimension is relatively single, it is difficult to reflect the hardware response behavior under real workloads, and there is a lack of a dynamic input task load-driven mechanism.

Method used

Build a computer hardware performance evaluation method based on simulation model, establish a hardware component model by analyzing the system structure description file, extract the task load model, and perform task scheduling under the event-driven mechanism, and record multi-index performance indicators, including task execution delay, data communication overhead and power consumption.

Benefits of technology

It realizes multi-index simulation evaluation of computer hardware performance, improves evaluation accuracy and interpretability, can quantify performance differences under real load conditions, and supports multi-dimensional performance evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353684A_ABST
    Figure CN120353684A_ABST
Patent Text Reader

Abstract

The invention provides a computer hardware performance evaluation method and system based on an analogue simulation model, and relates to the technical field of computer performance evaluation. According to the method, hardware component models such as a processor, a memory and I / O are constructed by analyzing a system structure description file, semantic analysis is performed on an instruction stream or a parallel computing graph of a to-be-evaluated application, and a task load model is formed. The model is input into an event-driven scheduling module, simulation is executed on the hardware component model, task-resource mapping is generated, and task time delay, communication traffic and power consumption are recorded. And calculating average response time delay, data bandwidth and unit energy consumption according to the simulation data, and outputting an evaluation report compared with the performance baseline. According to the method, real load-oriented multi-dimensional performance prediction can be realized, and the method can be used for architecture optimization and scheme selection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer performance evaluation, and particularly to a method and system for evaluating computer hardware performance based on a simulation model. Background Art

[0002] With the rapid evolution of computer architectures, the complexity of processors, memory, GPUs, and I / O subsystems has been continuously increasing. Traditional hardware performance evaluation methods mostly rely on actual measurements, that is, performance metrics are obtained by running standard performance testing software (such as SPEC, PassMark, etc.). Although such methods are intuitive, their adaptability and repeatability are limited, and it is difficult to conduct rapid comparative analysis in the early design stage or for multiple hardware configuration combinations.

[0003] In recent years, with the improvement of virtualization technology, system simulation, and modeling capabilities, researchers are increasingly inclined to use software-level simulation means to predict hardware performance. Especially in the design stage of system-on-chip (SoC) and high-performance server architectures, it has become a trend to conduct simulation analysis by constructing hardware behavior models. In addition, AI-assisted performance modeling methods have also begun to emerge to improve the evaluation accuracy and inference speed.

[0004] The application of current simulation models in hardware performance evaluation still faces several challenges: one is the lack of a high-fidelity performance mapping mechanism for general computing tasks, resulting in a large deviation between the model prediction value and the actual running value; the second is that the evaluation dimension is relatively single, often ignoring the collaborative constraints of multiple metrics (such as latency, bandwidth, power consumption, etc.); the third is the lack of a dynamic input task load driving mechanism, making it difficult to reflect the response behavior of the hardware under real workloads. Therefore, there is an urgent need for a simulation-based performance evaluation method that can comprehensively consider load characteristics, hardware model accuracy, and multi-metric output. Summary of the Invention

[0005] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide a method and system for evaluating computer hardware performance based on a simulation model, which realizes multi-metric simulation evaluation of hardware performance for real task loads and improves the accuracy and interpretability of performance prediction.

[0006] To achieve the above object, the present invention provides the following solutions: A method for evaluating computer hardware performance based on a simulation model, comprising: Based on the system structure description file of the target computer system, constructing hardware component models for each computing hardware unit; Based on the instruction stream or parallel computing graph of the application to be evaluated, extracting the task operation sequence, resource call behavior, and data dependency relationship to form a task load model; Input the task load model into the scheduling module, and use the scheduling module to start the task scheduling process under the event-driven mechanism to control the simulation execution of each task in the corresponding hardware component model, dynamically generate the resource mapping relationship between the task and the hardware component model, and record the task execution delay, data communication overhead, and power consumption index at the same time; Determine the resource mapping relationship, task execution delay, data communication overhead, and power consumption index as simulation data; Construct a multi-index performance evaluation result including average response delay, total data transmission bandwidth, and unit task energy consumption based on the simulation data, and horizontally compare the multi-index performance evaluation result with the preset performance baseline to obtain a performance evaluation report.

[0007] Preferably, it further includes: Use the performance evaluation report as a feedback input to iteratively optimize the key parameters of the hardware component model; the key parameters include: processing delay parameters, communication bandwidth parameters, power consumption model parameters, and resource capacity parameters.

[0008] Preferably, based on the system structure description file of the target computer system, construct the hardware component models of each computing hardware unit, including: Read the system structure description file of the target computer system; the system structure description file includes processor structure, memory hierarchy, interconnection topology, and I / O interface configuration parameters; Parse the system structure description file and extract the structure parameters and operating characteristic parameters of each computing hardware unit; the structure parameters and operating characteristic parameters include main frequency, cache capacity, bus width, access delay, and power consumption model coefficient; Establish corresponding hardware component models according to the extracted parameters; the hardware component models include processor models, memory models, I / O models, and interconnection structure models; Store the constructed hardware component models in the model library of the simulation platform in a modular form.

[0009] Preferably, based on the instruction stream or parallel computation graph of the application to be evaluated, extract the task operation sequence, resource call behavior, and data dependency relationship to form a task load model, including: Obtain the instruction stream or parallel computation graph of the application to be evaluated; the instruction stream is the instruction sequence generated after the application program is compiled or during operation; the parallel computation graph represents the control and data dependency relationship between multiple tasks; Perform semantic parsing on the instruction stream or parallel computation graph to extract the task operation types; the task operation types include arithmetic operations, logical judgments, memory access operations, and communication instructions; Analyze the calling relationship between the instruction stream or parallel computing graph and the hardware resources to determine the calling paths and occupancy cycles of each task for the processor, memory, cache, and I / O resources, and generate a resource calling behavior description; Construct a data dependency graph between tasks based on a preset read-write order to represent the pre-and post-execution constraints of tasks; Encapsulate the task operation type, the resource calling behavior description, and the pre-and post-execution constraints into structured data to form a task workload model.

[0010] Preferably, analyze the calling relationship between the instructions or nodes and the hardware resources to determine the calling paths and occupancy cycles of each task for the processor, memory, cache, and I / O resources, and generate a resource calling behavior description, including: Parse the task nodes in the instruction stream or parallel computing graph and identify the resource usage fields involved in each instruction or task node; the resource usage fields include register read-write, memory access, and I / O operations; According to the resource usage fields and the hierarchical structure of the hardware component model, establish a calling mapping path between the task operations and the target resources; the calling path includes the access order of the in-processor pipeline, caches at all levels, main memory channels, and I / O subsystems; Combine the preset processing delay parameters and bandwidth parameters in the hardware component model to estimate the occupancy cycles of each task on the resource calling path; Encapsulate the resource calling path and the corresponding occupancy cycle of each task into structured description information to form a resource calling behavior description.

[0011] Preferably, construct a data dependency graph between tasks based on a preset read-write order to represent the pre-and post-execution constraints of tasks, including: Extract the data access operations involved in the task operation sequence; the data access operations include read operations and write operations in each task, and identify the shared variables and data objects involved; According to the preset read-write order rules, analyze the access relationships of the same data object between different tasks, and identify data access dependencies such as read-after-write, write-after-read, or write-after-write; Based on the data access dependencies, establish directed dependency edges between tasks to represent the pre-and post-execution constraint relationships between tasks; Construct all the directed dependency edges into a data dependency graph; the data dependency graph is a directed acyclic graph, which is used to represent the executable order and parallel range of tasks while maintaining data consistency.

[0012] Preferably, input the task load model into the scheduling module, and use the scheduling module to start the task scheduling process under the event-driven mechanism to control the simulation execution of each task in the corresponding hardware component model, dynamically generate the resource mapping relationship between the task and the hardware component model, and record the task execution latency, data communication overhead, and power consumption metrics, including: Instantiate an event-driven scheduling module within the simulation platform; The scheduling module loads and registers the hardware component model from the model library of the simulation platform as needed; Create a global simulation clock and initialize the event queue to provide a unified time reference and event management structure for task scheduling; Write the task load model into the ready task list of the event queue, and set the initial scheduling event according to the data dependency graph; When the initial scheduling event is triggered, the scheduling module selects the currently executable task according to the preset scheduling policy and queries the real-time resource occupancy status of each hardware component model; Bind the selected currently executable task to the available hardware component model instance, generate and dynamically update the resource mapping relationship between the task and the hardware component model; During the task simulation execution, use the scheduling module to periodically send monitoring events, and collect and record the task execution latency, data communication overhead, and power consumption metrics according to the global simulation clock; When all tasks are completed or the simulation termination condition is reached, output the final resource mapping relationship and the accumulated task execution latency, data communication overhead, and power consumption metrics.

[0013] Preferably, construct a multi-index performance evaluation result including the average response latency, total data transmission bandwidth, and unit task energy consumption based on the simulation data, and make a horizontal comparison between the multi-index performance evaluation result and the preset performance baseline to obtain a performance evaluation report, including: Perform formatting processing on the simulation data, extract the single-task execution latency, data traffic, and power consumption records, and aggregate them according to the task number; Calculate the average response latency, total data transmission bandwidth, and unit task energy consumption respectively; where, average response latency = sum of all task execution latencies / number of tasks; total data transmission bandwidth = sum of all task data traffic / simulation duration; unit task energy consumption = sum of all task power consumptions / number of tasks; Determine the average response latency, the total data transmission bandwidth, and the unit task energy consumption as the multi-index performance evaluation result; Retrieve a set of baseline metrics corresponding to the current task load type from a database of preset performance baselines, and perform item-by-item difference calculations between the multi-metric performance evaluation results and the set of baseline metrics to obtain horizontal comparison data; Generate a performance evaluation report based on the multi-metric performance evaluation results and the horizontal comparison data according to a preset report template, and output it to the user interface.

[0014] A computer hardware performance evaluation system based on a simulation model, comprising: A hardware modeling unit, configured to construct hardware component models of each computing hardware unit based on a system structure description file of a target computer system; A task modeling unit, configured to extract a task operation sequence, resource call behavior, and data dependency relationship based on an instruction stream or a parallel computing graph of an application to be evaluated, and form a task load model; A simulation scheduling unit, configured to input the task load model into a scheduling module, and use the scheduling module to start a task scheduling process under an event-driven mechanism to control each task to be simulated and executed in the corresponding hardware component model, dynamically generate a resource mapping relationship between the task and the hardware component model, and record task execution latency, data communication overhead, and power consumption metrics; A simulation data generation unit, configured to determine the resource mapping relationship, task execution latency, data communication overhead, and power consumption metrics as simulation data; A performance evaluation unit, configured to construct multi-metric performance evaluation results including average response latency, total data transfer bandwidth, and energy consumption per unit task based on the simulation data, and perform a horizontal comparison between the multi-metric performance evaluation results and a preset performance baseline to obtain a performance evaluation report.

[0015] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention: The present invention overcomes the deficiencies of insufficient accuracy, lack of dynamics, and multi-dimensional performance coverage ability in the prior art evaluation means. By constructing hardware component models that strictly correspond to the target system structure and combining the task load model of the application to be evaluated, task scheduling simulation is realized under the event-driven mechanism, and the scheduling behavior and operating load of tasks on different resources can be accurately restored. The resource mapping relationship, execution latency, communication overhead, and power consumption metrics collected on this basis are uniformly defined as simulation data, which supports the construction of multi-metric performance evaluation results under a unified evaluation system and horizontal comparison with a preset performance baseline, realizing the quantification and visualization of performance differences of different computing architectures under real load conditions, and significantly improving the engineering adaptability and decision-making guiding value of the evaluation method. Description of the Drawings

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0017] Figure 1 It is the flowchart of the method provided by the embodiment of the present invention; Figure 2 It is the schematic structural diagram of the system provided by the embodiment of the present invention. Specific embodiments

[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0019] The purpose of the present invention is to provide a method and system for evaluating the performance of computer hardware based on a simulation model, which can realize the multi-index simulation evaluation of hardware performance for real task loads and improve the accuracy and interpretability of performance prediction.

[0020] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0021] Figure 1 It is the flowchart of the method provided by the embodiment of the present invention. As Figure 1 shown, the present invention provides a method for evaluating the performance of computer hardware based on a simulation model, including: Step 100: Based on the system structure description file of the target computer system, construct the hardware component models of each computing hardware unit; Step 200: Based on the instruction stream or parallel computing graph of the application to be evaluated, extract the task operation sequence, resource call behavior, and data dependency relationship to form a task load model; Step 300: Input the task load model into the scheduling module, and use the scheduling module to start the task scheduling process under the event-driven mechanism to control the simulated execution of each task in the corresponding hardware component model, dynamically generate the resource mapping relationship between the task and the hardware component model, and record the task execution delay, data communication overhead, and power consumption index at the same time; Step 400: Determine the resource mapping relationship, task execution delay, data communication overhead, and power consumption index as simulation data; Step 500: Construct a multi - metric performance evaluation result including average response delay, total data transfer bandwidth, and unit task energy consumption based on the simulation data, and make a horizontal comparison between the multi - metric performance evaluation result and a preset performance baseline to obtain a performance evaluation report.

[0022] Preferably, it further includes: Use the performance evaluation report as feedback input to iteratively optimize the key parameters of the hardware component model; the key parameters include: processing delay parameters, communication bandwidth parameters, power consumption model parameters, and resource capacity parameters.

[0023] Specifically, in this embodiment, the performance evaluation report is used as feedback input to iteratively optimize the key parameters of the hardware component model. Specifically, in this embodiment, the multi - metric performance evaluation results generated in the performance evaluation report are first parsed, the differences between the performance metrics and the preset performance baseline are calculated to obtain an error vector. Subsequently, an optimization objective function with the error vector as input is constructed, and an optimization algorithm applicable to multi - parameter non - linear problems is used for solution, such as a numerical optimization method based on gradient adjustment strategy or an evolutionary algorithm based on search strategy. The key parameter values for updating the hardware component model are obtained through the optimization process, including parameters such as processing delay, communication bandwidth, power consumption model, and resource capacity. In this embodiment, the updated parameters are written into the corresponding hardware component model, and the model is re - loaded into the simulation platform, and the simulation process is re - run under the same task load model conditions to generate a new performance evaluation report. This optimization process iterates continuously with the convergence of the error vector as the criterion, and finally obtains a hardware component model with high evaluation accuracy and general adaptability under the current load.

[0024] Preferably, based on the system structure description file of the target computer system, construct the hardware component models of each computing hardware unit, including: Read the system structure description file of the target computer system; the system structure description file includes processor structure, memory hierarchy, interconnection topology, and I / O interface configuration parameters; Parse the system structure description file, and extract the structure parameters and operating characteristic parameters of each computing hardware unit; the structure parameters and operating characteristic parameters include main frequency, cache capacity, bus width, access delay, and power consumption model coefficient; Establish corresponding hardware component models according to the extracted parameters; the hardware component models include processor model, memory model, I / O model, and interconnection structure model; Store each of the constructed hardware component models in the model library of the simulation platform in a modular form.

[0025] Specifically, in this embodiment, first, the system structure description file of the target computer system is read. This file can be sourced from the configuration list of the actual system, the system architecture diagram, or the hardware resource description information generated by an automated tool. The system structure description file contains the structural configuration data of multiple hardware components, including the processor structure, memory hierarchy, interconnection topology form, and the quantity and bandwidth configuration of I / O interfaces, providing a complete data basis for the subsequent modeling process.

[0026] Next, the system structure description file is parsed to extract the structural parameters and operating characteristic parameters of each computing hardware unit. In this embodiment, the extracted parameters include the processor main frequency, cache capacity, bus width, access latency of different levels of storage, and the power consumption model coefficients of each hardware module under typical loads. The above parameters can be confirmed through static configuration files, chip specification documents, or experimental data sheets, and are converted into structured inputs through parsing scripts for the model construction module to call. After the parameter extraction is completed, corresponding hardware component models are established according to the types and configuration requirements of different computing hardware units. In this embodiment, a processor model, a memory model, an I / O model, and an interconnection structure model are constructed. Each model encapsulates its structural and performance parameters in a modular form. The established hardware component models are uniformly stored in the model library of the simulation platform. The model library supports version management, on-demand loading, and parameterized instantiation of the models, and is used for the simulation scheduling and performance evaluation process calls of subsequent task loads.

[0027] Preferably, based on the instruction stream or parallel computation graph of the application to be evaluated, the task operation sequence, resource call behavior, and data dependency relationship are extracted to form a task load model, including: Obtain the instruction stream or parallel computation graph of the application to be evaluated; the instruction stream is the instruction sequence generated after the application is compiled or during the running process; the parallel computation graph represents the control and data dependency relationship between multiple tasks; Perform semantic parsing on the instruction stream or parallel computation graph to extract the task operation types therein; the task operation types include arithmetic operations, logical judgments, memory access operations, and communication instructions; Analyze the call relationship between the instruction stream or parallel computation graph and the hardware resources to determine the call paths and occupancy cycles of each task for the processor, memory, cache, and I / O resources, and generate a resource call behavior description; Based on the preset read / write order, construct a data dependency relationship graph between tasks to represent the front-to-back execution constraints of tasks; Encapsulate the task operation types, the resource call behavior description, and the front-to-back execution constraints into structured data to form a task load model.

[0028] In this embodiment, first, the instruction stream or parallel computation graph of the application to be evaluated is obtained. The instruction stream can be extracted from the target application by a static compilation tool or a runtime analyzer, reflecting the instruction sequence during its specific execution process; the parallel computation graph represents the control relationship and data dependency structure among various computing tasks, and is often generated by a program modeling tool, a data flow analysis module, or a task scheduler, and is used to reflect the possible concurrent execution patterns among tasks. The above two types of inputs serve as the basic data sources for task load modeling.

[0029] Subsequently, semantic parsing is performed on the instruction stream or parallel computation graph to extract the types of task operations contained therein, including arithmetic operations, logical judgments, memory accesses, and communication interactions, etc. In this embodiment, a structured syntax tree or an instruction classification table is used to identify and classify task operations, and in combination with the distribution structure of each hardware resource in the system, the call relationship between task instructions and the processor, memory, cache, and I / O resources is further analyzed. On this basis, in combination with the task granularity, instruction complexity, and the latency and bandwidth parameters in the preset resource model, the access path and occupancy cycle of each task operation to the corresponding resources are calculated, and a corresponding resource call behavior description is generated.

[0030] Finally, according to the extracted instruction access order, a data dependency graph between tasks is constructed to identify the front and back execution constraints formed by data read and write conflicts between tasks. This dependency relationship is represented by directed edges to represent task scheduling restrictions, forming a structured task scheduling graph. In this embodiment, the task operation type, resource call behavior description, and front and back execution constraints are uniformly encapsulated into structured data to form a complete task load model, which is used as the input of the simulation scheduling module for subsequent mapping scheduling and performance evaluation simulation in the hardware component model.

[0031] Preferably, analyze the call relationship between the instructions or nodes and the hardware resources to determine the call paths and occupancy cycles of each task to the processor, memory, cache, and I / O resources, and generate a resource call behavior description, including: Parse the task nodes in the instruction stream or parallel computation graph, and identify the resource usage fields involved in each instruction or task node; the resource usage fields include register read and write, memory access, and I / O operations; According to the resource usage fields and the hierarchical structure of the hardware component model, establish a call mapping path between task operations and target resources; the call path includes the access order of the in-processor pipeline, each level of cache, the main memory channel, and the I / O subsystem; In combination with the preset processing latency parameters and bandwidth parameters in the hardware component model, estimate the occupancy cycle of each task on the resource call path; Encapsulate the resource call path of each task and the corresponding occupancy cycle into structured description information to form a resource call behavior description.

[0032] In this embodiment, to analyze the calling relationship between each task and hardware resources in the application to be evaluated, the task nodes in the instruction stream or parallel computing graph are first parsed. Specifically, the system traverses each instruction or task node, extracts the field information related to resource access therein, and the resource usage fields include read and write operations of registers, various memory address access instructions, and I / O calls related to peripheral interaction. This parsing process can be implemented based on an instruction format template or a static analysis tool to extract structured resource usage data, laying a foundation for subsequent call path derivation.

[0033] After the extraction of resource usage fields is completed, in combination with the hierarchical structure of the hardware component model, a mapping path is established for the resource access process involved in each task operation. In this embodiment, according to the processor architecture and memory architecture, the access order including the processor pipeline, level-1 cache, level-2 cache, main memory interface channel, and peripheral I / O module is sequentially identified to form a complete call path from the task to the target resource. For each resource access operation, the system also marks the starting node (such as the computing core) and the ending node (such as the target memory block or I / O port) in the path, and retains the logical access order and data flow information in the path.

[0034] After the above call path is established, in combination with the processing delay parameters and bandwidth parameters of each resource module in the pre-constructed hardware component model, the access time of the task on the call path is estimated. The system comprehensively calculates the occupancy cycles of each instruction at each resource node based on factors such as the type of each instruction, the bandwidth limit, and concurrent conflicts on the mapped path. Finally, the resource call path and occupancy cycle corresponding to each task operation are integrated into a structured data description form to form a resource call behavior description, which is used as an important part of the task load model for subsequent scheduling and simulation modeling calls.

[0035] Preferably, a data dependency graph between tasks is constructed based on a preset read-write order to represent the execution constraints before and after tasks, including: extracting the data access operations involved in the task operation sequence; the data access operations include read operations and write operations in each task, and identifying the shared variables and data objects involved; analyzing the access relationship of the same data object between different tasks according to the preset read-write order rules, and identifying data access dependencies such as read-after-write, write-after-read, or write-after-write; establishing directed dependency edges between tasks based on the data access dependencies to represent the execution constraint relationship before and after tasks; Construct all the directed dependency edges into a data dependency graph; the data dependency graph is a directed acyclic graph, which is used to represent the executable order and parallel range of tasks while maintaining data consistency.

[0036] In this embodiment, to construct the data dependency graph between tasks, the system first scans the task operation sequences that have been extracted and completed, and identifies the data access operations involved in each task, including read operations and write operations. Through semantic parsing and variable mapping mechanisms, the system identifies the shared variables or data objects involved in the operations, such as global variables, shared memory blocks, or I / O buffers. This process can be executed based on the abstract syntax tree or the intermediate representation IR to ensure the accurate establishment of the pointing relationship between data objects and operations. Subsequently, the system pairwise matches the data access pairs in all tasks according to the preset read-write order rules, and analyzes whether there are data access dependency relationships. The dependency relationships include three typical types: read after write (RAW), write after read (WAR), and write after write (WAW). The system traverses the access situations of the same data object between task pairs. When it detects that the previous task performs a write operation on the object and the subsequent task has a read or write access behavior on the same object, it determines that there is an execution order constraint. The system generates a directed edge for this dependency relationship to identify the scheduling order required by data consistency.

[0037] After analyzing the access relationships of all task pairs, the system organizes and constructs the identified directed dependency edges into a data dependency graph according to task numbers. This relationship graph is a directed acyclic graph (DAG), where each node represents a specific task and each edge represents the front-back execution constraint formed by data dependency. Through this graph structure, the system can clarify the executable order and parallel execution range between tasks while maintaining data consistency, providing structured support for the subsequent scheduling module to initialize scheduling events and perform dependency checks.

[0038] Preferably, input the task load model into the scheduling module, and use the scheduling module to start the task scheduling process under the event-driven mechanism to control the simulation execution of each task in the corresponding hardware component model, dynamically generate the resource mapping relationship between the task and the hardware component model, and record the task execution latency, data communication overhead, and power consumption metrics, including: Instantiate an event-driven scheduling module within the simulation platform; The scheduling module loads and registers the hardware component model from the model library of the simulation platform as needed; Create a global simulation clock and initialize the event queue to provide a unified time reference and event management structure for task scheduling; Write the task load model into the ready task list of the event queue, and set the initial scheduling event according to the data dependency graph; When the initial scheduling event is triggered, the scheduling module selects the currently executable task according to the preset scheduling policy, and queries the real-time resource occupancy status of each hardware component model; Bind the selected currently executable task to the available hardware component model instance, and generate and dynamically update the resource mapping relationship between the task and the hardware component model; During the task simulation execution, the scheduling module periodically sends monitoring events, and collects and records the task execution delay, data communication overhead, and power consumption metrics according to the global simulation clock; When all tasks are completed or the simulation termination condition is reached, output the final resource mapping relationship and the accumulated task execution delay, data communication overhead, and power consumption metrics.

[0039] In this embodiment, the simulation platform first instantiates the event-driven scheduling module, and loads hardware component model instances such as processor models, memory models, I / O models, and interconnection structure models from the model library as needed through the interface. Subsequently, the platform creates a global simulation clock and initializes the event queue, pushes basic events such as system startup events and resource status update events into the queue, and establishes a unified time reference and event management structure based on this. After initialization, the system writes the constructed task load model into the ready task list of the event queue, and at the same time generates corresponding dependency detection events for each task according to the aforementioned data dependency graph to ensure that the subsequent scheduling process can correctly identify the task executable state.

[0040] When the first batch of dependency detection events in the ready queue are triggered, the scheduling module selects the currently executable task using a preset scheduling policy (such as based on priority or the principle of minimizing resource occupancy). The scheduling module queries the resource occupancy status of each hardware component model in real time, judges the availability of processor cores, cache levels, memory channels, and I / O interfaces, and binds the selected task to the hardware component model instance that meets the resource constraints. After the binding is completed, the scheduling module immediately updates the resource mapping table between the task and the hardware component model, and generates corresponding task start events and estimated completion events, and writes them into the event queue again to realize the closed-loop drive of the task scheduling and execution process.

[0041] During the task simulation execution, the scheduling module periodically triggers monitoring events as the global simulation clock advances, reads the status registers or counters of each hardware component model, and collects and accumulates the task execution delay, data traffic, and energy consumption statistics. If the monitoring event detects that all tasks have been executed, or the simulation reaches the preset termination condition, the scheduling module ends the clock advance and freezes the event queue; at this time, the system outputs the final resource mapping relationship, as well as the accumulated task execution delay, total data communication overhead, and power consumption metrics, as the complete simulation data results, providing an input basis for subsequent multi-metric performance evaluation.

[0042] Preferably, construct a multi-metric performance evaluation result including the average response delay, total data transmission bandwidth, and energy consumption per unit task based on the simulation data, and make a horizontal comparison between the multi-metric performance evaluation result and the preset performance baseline to obtain a performance evaluation report, including: Perform formatting processing on the simulation data, extract the single-task execution delay, data traffic, and power consumption records, and aggregate them according to the task number; Calculate the average response delay, total data transmission bandwidth, and energy consumption per unit task respectively; among them, average response delay = sum of all task execution delays / number of tasks; total data transmission bandwidth = sum of all task data traffic / simulation duration; energy consumption per unit task = sum of all task power consumptions / number of tasks; Determine the average response delay, the total data transmission bandwidth, and the energy consumption per unit task as the multi-metric performance evaluation result; Retrieve the baseline metric set corresponding to the current task load type from the database of the preset performance baseline, and perform item-by-item difference calculation between the multi-metric performance evaluation result and the baseline metric set to obtain horizontal comparison data; Generate a performance evaluation report based on the multi-metric performance evaluation result and the horizontal comparison data according to the preset report template, and output it to the user interface.

[0043] In this embodiment, after the simulation ends, the obtained simulation data is first subjected to formatting processing. The system reads the original records of task execution delay, data traffic, and power consumption according to the task number, aggregates the data generated by the same task in different time slices into a unified data table, and eliminates invalid or duplicate sampling points to obtain a structured single-task metric set. This set is stored in the form of a four-tuple of "task number - execution delay - traffic - power consumption", providing a clear data dimension for subsequent statistical operations.

[0044] Subsequently, the system calls an aggregation function in the statistics module to calculate global performance metrics: First, the sum of the execution latencies of all tasks is divided by the number of tasks to obtain the average response latency; Second, the sum of the data communication volumes of all tasks is divided by the simulation duration to obtain the total data transfer bandwidth; Third, the sum of the power consumption values of all tasks is divided by the number of tasks to obtain the energy consumption per task. The above three calculation results are encapsulated into a triple of "average response latency, total data transfer bandwidth, energy consumption per task", which constitutes the multi-metric performance evaluation result, and is written into the performance result table together with the task number and simulation configuration.

[0045] After generating the multi-metric performance evaluation result, the system connects to a preset performance baseline database, retrieves the baseline metric set corresponding to the current task load type, and calculates the difference values item by item in the same dimension to obtain the horizontal comparison data. Finally, the system fills the multi-metric performance evaluation result and the horizontal comparison data into the report structure according to the preset report template, automatically generates a performance evaluation report including a performance radar chart, a difference bar chart, and an index summary table, and presents it through the user interface or exports it as a PDF file for developers and hardware designers to quickly compare the performance differences of different architectures or parameter configurations.

[0046] Corresponding to the above method, as Figure 2 shown, this embodiment also provides a computer hardware performance evaluation system based on a simulation model, including: A hardware modeling unit for constructing hardware component models of each computing hardware unit based on the system structure description file of the target computer system; A task modeling unit for extracting task operation sequences, resource call behaviors, and data dependency relationships based on the instruction stream or parallel computing graph of the application to be evaluated, and forming a task load model; A simulation scheduling unit for inputting the task load model into the scheduling module, and using the scheduling module to start the task scheduling process under the event-driven mechanism to control the simulated execution of each task in the corresponding hardware component model, dynamically generating the resource mapping relationship between the task and the hardware component model, and recording the task execution latency, data communication overhead, and power consumption metrics at the same time; A simulation data generation unit for determining the resource mapping relationship, task execution latency, data communication overhead, and power consumption metrics as simulation data; A performance evaluation unit for constructing a multi-metric performance evaluation result including average response latency, total data transfer bandwidth, and energy consumption per task according to the simulation data, and making a horizontal comparison between the multi-metric performance evaluation result and a preset performance baseline to obtain a performance evaluation report.

[0047] The beneficial effects of the present invention are as follows: (1) By parsing the target computer system description file into modular hardware component models such as processor models, memory models, I / O models, and interconnection structure models, the present invention realizes a high-fidelity simulation environment that corresponds one-to-one with the real hardware structure. Compared with the traditional evaluation methods that rely on benchmark scores, it can accurately reflect the latency, bandwidth, and power consumption characteristics of different hardware units at the early stage of design, thereby significantly improving the accuracy of performance prediction and shortening the hardware prototype verification cycle.

[0048] (2) The present invention converts the instruction stream or parallel computing graph of the application to be evaluated into a task load model, and simulates the dynamic interaction between tasks and hardware resources within an event-driven scheduling framework, generating the task-resource mapping relationship in real time and synchronously recording the execution latency, communication overhead, and power consumption metrics, breaking through the limitations of the existing technology in terms of single evaluation dimension and lack of runtime data association, and realizing multi-dimensional performance collection related to the load and traceable resources.

[0049] (3) By calculating the average response latency, total data transfer bandwidth, and unit task energy consumption of the simulation data and making a horizontal comparison with the preset performance baseline, the performance evaluation report output by the present invention can intuitively display the advantages and bottlenecks of each hardware configuration under the real load in the form of quantitative differences. This report, combined with visualization charts, provides a quick decision-making basis for chip designers and system integrators, reducing the manpower and experimental costs of multi-scheme comparison.

[0050] (4) The present invention further introduces a closed-loop feedback mechanism for the performance evaluation results, and uses an optimization algorithm to iteratively adjust key parameters such as processing delay, communication bandwidth, power consumption model, and resource capacity, continuously correcting the accuracy of the hardware component model. This self-learning iterative process enables the simulation environment to continuously adapt to the evolution of the hardware and the change of the load, enhancing the generality and sustainable application value of the method, and laying a foundation for subsequent automated exploration of the hardware design space.

[0051] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0052] Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those of ordinary skill in the art, there will be changes in the specific implementation manner and application scope according to the idea of the present invention. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for evaluating the performance of computer hardware based on a simulation model, characterized in that Including: Based on the system structure description file of the target computer system, construct the hardware component models of each computing hardware unit; Based on the instruction stream or parallel computing graph of the application to be evaluated, extract the task operation sequence, resource call behavior, and data dependency relationship to form a task workload model; Input the task workload model into the scheduling module, and use the scheduling module to start the task scheduling process under the event-driven mechanism to control each task to be simulated and executed in the corresponding hardware component model, dynamically generate the resource mapping relationship between the task and the hardware component model, and record the task execution latency, data communication overhead, and power consumption metrics at the same time; Determine the resource mapping relationship, task execution latency, data communication overhead, and power consumption metrics as simulation data; Construct a multi-metric performance evaluation result including the average response latency, total data transfer bandwidth, and energy consumption per unit task based on the simulation data, and make a horizontal comparison between the multi-metric performance evaluation result and the preset performance baseline to obtain a performance evaluation report.

2. The method for evaluating computer hardware performance based on a simulation model according to claim 1, wherein Also including: Use the performance evaluation report as feedback input to iteratively optimize the key parameters of the hardware component model; The key parameters include: processing delay parameter, communication bandwidth parameter, power consumption model parameter, and resource capacity parameter.

3. The method for evaluating computer hardware performance based on a simulation model according to claim 1, characterized in that Based on the system structure description file of the target computer system, construct the hardware component models of each computing hardware unit, including: Read the system structure description file of the target computer system; the system structure description file includes processor structure, memory hierarchy, interconnection topology, and I / O interface configuration parameters; Parse the system structure description file to extract the structure parameters and operating characteristic parameters of each computing hardware unit; the structure parameters and operating characteristic parameters include main frequency, cache capacity, bus width, access delay, and power consumption model coefficient; Establish corresponding hardware component models according to the extracted parameters; the hardware component models include processor model, memory model, I / O model, and interconnection structure model; Store each of the constructed hardware component models in the model library of the simulation platform in a modular form.

4. The method for evaluating the performance of computer hardware based on a simulation model according to claim 1, wherein Based on the instruction stream or parallel computing graph of the application to be evaluated, extract the task operation sequence, resource call behavior, and data dependency relationship to form a task workload model, including: Obtain the instruction stream or parallel computing graph of the application to be evaluated; the instruction stream is an instruction sequence generated after the application program is compiled or during the running process; the parallel computing graph represents the control and data dependency relationship between multiple tasks; Perform semantic parsing on the instruction stream or parallel computing graph to extract the task operation types; the task operation types include arithmetic operations, logical judgments, memory access operations, and communication instructions; Analyze the call relationship between the instruction stream or parallel computing graph and the hardware resources to determine the call paths and occupancy cycles of each task for the processor, memory, cache, and I / O resources, and generate a resource call behavior description; Based on the preset read / write order, construct a data dependency relationship graph between tasks to represent the front-back execution constraints of tasks; Package the task operation types, the resource call behavior description, and the front-back execution constraints into structured data to form a task workload model.

5. The method for evaluating the performance of computer hardware based on a simulation model according to claim 4, wherein Analyze the calling relationship between the instructions or nodes and the hardware resources to determine the calling paths and occupancy cycles of each task for the processor, memory, cache, and I / O resources, and generate a description of the resource calling behavior, including: Parse the task nodes in the instruction stream or parallel computing graph, and identify the resource usage fields involved in each instruction or task node; the resource usage fields include register reading and writing, memory access, and I / O operations; According to the resource usage fields and the hierarchical structure of the hardware component model, establish a calling mapping path between the task operations and the target resources; the calling path includes the access order of the in-processor pipeline, caches at all levels, main memory channels, and I / O subsystems; Combined with the preset processing delay parameters and bandwidth parameters in the hardware component model, estimate the occupancy cycles of each task on the resource calling path; Package the resource calling path and the corresponding occupancy cycle of each task into structured description information to form a description of the resource calling behavior.

6. The method for evaluating computer hardware performance based on a simulation model according to claim 4, characterized in that Based on the preset read-write order, construct a data dependency graph between tasks to represent the execution constraints before and after tasks, including: Extract the data access operations involved in the task operation sequence; the data access operations include read operations and write operations in each task, and identify the shared variables and data objects involved; According to the preset read-write order rules, analyze the access relationships of the same data object between different tasks, and identify data access dependencies such as read-after-write, write-after-read, or write-after-write; Based on the data access dependencies, establish directed dependency edges between tasks to represent the execution constraint relationships before and after tasks; Construct all the directed dependency edges into a data dependency graph; the data dependency graph is a directed acyclic graph, which is used to represent the executable order and parallel range of tasks while maintaining data consistency.

7. The method for evaluating computer hardware performance based on a simulation model according to claim 4, characterized in that Input the task load model into the scheduling module, and use the scheduling module to start the task scheduling process under the event-driven mechanism to control the simulation execution of each task in the corresponding hardware component model, dynamically generate the resource mapping relationship between the task and the hardware component model, and record the task execution latency, data communication overhead, and power consumption metrics, including: Instantiate an event-driven scheduling module in the simulation platform; The scheduling module loads and registers the hardware component model as needed from the model library of the simulation platform; Create a global simulation clock and initialize the event queue to provide a unified time reference and event management structure for task scheduling; Write the task load model into the ready task list of the event queue, and set the initial scheduling event according to the data dependency graph; When the initial scheduling event is triggered, the scheduling module selects the currently executable task according to the preset scheduling strategy and queries the real-time resource occupancy status of each hardware component model; Bind the selected currently executable task to the available hardware component model instance, and generate and dynamically update the resource mapping relationship between the task and the hardware component model; During the task simulation execution, the scheduling module is used to periodically send monitoring events, and the task execution delay, data communication overhead, and power consumption metrics are collected and recorded based on the global simulation clock; When all tasks are completed or the simulation termination condition is reached, the final resource mapping relationship, as well as the cumulative task execution delay, data communication overhead, and power consumption metrics, are output.

8. The method for evaluating computer hardware performance based on a simulation model according to claim 1, characterized in that Based on the simulation data, a multi-metric performance evaluation result including average response delay, total data transmission bandwidth, and energy consumption per task is constructed, and the multi-metric performance evaluation result is horizontally compared with a preset performance baseline to obtain a performance evaluation report, including: The simulation data is formatted, the single-task execution delay, data communication volume, and power consumption records are extracted, and aggregated according to the task number; The average response delay, total data transmission bandwidth, and energy consumption per task are calculated respectively; among them, average response delay = sum of all task execution delays / number of tasks; total data transmission bandwidth = sum of all task data communication volumes / simulation duration; energy consumption per task = sum of all task power consumptions / number of tasks; The average response delay, the total data transmission bandwidth, and the energy consumption per task are determined as the multi-metric performance evaluation result; The baseline metric set corresponding to the current task load type is retrieved from the database of the preset performance baseline, and the multi-metric performance evaluation result is calculated item by item with the baseline metric set to obtain horizontal comparison data; According to the preset report template, the multi-metric performance evaluation result and the horizontal comparison data are used to generate a performance evaluation report and output it to the user interface.

9. A computer hardware performance evaluation system based on a simulation model, characterized in that, Including: A hardware modeling unit for constructing hardware component models of each computing hardware unit based on the system structure description file of the target computer system; A task modeling unit for extracting task operation sequences, resource call behaviors, and data dependencies based on the instruction stream or parallel computing graph of the application to be evaluated, and forming a task load model; A simulation scheduling unit for inputting the task load model into the scheduling module, and using the scheduling module to start the task scheduling process under the event-driven mechanism to control each task to be simulated and executed in the corresponding hardware component model, dynamically generating the resource mapping relationship between the task and the hardware component model, and simultaneously recording the task execution delay, data communication overhead, and power consumption metrics; A simulation data generation unit for determining the resource mapping relationship, task execution delay, data communication overhead, and power consumption metrics as simulation data; A performance evaluation unit for constructing a multi-metric performance evaluation result including average response delay, total data transmission bandwidth, and energy consumption per task based on the simulation data, and horizontally comparing the multi-metric performance evaluation result with a preset performance baseline to obtain a performance evaluation report.

Citation Information

Patent Citations

  • System modeling evaluation method and device, electronic equipment and storage medium

    CN116501594A

  • Method and electronic device for target system performance estimation

    CN118350168A

  • Autonomous device optimization architecture generation method

    CN118760522A

  • Wafer level chip system design space construction and rapid parameter search method

    CN120163113A

  • Performance Modeling and Analysis of Artificial Intelligence (AI) Accelerator Architectures

    US20220366267A1

Cited By

  • Method, computing device, medium and program product for predicting performance of computing system executed by application program

    CN120631705A

  • Methods, computing devices, media, and program products for predicting performance of a computing system on which an application is run

    CN120631705B

  • Communication network real-time simulation method and system of large model enabling OMNEST

    CN121396370A

  • Performance evaluation method and device of AI processor, equipment and medium

    CN122432006A