Comprehensive performance evaluation device and method for domestic credit and creativity platform hardware
Through the customized instruction set and hybrid load generation technology that adapts to the domestic information innovation platform, the problems of single granularity and insufficient scenario adaptability in the existing technology are solved, and a comprehensive performance evaluation of the domestic information innovation platform hardware is achieved, and intuitive performance analysis tools are provided.
Patent Information
- Application Number
- CN202510506116.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-08-01
AI Technical Summary
When evaluating the hardware of the domestic information innovation platform, the existing technology has a single indicator dimension, insufficient scenario adaptability, and extensive evaluation granularity. It cannot fully reflect the performance of the hardware in complex scenarios, and lacks support for general computing tasks such as cloud computing and big data processing.
A performance evaluation device for domestic processors, storage devices and network components was designed. By adapting to its customized instruction set, hybrid load generation technology was introduced to build a multi-dimensional evaluation method, and combined with hardware performance portrait methods, intuitive comprehensive performance evaluation results were generated.
It has achieved a comprehensive performance evaluation of the hardware of the domestic information innovation platform, improved the scientificity and accuracy of the evaluation, and can truly reflect the responsiveness and resource collaboration efficiency of the hardware in complex scenarios, providing intuitive performance analysis tools.
Smart Images

Figure CN120407304A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of electro-digital data processing, and further relates to a multi-dimensional and scenario-based comprehensive performance evaluation device and imaging method for domestic Xinchuang platform hardware in the technical field of detecting or locating faulty hardware through testing. The present invention can be used for processors, storage devices and network hardware of domestic Xinchuang platforms to simulate the mixed load of real business scenarios, and evaluate the performance quantification of the response robustness and resource cooperation efficiency of domestic Xinchuang platform hardware in high-complexity scenarios. Background Art
[0002] With the in-depth promotion of the domestic information technology application innovation strategy, cloud computing infrastructures, artificial intelligence training, and big data processing clusters built on autonomous and controllable hardware platforms have become the core carriers to support digital transformation. Domestic processors (such as Ascend, Haiguang, Loongson), storage devices, and network components show significant differences in instruction set architectures, accelerator module designs, and energy efficiency management mechanisms, and their performance directly affects the stability of business systems and resource utilization efficiency in key fields such as smart cities, fintech, and intelligent manufacturing. The accuracy of hardware performance evaluation is directly related to the ability to truly quantify the performance of hardware in actual application scenarios, thus affecting the scientific decision-making of hardware selection and system optimization.
[0003] However, the current mainstream hardware performance evaluation system still heavily relies on international general benchmark test tools, and these methods expose systematic defects when facing the domestic hardware ecosystem: First, traditional test models are designed based on general architectures such as x86 / ARM, and fail to adapt to the customized instruction extensions unique to domestic chips, resulting in significant deviations in the quantification of core indicators such as the true computing power density and energy efficiency ratio of hardware; Second, most existing evaluation processes adopt a static load mode (such as floating-point operation stress tests with a fixed number of threads), lacking the ability to simulate the mixed load of real business scenarios such as the elastic scaling of cloud computing virtualization resources, the dynamic adjustment of batch processing scale in artificial intelligence training, and the high-concurrency writing in big data processing, making it difficult for test results to reflect the resource cooperation efficiency of hardware under complex multi-task competition; Third, international benchmark tests overly focus on peak performance indicators (such as the peak TFLOPS of single-precision floating-point operations and sequential read / write bandwidth), ignoring systematic efficiency decay problems such as power consumption drift and cache coherence bottlenecks during long-term high-load operation of hardware, resulting in a serious disconnect between evaluation data and the actual performance of hardware in the production environment.
[0004] Xi'an Chaoyue Shentai Information Technology Co., Ltd. discloses a training performance testing method and system for domestic heterogeneous platform artificial intelligence acceleration cards in its patent document "A Performance Testing Method and System for Domestic Heterogeneous Platform Artificial Intelligence Acceleration Cards" (application number CN 202311431943.5, publication number CN117370088A). The method includes the following steps: confirming test indicators, which is the speed of training the model to the target quality indicator in the terminal system; building a test environment and preparing the software and hardware platform; using the performance comparison under a fixed AI framework and model as the evaluation indicator; when the training result reaches the accuracy indicator, recording the time, performance result, and power consumption result used for model training to complete the test. This invention can provide a general test method for the training performance testing of different artificial intelligence acceleration cards under domestic heterogeneous computing platforms. However, this method still has the following two deficiencies: First, this invention only conducts performance testing on a single hardware of the artificial intelligence acceleration card and cannot perform performance testing and comprehensive evaluation on the platform's general-purpose processor, memory device, and network device. Second, this invention only focuses on artificial intelligence training tasks and cannot perform performance testing on general computing task scenarios such as cloud computing and big data processing. Summary of the Invention
[0005] The purpose of the present invention is to provide an evaluation device and method for the comprehensive performance of hardware in domestic Xinchuang platforms in view of the above-mentioned deficiencies of the prior art, aiming to solve the technical bottleneck problems such as single-dimensional test indicators, insufficient scenario adaptability, and rough evaluation granularity in traditional test methods.
[0006] The technical idea for achieving the object of the present invention is that, in view of the problem that traditional test tools are designed for general architectures such as x86 / ARM and fail to adapt to the specific customized instruction extensions of domestic chips, the present invention designs a performance evaluation load task generation mechanism for the instruction set architectures of domestic processors, storage devices and network components. By adapting the software architecture to the specific instruction sets of domestic hardware devices, the evaluation deviation caused by architecture mismatch is eliminated, so as to solve the problem that the specific customized instruction extensions of domestic chips cannot be adapted, and ensure the accuracy and reliability of test results. In view of the limitation of the existing static load test that it cannot reflect the hardware performance in complex scenarios, the present invention introduces a hybrid load generation technology. By configuring the load pressure of various hardware, the hybrid dynamic load on each hardware in the real business scenario is reproduced, and the comprehensive performance of the hardware platform in actual applications is more realistically reflected. In view of the deficiency of the existing evaluation indicators that they overly focus on peak performance while ignoring energy efficiency and stability, the present invention constructs a multi-dimensional hardware performance evaluation method. By analyzing the data of the domestic information and communication technology (ICT) platform to be evaluated during task operation, considering the changing trend of hardware performance over time, and excluding the index deviation caused by systematic efficiency decay phenomena, a comprehensive hardware performance evaluation result is provided. In view of the limitation of the existing technology that only the performance of a single hardware, i.e., the artificial intelligence acceleration card, is tested and the general-purpose processor, memory device and network device of the platform cannot be comprehensively evaluated, the evaluation device of the present invention covers the general-purpose processor, memory device and network device to achieve the comprehensive performance test and evaluation of the domestic ICT platform. In view of the problem that the existing technology only focuses on artificial intelligence training tasks and cannot adapt to general computing task scenarios such as cloud computing and big data processing, the evaluation method of the present invention accurately simulates diversified scenarios such as cloud computing and big data processing to ensure the versatility and scenario adaptability of performance testing. Finally, the present invention uses a hardware comprehensive performance profiling method. By analyzing the data of the domestic ICT platform to be evaluated during task operation, considering the changing trend of hardware performance over time, excluding the index deviation caused by systematic efficiency decay phenomena, and using the benchmark comparison method for quantitative scoring, an intuitive radar chart of the hardware comprehensive performance evaluation of the domestic ICT platform is generated, which helps users quickly identify the performance advantages and potential bottlenecks of the hardware platform in different scenarios and provides a scientific basis for hardware selection and system optimization.
[0007] To achieve the above object, the evaluation device of the present invention includes a task construction module, a backend parsing module, a data storage module, a load execution module, a real-time monitoring module, and a hardware platform performance profiling module; where:
[0008] The task construction module executes customized hybrid load tasks according to the characteristics of different load scenarios; according to the initialization parameters of the evaluation task configured by the evaluator, supports Web page operations and terminal command line operations, generates a description file of the hybrid load task for hardware platform evaluation, and sends the description file to the backend parsing module;
[0009] After receiving the hardware platform evaluation hybrid load task description file sent by the task construction module, the backend parsing module parses the hardware platform evaluation hybrid load task description file into specific complex task instructions for the domestic Xinchuang platform to be evaluated and sends them to the load execution module; at the same time, the backend parsing module also receives the hardware platform evaluation hybrid load task execution process information and execution results returned by the load execution module, and stores the hardware platform evaluation hybrid load task execution process information and execution results in the data storage module;
[0010] The data storage module stores the execution results of all hardware platform evaluation hybrid load tasks for access and subsequent analysis by the hardware platform performance profiling module;
[0011] The load execution module is deployed on the domestic Xinchuang platform to be evaluated, receives the hardware platform evaluation hybrid load task sent by the backend parsing module and executes it, and at the same time returns the hardware platform evaluation hybrid load task execution process information and execution results to the backend parsing module;
[0012] The real-time monitoring module is used to monitor the execution situation and intermediate state of the hardware platform evaluation hybrid load task. By reading the hardware platform evaluation hybrid load task execution process information stored in the data storage module, it real-time feedbacks the task execution process situation to help the evaluator understand the intermediate process of the task execution;
[0013] The hardware platform performance profiling module is used to, after the execution of the hardware platform evaluation hybrid load task is completed, read the execution result data of the hybrid load task from the data storage module. By analyzing the data of the domestic Xinchuang platform to be evaluated during the task operation, considering the change trend of hardware performance over time, excluding the index deviation caused by the systematic efficiency decay phenomenon, using the benchmark comparison method for quantitative scoring, and presenting the hardware performance profile in the form of a radar chart to visually present the key dimensions of computing performance, memory usage efficiency, storage bandwidth, and network throughput.
[0014] The specific steps of a comprehensive performance evaluation method for the hardware of a domestic Xinchuang platform by the evaluation device of the present invention are as follows:
[0015] Step 1, the task construction module executes customized hybrid load tasks according to the characteristics of different load scenarios;
[0016] Step 2, the task construction module generates a description file of the hardware platform evaluation hybrid load task according to the evaluation task initialization parameters configured by the evaluator, and sends the description file to the backend execution module;
[0017] Step 3: The backend parsing module parses the task description file for the domestic Xinchuang platform to be evaluated and sends it to the load execution module; meanwhile, it receives the execution process information and execution results of the hardware platform evaluation hybrid load task returned by the load execution module and stores them in the data storage module.
[0018] Step 4: The load execution module deployed on the domestic Xinchuang platform to be evaluated executes the hardware platform evaluation hybrid load task instruction and returns the evaluation hybrid load task execution process information and execution results to the backend parsing module.
[0019] Step 5: The real-time monitoring module reads the hardware platform evaluation hybrid load task execution process information stored in the data storage module and provides real-time feedback on the task execution process to help the evaluator understand the intermediate process of task execution.
[0020] Step 6: The hardware platform performance profiling module reads the task execution process information and execution result data stored in the data storage module, analyzes the data of the domestic Xinchuang platform to be evaluated during task operation, considers the change trend of hardware performance over time, eliminates the index deviation caused by systematic efficiency decay, conducts quantitative scoring using the benchmark comparison method, and uses a radar chart to display the hardware performance profile, intuitively presenting the key dimensions of computing performance, memory usage efficiency, storage bandwidth, and network throughput.
[0021] Furthermore, the hybrid load task refers to mapping scalar operation instructions to the vectorized special instruction set of domestic processors, replacing traditional parallel vector instructions with Ascend vmm.4x4f32 matrix multiplication instructions, and injecting instruction gap control parameters that conform to the characteristics of domestic chips; for storage devices, by adapting to the scheduling strategy characteristics of domestic storage hardware, integrating a data sharding verification mechanism based on national cryptography algorithms into the load task; in network component testing, by adapting to the jumbo frame acceleration and zero-copy mechanisms of domestic network hardware, constructing a test traffic model that conforms to the optimized protocol of the domestic Xinchuang platform, and performing traffic coloring by inserting a custom field containing device fingerprint identification information into the data link layer header to achieve the localization feature marking of the protocol header.
[0022] Furthermore, the evaluation task initialization parameters include the current hardware platform to be tested, the test items to be executed, and the specific execution parameters for each test item; the test items to be executed include the CPU computing resource load test item, the memory pressure load test item, the network card hardware pressure load test item, and the storage device IO pressure load test item; the specific execution parameters for each test item include: the computing task mode and maximum pressure in the CPU computing resource load test item; the total memory access read and write volume and the proportion of sequential and random read and write in the memory pressure load test item.
[0023] Further, the description file includes: hardware performance evaluation load, test conditions, test metrics, and execution order information.
[0024] Further, the instruction for executing the hybrid load task for hardware platform evaluation means that for a general-purpose processor, the load execution module runs high-concurrency computing tasks and applies high load pressure to evaluate the performance of the general-purpose processor in executing tasks with strong computing requirements; for a memory device, the load execution module performs large-scale data read and write operations to comprehensively test the key metrics of the sequential read and write speed and random read and write speed of the memory device, so as to evaluate the performance of the memory device in tasks with strong memory access requirements; for a network device, the load execution module generates a high-traffic packet network load, examines the transmission efficiency and reliability of the network load, and thus evaluates its performance in tasks with strong communication requirements.
[0025] The running of high-concurrency computing tasks means that by creating multiple thread or process execution units, intensive mathematical operations are performed simultaneously to fully occupy the computing resources of the processor.
[0026] The key metrics for testing the sequential read and write speed and random read and write speed of the memory device refer to: recording the key metrics of the read and write speed and latency of the memory device during the sequential read and write and random read and write operations of the memory device, that is, by continuously reading and writing large data blocks, measuring the sustained throughput and latency of the memory device; the random read and write test is to randomly select positions in the memory address space for small data read and write operations to evaluate the response speed and efficiency of the memory device in non-sequential access mode.
[0027] The transmission efficiency and reliability of the network load refer to that the load execution module generates a high-traffic packet to construct a network load scenario, monitors the key parameters such as throughput, error rate, and connection stability of the network device during the test, and evaluates the reliability and performance of the network device in tasks with strong communication requirements.
[0028] Further, the elimination of the index offset caused by the systematic performance attenuation phenomenon means that through workload measurement independent of the hardware platform, the workload of the load task is uniformly quantified at the application layer. By collecting the time-series data of the processor occupancy rate, memory bandwidth consumption, and network throughput metrics during the execution of the load task, which change intermittently with the execution time of the load task, and adopting the time window idea, analyzing the change law of the performance data within a fixed time interval, decomposing the collected hardware metrics into three parts: temporary fluctuation, periodic fluctuation, and long-term attenuation, and quantitatively offsetting the hardware behavior differences between the original performance metrics and the reference architecture, finally generating an evaluation load task adapted to the hardware instruction set of the domestic Xinchuang platform.
[0029] Furthermore, the quantitative scoring using the benchmark comparison method means that after the execution of the hybrid load simulation task on the hardware platform is completed, the hardware platform performance profiling module reads the execution result data of this hybrid load simulation task from the data storage module. A group of standard hardware configurations is pre-selected as the benchmark test environment, and the hardware performance metric values in the same scenario are recorded as the baseline. For the domestic Xinchuang platform to be evaluated, the ratios of its various performance metrics to the baseline values are calculated, and the corresponding scoring functions are applied according to the positive or negative nature of the metrics to calculate the scores.
[0030] The present invention has the following advantages compared with the prior art:
[0031] First, by deeply analyzing the customized instruction extension characteristics of the domestic Xinchuang platform, the present invention designs a matching performance evaluation load generation mechanism, eliminating the performance evaluation deviation caused by the failure of traditional evaluation tools to adapt to the unique instruction extensions and heterogeneous computing resource scheduling strategies when evaluating domestic processors, storage devices, and network components, and improving the scientificity and accuracy of the comprehensive hardware performance evaluation of the domestic Xinchuang platform.
[0032] Second, by introducing the hybrid load simulation technology, the present invention breaks through the limitations of traditional static load testing, and can truly reproduce the dynamic load requirements in high-complexity scenarios such as cloud computing, artificial intelligence training, and big data processing, so that the present invention can comprehensively evaluate the response robustness and resource collaboration efficiency of the hardware under complex multi-task competition and sudden traffic impact.
[0033] Third, through the hardware comprehensive performance profiling method, combined with the hardware performance quantification algorithm and visualization technology, the present invention provides an intuitive and practical performance analysis tool. Traditional test results are often presented in the form of complex data, making it difficult for users to quickly understand the performance advantages and potential bottlenecks of the hardware. The performance profiling radar chart generated by the present invention can intuitively display multi-dimensional metrics such as computing performance, memory usage efficiency, storage bandwidth, and network throughput, helping users quickly identify the performance characteristics of the hardware in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is the working system architecture diagram of the device of the present invention;
[0035] Figure 2 It is the flowchart of the implementation steps of the method of the present invention;
[0036] Figure 3 It is the example diagram of the hardware comprehensive performance evaluation profiling of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] The following combines the drawings and embodiments to make a further detailed description of the present invention:
[0038] Embodiments of the present invention design a performance evaluation load task generation mechanism for the instruction set architectures of processors, storage devices, and network components in domestic information and communication technology (ICT) innovation platforms. Specifically, based on the instruction set architecture documents and hardware behavior data provided by domestic processor manufacturers, a load task generation unit covering Ascend matrix operation instructions, Hygon security coprocessor instructions, and Loongson self-extended instructions is established.
[0039] When generating load tasks, embodiments of the present invention map scalar operation instructions to the vectorized special instruction set of domestic processors, including but not limited to replacing traditional parallel vector instructions with Ascend vmm.4x4f32 matrix multiplication instructions and injecting instruction gap control parameters that conform to the characteristics of domestic chips. For storage devices, by adapting to the scheduling strategy characteristics of domestic storage hardware, a data sharding verification mechanism based on national cryptography algorithms is integrated into the load tasks. In network component testing, embodiments of the present invention construct a test traffic model that conforms to the optimized protocol of the domestic ICT innovation platform by adapting to the jumbo frame acceleration and zero-copy mechanisms of domestic network hardware, and perform traffic coloring by inserting a custom field containing device fingerprint identification information into the data link layer header to achieve the localization feature marking of the protocol header. To eliminate the evaluation deviation caused by architectural differences, in embodiments of the present invention, through workload metrics independent of the hardware platform, the workload of load tasks is uniformly quantified at the application layer, and the hardware behavior differences between the original performance metrics and the benchmark architecture are quantitatively offset, finally generating evaluation load tasks adapted to the hardware instruction set of the domestic ICT innovation platform.
[0040] Embodiments of the present invention achieve hybrid load simulation and systematic efficiency decay elimination through a dynamic load regulation mechanism and timing data analysis technology. In the hybrid load generation strategy, the system will dynamically construct heterogeneous task combinations that adapt to different hardware resource requirements based on the characteristics of cloud computing and big data processing scenarios configured by users. Specifically, for the characteristics of the big data processing scenario where CPU-intensive computing dominates and network communication requirements are low, the system will automatically increase the load distribution weight of general-purpose processor resources through a weighted algorithm, and at the same time proportionally lower the load pressure threshold of communication modules such as network interface devices, so as to achieve the optimal scheduling and load balancing of hardware resources. The load task execution module coordinates the general-purpose processor-intensive computing threads, memory sequential / random read and write operations, and network traffic load, so that the processor occupancy rate, memory bandwidth consumption, and network throughput metrics dynamically match the changes over task time in cloud computing and big data processing scenarios.
[0041] In terms of eliminating performance degradation, in the embodiments of the present invention, by collecting the time-series data of the processor occupancy rate, memory bandwidth consumption, network throughput, etc. during the execution of the load task changing with the time interval of the load task execution, and adopting the idea of a time window, analyzing the variation law of the performance data within a fixed time interval, the collected hardware metrics are decomposed into three parts: temporary fluctuations, periodic fluctuations, and long-term attenuation. Temporary fluctuations represent instantaneous anomalies caused by external environmental disturbances, such as sudden spikes in CPU occupancy rate caused by operating system interrupts or instantaneous jitters in throughput caused by memory bus contention; periodic fluctuations reflect the regular oscillations generated by the inherent behavior patterns of the load, such as the sinusoidal fluctuations in memory occupancy rate triggered by the garbage collection mechanism or the fluctuations in disk usage rate formed by the batch processing task queue; long-term attenuation refers to the continuous degradation trend caused by insufficient hardware heat dissipation conditions or resource leakage during long-term task operation.
[0042] When using the benchmark comparison method for quantitative scoring, only the temporary fluctuation data and periodic fluctuation data reflecting the true capabilities of the hardware are used to eliminate the interference of systematic attenuation factors such as hardware heat dissipation decline and component aging on the evaluation results, and avoid the impact of systematic performance degradation on the performance evaluation of the hardware platform.
[0043] Refer to Figure 1 , the detection device of the present invention is composed of six core modules: a task construction module responsible for generating customized hybrid load tasks; a backend parsing module responsible for parsing task description files; a data storage module responsible for storing intermediate data and final results during task execution; a load execution module deployed on the domestic information and communication technology (ICT) platform to be evaluated for executing hybrid load tasks; a real-time monitoring module responsible for dynamically feedbacking the progress of task execution; a hardware platform performance profiling module that uses the benchmark comparison method for quantitative scoring and generates a radar chart to display multi-dimensional performance profiles. The specific descriptions of the core modules are as follows:
[0044] The described task construction module: is responsible for constructing hybrid load tasks according to the characteristics of different load scenarios. The hybrid load task refers to a customized hybrid load task that applies load pressure to the general-purpose processor, memory device, and network device simultaneously during the evaluation process to match the multi-dimensional hardware resource requirements in diverse actual application scenarios such as cloud computing and big data processing, and provides support for the evaluator's Web page operations and terminal command line operations. The task construction module generates a description file for the hybrid load task for hardware platform evaluation based on the initialization parameters of the evaluation task configured by the evaluator, and sends the description file to the backend parsing module.
[0045] Backend parsing module: Connected to the load execution module and data storage module deployed on the domestic Xinchuang platform to be evaluated. After receiving the hardware platform evaluation hybrid load task description file sent by the task construction module, it parses the hardware platform evaluation hybrid load task description file into specific complex task instructions for the domestic Xinchuang platform to be evaluated and issues them to the load execution module. It receives the hardware platform evaluation hybrid load task execution process information and execution results returned by the load execution module, and stores the hardware platform evaluation hybrid load task execution process information and execution results in the data storage module;
[0046] Data storage module: Stores the execution results of all hardware platform evaluation hybrid load tasks for access and subsequent analysis by the hardware platform performance profiling module;
[0047] Load execution module: Deployed on the domestic Xinchuang platform to be evaluated, with [internal details not provided in the original]. It accepts the hardware platform evaluation hybrid load task sent by the backend parsing module and executes it, and at the same time returns the hardware platform evaluation hybrid load task execution process information and execution results to the backend parsing module;
[0048] Real-time monitoring module: Used to monitor the execution status and intermediate states of the hardware platform evaluation hybrid load tasks. By reading the hardware platform evaluation hybrid load task execution process information stored in the data storage module, it provides real-time feedback on the task execution process to help the evaluator understand the intermediate process of task execution;
[0049] Hardware platform performance profiling module: After the hardware platform evaluation hybrid load task is completed, it reads the execution result data of the hybrid load task from the data storage module. By analyzing the data of the domestic Xinchuang platform to be evaluated during task operation, considering the changing trend of hardware performance over time, eliminating the index deviation caused by systematic efficiency decay phenomena, it uses the benchmark comparison method for quantitative scoring and presents the hardware performance profile in the form of a radar chart, intuitively showing the key dimensions of computing performance, memory usage efficiency, storage bandwidth, and network throughput.
[0050] Refer to Figure 2 , and further describe the steps of the evaluation method of the present invention.
[0051] Step 1, the task construction module executes customized hybrid load tasks according to the characteristics of different load scenarios;
[0052] In the embodiment of the present invention, the hybrid load task refers to a customized load task that simultaneously applies load pressure to the general-purpose processor, memory device, and network device during the evaluation process to match the multi-dimensional hardware resource requirements in the actual application scenarios of cloud computing and big data processing. By running multiple load tasks, the performance of different hardware components of the domestic Xinchuang platform to be evaluated can be comprehensively evaluated.
[0053] Computing performance test: Execute compute-intensive tasks, such as mathematical operations or data processing, to test the computing power and efficiency of the processor. Pay attention to the performance of the processor under high load, for example, the task completion speed and resource utilization.
[0054] Memory performance test: Run memory-related tasks to measure the bandwidth and latency of the memory. Test the read and write speeds of the memory and its stability in a multi-tasking scenario.
[0055] Disk performance test: Perform various read and write operations, including sequential and random reads and writes, to evaluate the input / output capabilities of the storage device. Check the response speed and data transfer efficiency of the disk under different load conditions.
[0056] Network performance test: Simulate network traffic, such as high-concurrency connections or large-volume transmissions, to evaluate the bandwidth and latency of the network device. Test the stability and transmission quality of the network in different scenarios.
[0057] Step 2, the task construction module initializes the parameters according to the evaluation tasks configured by the evaluator.
[0058] The initialized parameters of the evaluation task include but are not limited to which systems under test need to be tested in this test (such as domestic Shenwei hardware platform, domestic Feiteng hardware platform), which test items need to be executed (such as CPU computing resource load test item, memory stress load test item, network card hardware stress load test item, storage device IO stress load test item, etc.), and the specific execution parameters of each test item (such as the computing task mode and maximum pressure of the CPU computing resource load test item, the total amount of memory access read and write and the proportion of sequential and random reads and writes of the memory stress load test item, etc.). After the initialization parameters of the evaluation task are configured, the task construction module generates a description file for the hybrid load task of the hardware platform evaluation based on the configured initialization parameters of the evaluation task and sends the description file to the backend execution module.
[0059] Step 3, the backend parsing module parses the task description file into a specific hybrid load task for the domestic IT application platform to be evaluated and issues it to the load execution module; at the same time, it receives the execution process information and execution results of the hybrid load task of the hardware platform evaluation returned by the load execution module and stores them in the data storage module.
[0060] The parsing process first extracts core information such as load configuration, test conditions, test metrics, and execution order from the task description file. Based on the extracted information, the backend parsing module dynamically generates a set of targeted hybrid workload tasks in combination with the hardware characteristics and resource configuration of the platform to be evaluated. For high-concurrency computing tasks on general-purpose processors, appropriate multi-threaded operation instructions will be generated according to the computing power of the platform; for large-scale data read and write operations on memory devices, a data access pattern that conforms to the memory bandwidth and latency characteristics will be constructed; for strong requirements on network devices, high-traffic data packets will be generated to simulate stress tests in a real network environment.
[0061] Step 4: The load execution module deployed on the domestic IT infrastructure platform to be evaluated executes the hybrid workload task instructions for hardware platform evaluation and returns the information and results of the hybrid workload task execution process to the backend parsing module.
[0062] The load execution module runs a series of domestically adapted test components, which can comprehensively evaluate various performance metrics of the domestic hardware platform. The specific performance metrics are as follows:
[0063] Computing performance metrics: including single-core and multi-core processing capabilities, integer and floating-point operation capabilities, processing performance under floating loads, etc. The computing performance test executes compute-intensive tasks such as mathematical operations, data processing, and scientific computing to test the computing power and efficiency of the processor under different loads. Focus on the performance of the processor under high loads, evaluate the task completion speed, resource utilization of the processor (such as CPU occupancy rate, context switching frequency, etc.), and the stability and response time of the system when executing complex tasks.
[0064] Memory performance metrics: including memory bandwidth, memory latency, sequential and random access performance of memory, memory usage efficiency, etc. The memory performance test runs memory-related tasks such as data-intensive operations and memory access in multi-task scenarios to measure the bandwidth and latency of memory. Test the read and write speed of memory, sequential and random access performance, and the stability and reliability of memory under high-concurrency loads. In addition, the memory usage efficiency will also be evaluated, such as memory fragmentation, cache hit rate, etc.
[0065] Disk performance metrics: including disk read and write speed, IOPS, latency, sequential and random read and write performance, storage bandwidth, etc. The disk performance test performs various read and write operations, including sequential and random reads and writes, to evaluate the input and output capabilities of storage devices. Test the response speed of the disk under different load conditions, data transfer efficiency, and performance under high loads, such as disk queue length, response latency, etc. The performance bottlenecks and stability of the disk in scenarios such as large data writing and data backup will also be analyzed.
[0066] Network performance metrics: including network bandwidth, latency, throughput, packet loss rate, connection stability, etc. Network performance tests simulate different network traffic scenarios, such as high-concurrency connections, large-volume data transfers, or network performance under long-term high loads, and evaluate key metrics of network devices such as bandwidth, latency, throughput, and packet loss rate. Test the stability and transmission quality of the network in different scenarios, including the performance of issues such as packet loss, latency fluctuations, and connection timeouts during data transmission. Also test the robustness of network devices during high-load and long-term operation to ensure their stability and reliability.
[0067] Step 5: The real-time monitoring module reads the information on the execution process of the hybrid load task stored in the data storage module and provides real-time feedback on the task execution process to help the evaluator understand the intermediate process of task execution.
[0068] The information on the running process of the load task covers a variety of key time-series data, including but not limited to CPU utilization rate, memory usage rate, network traffic volume, disk throughput, and hardware temperature. These data are presented in a time-series format, which can clearly show the changing trends of hardware resources during task execution. To achieve efficient data visualization, the real-time monitoring module adopts Prometheus, an open-source monitoring solution. Prometheus can not only collect and store these time-series data in real time but also support flexible query and display functions, allowing evaluators to observe details such as the temperature fluctuations of the processor under high loads, the read / write pressure changes of memory devices, and the traffic fluctuations of network devices through intuitive charts.
[0069] Step 6: The hardware platform performance profiling module reads the information on the task execution process and the execution result data stored in the data storage module. By analyzing the data of the domestic IT infrastructure platform to be evaluated during task operation, it analyzes the changing patterns of performance data within a time window, decomposes metrics such as processor occupancy rate and memory throughput into three parts: temporary fluctuations, periodic fluctuations, and long-term decay, fully considering the changing trends of hardware performance over time, eliminating the metric offsets caused by systematic performance decay phenomena, and uses the benchmark comparison method for quantitative scoring. It uses a radar chart to display the hardware performance profile, visually presenting key dimensions such as computing performance, memory usage efficiency, storage bandwidth, and network throughput.
[0070] The specific process of quantitative scoring using the benchmark comparison method is as follows:
[0071] Based on the test results of calculating the computing load, memory load, disk load, network load, etc. described in Step 4, taking the test results of a certain test environment as a benchmark, using a scoring function to calculate the result gap between the current test environment and the benchmark test environment, obtaining a result score, and finally accumulating and averaging the scores as the scores of the computing performance, memory performance, disk performance, and network performance of the current test environment, as a hardware performance evaluation scheme.
[0072] For positive indicators where the larger the observed value, the better the performance. The score calculation formula is:
[0073]
[0074] Among them, z j represents the specific quantitative score of the j-th test index item, and x j represents the j-th index value, and x (std)j represents the j-th index baseline value.
[0075] For reverse indicators where the smaller the observed value, the better the performance. The score calculation formula is:
[0076]
[0077] Among them, in the embodiments of the present invention, the score corresponding to the baseline is 100.
[0078] After obtaining the specific quantitative scores of each index according to the above method, the hardware platform performance profiling module draws a radar chart for the comprehensive performance evaluation of domestic Xinchuang platform hardware as shown in Figure 3 This radar chart uses a multi-dimensional polar coordinate system to intuitively present the comprehensive performance comparison of the two major domestic Xinchuang hardware platforms of Shenwei and Feiteng.
[0079] Figure 3 The graphic layout in Figure 3 is composed of 25 equally spaced radial coordinate axes, covering five major technical dimensions of computing performance, memory efficiency, storage capacity, network characteristics, and the effectiveness of artificial intelligence acceleration cards. Each axis corresponds to a specific performance index, forming a closed-loop evaluation system. The performance data of each platform is mapped to the corresponding coordinate axis after normalization processing, and adjacent nodes are connected by straight lines to form a characteristic contour. In Figure 3 , the Shenwei platform is covered with a blue polygon, and the Feiteng platform is covered with an orange polygon, intuitively presenting the hardware performance of the domestic Xinchuang platform to be evaluated in key dimensions such as computing performance, memory usage efficiency, storage bandwidth, and network throughput.
Claims
1. An evaluation device for the comprehensive performance of domestic information and communication technology innovation platform hardware, characterized in that, The evaluation device includes a task construction module, a backend parsing module, a data storage module, a load execution module, a real-time monitoring module, and a hardware platform performance profiling module; among which: The task construction module executes customized hybrid load tasks according to the characteristics of different load scenarios; according to the initialization parameters of the evaluation tasks configured by the evaluator, it supports Web page operations and terminal command line operations, generates a description file for the hardware platform to evaluate hybrid load tasks, and sends the description file to the backend parsing module; The backend parsing module, after receiving the description file of the hardware platform to evaluate hybrid load tasks sent by the task construction module, parses the description file of the hardware platform to evaluate hybrid load tasks into specific complex task instructions for the domestic Xinchuang platform to be evaluated and issues them to the load execution module; at the same time, the backend parsing module also receives the execution process information and execution results of the hardware platform to evaluate hybrid load tasks returned by the load execution module, and stores the execution process information and execution results of the hardware platform to evaluate hybrid load tasks in the data storage module; The data storage module stores the execution results of all hardware platform to evaluate hybrid load tasks for the hardware platform performance profiling module to access and subsequent analysis; The load execution module is deployed on the domestic Xinchuang platform to be evaluated, receives the hardware platform to evaluate hybrid load tasks sent by the backend parsing module and executes them, and at the same time returns the execution process information and execution results of the hardware platform to evaluate hybrid load tasks to the backend parsing module; The real-time monitoring module is used to monitor the execution situation and intermediate status of the hardware platform to evaluate hybrid load tasks. By reading the execution process information of the hardware platform to evaluate hybrid load tasks stored in the data storage module, it real-time feedbacks the task execution process situation to help the evaluator understand the intermediate process of task execution; The hardware platform performance profiling module is used to, after the execution of the hardware platform to evaluate hybrid load tasks is completed, read the execution result data of the hybrid load tasks from the data storage module, analyze the data of the domestic Xinchuang platform to be evaluated during task operation, consider the change trend of hardware performance over time, exclude the index deviation caused by the systematic efficiency decay phenomenon, use the benchmark comparison method for quantitative scoring, and use a radar chart to display the hardware performance profile, visually presenting the key dimensions of computing performance, memory usage efficiency, storage bandwidth, and network throughput.
2. The comprehensive performance evaluation method for domestic Xinchuang platform hardware of the evaluation device according to claim 1, characterized in that, The specific steps of this evaluation method are as follows: Step 1, the task construction module executes customized hybrid load tasks according to the characteristics of different load scenarios; Step 2, the task construction module generates a description file for the hardware platform to evaluate hybrid load tasks according to the initialization parameters of the evaluation tasks configured by the evaluator, and sends the description file to the backend parsing module; Step 3, the backend parsing module parses the task description file into specific hardware platform to evaluate hybrid load tasks for the domestic Xinchuang platform to be evaluated, and issues them to the load execution module; at the same time, it receives the execution process information and execution results of the hardware platform to evaluate hybrid load tasks returned by the load execution module, and stores them in the data storage module; Step 4: The load execution module deployed on the domestic IT infrastructure platform to be evaluated executes the hybrid load task instruction for hardware platform evaluation and returns the information and results of the hybrid load task execution process to the backend parsing module. Step 5: The real-time monitoring module reads the information on the execution process of the hybrid load task for hardware platform evaluation stored in the data storage module and provides real-time feedback on the task execution process to help the evaluator understand the intermediate process of task execution. Step 6: The hardware platform performance profiling module reads the information on the task execution process and the result data stored in the data storage module, analyzes the data of the domestic IT infrastructure platform to be evaluated during task operation, considers the changing trend of hardware performance over time, eliminates the index deviation caused by systematic efficiency decay, performs quantitative scoring using the benchmark comparison method, and uses a radar chart to display the hardware performance profile, visually presenting the key dimensions of computing performance, memory usage efficiency, storage bandwidth, and network throughput.
3. The evaluation method according to claim 2, characterized in that The hybrid load task described in Step 1 refers to mapping scalar operation instructions to the vectorized special instruction set of domestic processors, replacing traditional parallel vector instructions with the Ascend vmm.4x4f32 matrix multiplication instruction, and injecting instruction gap control parameters that conform to the characteristics of domestic chips; for storage devices, integrating a data sharding verification mechanism based on the national cryptography algorithm into the load task by adapting to the scheduling strategy characteristics of domestic storage hardware; in the network component test, constructing a test traffic model that conforms to the optimized protocol of the domestic IT infrastructure platform by adapting to the jumbo frame acceleration and zero-copy mechanisms of domestic network hardware, and performing traffic coloring by inserting a custom field containing device fingerprint identification information into the data link layer header to achieve the localization feature marking of the protocol header.
4. The evaluation method according to claim 2, wherein The evaluation task initialization parameters described in Step 2 include the hardware platform to be tested currently, the test items to be executed, and the specific execution parameters of each test item; the test items to be executed include the CPU computing resource load test item, the memory stress load test item, the network card hardware stress load test item, and the storage device IO stress load test item. The specific execution parameters of each test item include: the computing task mode and maximum stress in the CPU computing resource load test item. The total amount of memory access read and write and the ratio of sequential to random read and write in the memory stress load test item.
5. The evaluation method according to claim 2, characterized in that The execution of the hybrid load task instruction for hardware platform evaluation described in Step 4 means that for general-purpose processors, the load execution module runs high-concurrency computing tasks and applies high load pressure to evaluate the performance of general-purpose processors in executing tasks with strong computing requirements. For memory devices, the load execution module performs large-scale data read and write operations to comprehensively test the key indicators of the sequential and random read and write speeds of memory devices, thereby evaluating the performance of memory devices in tasks with strong memory access requirements; for network devices, the load execution module generates a high-traffic packet network load to examine the transmission efficiency and reliability of the network load, thereby evaluating its performance in tasks with strong communication requirements.
6. The evaluation method according to claim 5, wherein The operation of high-concurrency computing tasks means that by creating multiple thread or process execution units, intensive mathematical operations are performed simultaneously to fully utilize the computing resources of the processor.
7. The evaluation method according to claim 5, wherein The key indicators for testing the sequential read / write speed and random read / write speed of the memory device refer to: recording the read / write speed and latency key indicators of the memory device during the sequential read / write and random read / write operations of the memory device, that is, by continuously reading and writing large data blocks, measuring the sustained throughput and latency of the memory device; the random read / write test is to randomly select positions in the memory address space to perform read / write operations on small data blocks, and evaluate the response speed and efficiency of the memory device in the non-continuous access mode.
8. The evaluation method according to claim 5, wherein The transmission efficiency and reliability of the network load refer to that the load execution module constructs a network load scenario by generating high-traffic data packets, monitors the key parameters of the throughput, error rate, and connection stability of the network device during the test, and evaluates the reliability and performance of the network device in communication-intensive demand tasks.
9. The evaluation method according to claim 2, wherein The elimination of the index offset caused by the systematic performance attenuation phenomenon in step 6 means that through workload measurement independent of the hardware platform, the workload of the load task is uniformly quantified at the application layer. By collecting the time-series data of the processor occupancy rate, memory bandwidth consumption, and network throughput index during the execution of the load task, which vary with the execution time of the load task, and adopting the idea of a time window, the variation law of the performance data within a fixed time interval is analyzed. The collected hardware indicators are decomposed into three parts: temporary fluctuation, periodic fluctuation, and long-term attenuation. The hardware behavior differences between the original performance indicators and the reference architecture are quantitatively offset, and finally an evaluation load task adapted to the hardware instruction set of the domestic IT infrastructure innovation platform is generated.
10. The evaluation method according to claim 2, wherein The quantitative scoring using the reference comparison method in step 6 means that after the execution of the mixed load task on the hardware platform is completed, the hardware platform performance profiling module reads the execution result data of this mixed load task from the data storage module. A group of standard hardware configurations are pre-selected as the reference test environment, and the hardware performance indicator values in the same scenario are recorded as the baseline. For the domestic IT infrastructure innovation platform to be evaluated, the ratio of its various performance indicators to the baseline value is calculated, and the corresponding scoring function is applied according to the positive or reverse nature of the indicators to calculate the score.
Citation Information
Patent Citations
Performance test method and system for artificial intelligence accelerator card of domestic heterogeneous platform
CN117370088A
Cited By
Network security operation and maintenance management system and method
CN120811762A
Computing center performance evaluation method and electronic equipment
CN121434038A
A method for evaluating performance of a computing center and an electronic device
CN121434038B