Method and device for evaluating computing power performance based on vector calculation scene
Patent Information
- Application Number
- CN202611208369.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-11
- Publication Date
- 2026-09-08
AI Technical Summary
[0003]上述基准测试集SPEC CPU 2017作为综合性测试套件,其设计初衷是评估处理器整体性能,而非专项测试向量计算能力,因其控制流图结构复杂,容易导致编译器自动向量化成功率低下的问题;且其完整程序测试项中核心计算模块的向量化加速效果会被任务调度、I/O等待等系统级开销严重稀释,即便局部函数获得显著SIMD加速,在宏观测试结果中也难以体现
该算力性能评估方法通过获取用户配置参数,基于屋顶线模型获取待测处理器在用户配置参数下的性能标定参数,生成屋顶线图。并基于用户配置参数获取多个标准化测试负载对应的性能参数与微架构指标参数,得到多维测试数据集;其中,多个标准化测试负载至少包括缓存阶梯扫描负载、DRAM访存带宽压力负载、算术强度扫描负载以及专项测试负载;专项测试负载用于表征跨Die行为和/或超线程对应的测试负载。随后将多维测试数据集下各测试点映射至屋顶线图,得到多个映射测试点,并按照预设划分规则将多个映射测试点划分至对应性能瓶颈类别的测试集合;每个测试集合对应一个瓶颈判定规则;每个测试集合基于自身对应的瓶颈判定规则进行瓶颈分析,得到瓶颈分析报告。本申请基于超线程、系统架构、CPU缓存、DDR 带宽、主频以及跨Die行为等向量性能影响因素构建多个标准化测试负载,得到待测处理器对应的多维测试数据集,并基于多维测试数据集自动化进行瓶颈性能分析,从而在评估CPU向量计算算力的基础上,自动化实现性能瓶颈分析。
Smart Images

Figure CN122711501A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computing power evaluation technology, and more specifically, to a method and apparatus for evaluating computing power performance based on vector computing scenarios. Background Technology
[0002] With the widespread adoption of Single Instruction Multiple Data (SIMD) architecture in modern processors, vector computation performance has become a core metric in key application areas such as scientific computing, audio and video encoding / decoding, machine learning, and vector databases. Current technologies primarily rely on the general-purpose benchmark set SPEC CPU 2017 and the dedicated matrix operation GEMM as alternatives to evaluate vector computation capabilities.
[0003] The SPEC CPU 2017 benchmark suite, as a comprehensive test suite, was designed to evaluate overall processor performance rather than specifically test vector computation capabilities. Its complex control flow graph structure makes it prone to low success rates in automatic vectorization by the compiler. Furthermore, the vectorization acceleration effect of core computation modules in its complete program test items is severely diluted by system-level overhead such as task scheduling and I / O waits; even if local functions achieve significant SIMD acceleration, it is difficult to reflect in the macroscopic test results. While dedicated matrix operations (GEMM), as a core linear algebra operation, have extremely high computational intensity and can effectively eliminate memory bandwidth bottlenecks, this idealized computational model cannot reflect the actual throughput capabilities in real-world applications, including memory-intensive and sparse operations, as well as complex vector instruction scenarios such as permutations, aggregations, and scattering. Additionally, its extreme performance relies on manual assembly optimizations for specific microarchitectures, which can easily obscure the compiler's automatic vectorization capabilities and the hardware's ease of use under general load conditions. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a computing performance evaluation method and apparatus based on vector computing scenarios, which can automatically perform performance bottleneck analysis based on the evaluation of CPU vector computing power, and effectively reveal performance influencing factors such as processor core computing power, memory bandwidth, cross-die, and hyper-threading.
[0005] To achieve the above objectives, the technical solutions adopted in the embodiments of the present invention are as follows: In a first aspect, the present invention provides a method for evaluating computing power performance based on a vector computing scenario, the method comprising: Obtain user configuration parameters, and based on the roofline model, obtain the performance calibration parameters of the processor under test under the user configuration parameters to generate a roofline diagram; the performance calibration parameters are used to characterize the computing power and memory bandwidth of the processor under test. Based on user-configured parameters, performance parameters and micro-architecture metrics corresponding to multiple standardized test loads are obtained to create a multi-dimensional test dataset. The multiple standardized test loads include at least cache ladder scan load, DRAM memory access bandwidth pressure load, arithmetic strength scan load, and specialized test loads. The specialized test loads are used to characterize cross-die behavior and / or test loads corresponding to hyper-threading. The test points in the multidimensional test dataset are mapped to the roofline diagram to obtain multiple mapped test points. These multiple mapped test points are then divided into test sets corresponding to the performance bottleneck categories according to preset partitioning rules. Each test set corresponds to a bottleneck determination rule. Each test set performs bottleneck analysis based on its own corresponding bottleneck determination rules, and obtains a bottleneck analysis report.
[0006] Optionally, the performance calibration parameters include: when calculating the actual roof value, the theoretical calculated peak value, the measured DRAM bandwidth value, and the cache capacity boundaries corresponding to each level of cache, the steps for obtaining the performance calibration parameters of the processor under test under user-configured parameters based on the roofline model and generating the roofline diagram include: The physical topology binding information of the processor under test is obtained based on the user configuration parameters, and the thread distribution characteristics of the processor under test under the user configuration parameters are obtained. Affinity constraints are constructed based on physical topology binding information; Based on preset rules, the actual calculated roof value, theoretical calculated peak value, DRAM measured bandwidth value, and cache capacity boundary corresponding to each level of cache are obtained for the processor under test under affinity constraints. A roofline diagram is constructed based on actual calculated roof values, theoretically calculated peak values, and measured DRAM bandwidth values. Optionally, the steps for generating a roofline map include constructing a roofline coordinate system based on actual calculated roof values, theoretically calculated peak values, and measured DRAM bandwidth values: Construct a coordinate system with arithmetic strength as the x-axis and measured floating-point performance as the y-axis; Two horizontal lines are drawn in the coordinate system based on the actual calculated roof value and the theoretical calculated peak value, respectively serving as the actual calculated roof line and the theoretical calculated roof line. Draw a diagonal line in the coordinate system with the measured bandwidth of DRAM as the slope, and use it as the bandwidth roofline.
[0007] Optionally, the steps to obtain a multidimensional test dataset by acquiring performance parameters and microarchitecture metrics corresponding to multiple standardized test loads based on user configuration parameters include: Multiple standardized test loads are constructed based on user configuration parameters and performance calibration parameters; Obtain the performance parameters and microarchitecture metrics of each standardized test load in any runtime cycle. The performance parameters include effective frequency, package power consumption, measured floating-point performance, SIMD instruction width distribution, and the miss rate corresponding to each level of cache data. The microarchitecture metrics include TLB miss rate, front-end stall rate, and branch prediction failure rate. Based on the timestamp and the affinity label corresponding to each standardized test load, the performance parameters and microarchitecture metrics of each standardized test load are encapsulated into a multidimensional test dataset.
[0008] Optionally, when the roofline diagram includes the actual calculated roofline, the theoretical calculated roofline, the bandwidth roofline, and the cache capacity boundaries corresponding to each level of cache, and the performance bottleneck categories include computational bottlenecks, cache capacity overflow bottlenecks, DRAM bandwidth bottlenecks, and microarchitecture bottlenecks, the steps of dividing multiple mapped test points into test sets corresponding to the performance bottleneck categories according to preset partitioning rules include: Select the first target test point from multiple mapping test points to construct a computational bottleneck test set; the first target test point is used to characterize that the coordinates of the current mapping test point are within the preset boundary threshold range of the actual calculated roof line; A second target test point is selected from multiple mapped test points to construct a DRAM bandwidth bottleneck test set; the second target test point is used to characterize that the coordinates of the current mapped test point are within the preset boundary threshold range of the bandwidth roofline; A third target test point is selected from multiple mapping test points to construct a cache capacity overflow bottleneck test set. The third target test point is used to characterize the deviation condition that the current mapping test point's own coordinates are lower than the actual calculated roof line and it has not reached the boundary constraint of the bandwidth roof line. At the same time, its own working set size exceeds the target-level cache capacity boundary and its miss rate at the target level exceeds the first preset failure rate threshold. A fourth target test point is selected from multiple mapping test points to construct a microarchitecture bottleneck test set. The fourth target test point is used to characterize that the coordinates of the current mapping test point are far away from the actual computational roofline and bandwidth roofline, and the miss rate of its own cache at all levels is lower than the second preset miss rate threshold.
[0009] Optionally, when the test set includes a bottleneck test set, the steps for each test set to perform bottleneck analysis based on its corresponding bottleneck determination rules and obtain a bottleneck analysis report include: For each first target test point under the computing bottleneck test set, determine whether the effective frequency of the first target test point itself is lower than the nominal turbo frequency; If so, then the current first target test point is determined to be the frequency loss bottleneck; If the frequency is not lower than the nominal turbo frequency, then the SIMD instruction width distribution of the current first target test point is used to determine whether the SIMD width utilization exceeds the preset vector width rule. If so, the current first target test point is determined to be underutilized in terms of vector width. If the preset vector width rule is met, then determine whether the number of instructions executed by the current first target test point in each cycle meets the preset IPC rule. If yes, then determine that the current first target test point is a pure computational bottleneck; otherwise, determine that the current first target test point is insufficient in instruction-level parallelism.
[0010] Optionally, when the performance calibration parameters include the cache capacity boundaries corresponding to each level of cache, and the test set includes a cache capacity overflow bottleneck test set, the steps for each test set to perform bottleneck analysis based on its own corresponding bottleneck determination rules and obtain a bottleneck analysis report include: For each third target test point under the cache capacity overflow bottleneck test set, if it is determined that the working set of the third target test point itself exceeds the L1 cache capacity boundary and the L1 cache miss rate exceeds the preset burst judgment rule, then the current third target test point is determined to meet the L1 cache capacity overflow. If it is determined that the working set of the third target test point exceeds the L2 cache capacity boundary, and the L2 cache miss rate exceeds the preset burst judgment rule, then the current third target test point is determined to meet the L2 cache capacity overflow condition. If it is determined that the working set of the third target test point exceeds the capacity boundary of the L3 cache, and the miss rate of the L3 cache exceeds the preset burst judgment rule, then the current third target test point is determined to meet the L3 cache capacity overflow.
[0011] Optionally, when the performance calibration parameters include the measured bandwidth value of DRAM, and the test set includes a DRAM bandwidth bottleneck test set, the steps for each test set to perform bottleneck analysis based on its corresponding bottleneck determination rules and obtain a bottleneck analysis report include: For each second target test point under the DRAM bandwidth bottleneck test set, determine whether the DRAM bandwidth utilization of the second target test point meets the preset range of the measured DRAM bandwidth value. If it does, determine that the current second target test point belongs to the DRAM bandwidth bottleneck; otherwise, determine that the current second target test point does not belong to the DRAM bandwidth bottleneck.
[0012] Optionally, when the microarchitecture metrics include TLB miss rate, frontend stall rate, and branch prediction failure rate, and the test set includes a microarchitecture bottleneck test set, the steps for each test set to perform bottleneck analysis based on its corresponding bottleneck determination rules and obtain a bottleneck analysis report include: For each fourth target test point under the microarchitecture bottleneck test set, determine whether the TLB miss rate of the fourth target test point exceeds the preset miss threshold. If so, determine that the current fourth target test point belongs to the TLB-type microarchitecture bottleneck. Alternatively, determine whether the front-end stall rate of the fourth target test point exceeds the preset stall rate. If so, determine that the current fourth target test point belongs to the front-end stall type microarchitectural bottleneck. Alternatively, determine whether the branch prediction failure rate of the fourth target test point exceeds the preset failure rate. If so, determine that the current fourth target test point belongs to the branch prediction failure type of microarchitecture bottleneck.
[0013] In a second aspect, the present invention also provides a computing power performance evaluation device, applied to the computing power performance evaluation method of any one of the first aspects above, the computing power performance evaluation device comprising: The performance calibration module is used to obtain user configuration parameters, acquire performance calibration parameters of the processor under test under user configuration parameters based on the roofline model, and generate a roofline diagram; the performance calibration parameters are used to characterize the computing power and memory bandwidth of the processor under test. The data acquisition module is used to obtain performance parameters and micro-architecture metrics corresponding to multiple standardized test loads based on user-configured parameters, and to obtain a multi-dimensional test dataset. The multiple standardized test loads include at least cache ladder scan load, DRAM memory access bandwidth pressure load, arithmetic intensity scan load, and special test loads. The special test loads are used to characterize cross-die behavior and / or test loads corresponding to hyper-threading. The bottleneck segmentation module is used to map each test point in the multidimensional test dataset to the roofline diagram, resulting in multiple mapped test points. These mapped test points are then divided into test sets corresponding to the performance bottleneck categories according to preset segmentation rules. Each test set corresponds to a bottleneck determination rule. The bottleneck analysis output module is used to perform bottleneck analysis on each test set based on its own corresponding bottleneck determination rules and generate a bottleneck analysis report.
[0014] The computing performance evaluation method and apparatus based on vector computing scenarios provided in this invention have the following beneficial effects: This computing performance evaluation method obtains user configuration parameters and, based on a roofline model, acquires the performance calibration parameters of the processor under test under these parameters, generating a roofline diagram. It then obtains performance parameters and microarchitectural metrics corresponding to multiple standardized test loads based on the user configuration parameters, resulting in a multidimensional test dataset. These standardized test loads include at least cache ladder scan load, DRAM memory access bandwidth pressure load, arithmetic intensity scan load, and specialized test loads. The specialized test loads characterize cross-die behavior and / or the test loads corresponding to hyper-threading. Subsequently, each test point in the multidimensional test dataset is mapped to the roofline diagram, resulting in multiple mapped test points. These mapped test points are then divided into test sets corresponding to performance bottleneck categories according to preset partitioning rules. Each test set corresponds to a bottleneck determination rule. Each test set performs bottleneck analysis based on its own bottleneck determination rule, generating a bottleneck analysis report. This application constructs multiple standardized test loads based on vector performance influencing factors such as hyper-threading, system architecture, CPU cache, DDR bandwidth, clock speed, and cross-die behavior to obtain a multi-dimensional test dataset corresponding to the processor under test. Based on the multi-dimensional test dataset, bottleneck performance analysis is automatically performed, thereby automatically realizing performance bottleneck analysis based on the evaluation of CPU vector computing power.
[0015] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 A flowchart illustrating the steps of the computing power performance evaluation method provided in an embodiment of the present invention is shown. Figure 2 A flowchart of step 100 provided in an embodiment of the present invention is shown; Figure 3 A flowchart of step 103 provided in an embodiment of the present invention is shown; Figure 4 A flowchart of step 200 provided in an embodiment of the present invention is shown; Figure 5 A flowchart of step 300 provided in an embodiment of the present invention is shown; Figure 6 One of the step flowcharts for step 400 provided in an embodiment of the present invention is shown; Figure 7 This illustrates a second flowchart of step 400 provided in an embodiment of the present invention; Figure 8 This illustrates a third step flowchart of step 400 provided in an embodiment of the present invention; Figure 9 The fourth step flowchart of step 400 provided in the embodiment of the present invention is shown; Figure 10 A block diagram of the computing power performance evaluation device provided in an embodiment of the present invention is shown; Figure 11 A block diagram of the computing power performance evaluation system provided in an embodiment of the present invention is shown.
[0018] Icons: 10-Computing power performance evaluation device; 11-Performance calibration module; 12-Data acquisition module; 13-Bottleneck identification module; 14-Bottleneck analysis output module; 20-Computing power performance evaluation system; 21-User configuration layer; 22-Scheduling and execution engine layer; 23-Platform abstraction layer; 24-System monitoring layer; 25-Output layer. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0020] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0021] It should be noted that relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0022] As described in the background section, existing technologies lack dedicated benchmark tests for evaluating vector computing power. In practice, alternative methods such as SPEC CPU 2017 or GEMM matrix operations are often used. However, since vector computing performance is affected by many factors, these alternative methods cannot accurately obtain CPU vector computing power and cannot automatically locate specific performance bottlenecks based on the current hardware during the evaluation process.
[0023] Based on this, this application provides a computing performance evaluation scheme for vector computing scenarios, which takes into account many interference factors, generates standardized test loads in a customized manner, and automatically evaluates the potential performance bottlenecks based on multi-dimensional test datasets during the evaluation process, thus making up for the lack of a comprehensive evaluation scheme for vector computing scenarios in the existing technology.
[0024] Please refer to Figure 1 , Figure 1 A flowchart of the computing power performance evaluation method provided in this embodiment of the invention is shown; the computing power performance evaluation method in this embodiment includes steps 100 to 400.
[0025] Step 100: Obtain user configuration parameters, obtain the performance calibration parameters of the processor under test under the user configuration parameters based on the roofline model, and generate the roofline diagram.
[0026] In this embodiment, the performance calibration parameters are used to characterize the computing power and memory bandwidth of the processor under test.
[0027] Step 200: Based on the user configuration parameters, obtain the performance parameters and micro-architecture metrics corresponding to multiple standardized test loads to obtain a multi-dimensional test dataset.
[0028] In this embodiment, the standardized test loads include at least cache ladder scan load, DRAM memory access bandwidth pressure load, arithmetic strength scan load, and special test loads; wherein, the special test loads are used to characterize cross-die behavior and / or test loads corresponding to hyper-threading.
[0029] Based on this, this application can incorporate factors such as differences in SIMD instruction sets of different architectures, hyper-threading factors, CPU cache factors, DDR bandwidth factors, clock speed, and cross-die behavior factors into the evaluation, thereby determining the corresponding test items based on the above-mentioned influencing factors and improving the performance evaluation.
[0030] Step 300: Map each test point in the multidimensional test dataset to the roofline diagram to obtain multiple mapped test points, and divide the multiple mapped test points into test sets corresponding to the performance bottleneck categories according to the preset division rules.
[0031] It should be noted that in this embodiment, each test set corresponds to a bottleneck determination rule. Step 400: Each test set performs bottleneck analysis based on its own corresponding bottleneck determination rules to obtain a bottleneck analysis report.
[0032] In this embodiment, the performance calibration parameters include at least the actual calculated peak value, the theoretical calculated peak value, the measured DRAM bandwidth value, and the cache capacity boundaries corresponding to each level of cache. Microarchitectural metrics parameters include at least the TLB miss rate, front-end stall rate, and branch prediction failure rate. Performance parameters include at least the effective frequency, package power consumption, measured floating-point performance, SIMD instruction width distribution, and the miss rate corresponding to each level of cache data.
[0033] In summary, this application constructs multiple standardized test loads by considering vector performance influencing factors such as hyper-threading, system architecture, CPU cache, DDR bandwidth, clock speed, and cross-die behavior, to obtain a multi-dimensional test dataset corresponding to the processor under test. This improves upon the shortcomings of insufficient analytical granularity in traditional evaluation methods and automates bottleneck performance analysis based on the multi-dimensional test dataset. Thus, based on the evaluation of CPU vector computing power, it automatically achieves performance bottleneck analysis, effectively revealing performance influencing factors such as processor core computing power, memory bandwidth, cross-die behavior, and hyper-threading.
[0034] In this embodiment, the multidimensional test dataset includes the configuration parameters, performance parameters, and microarchitecture metrics of each test point. The multidimensional test dataset includes, but is not limited to, effective frequency, package power consumption, measured floating-point performance, SIMD instruction width distribution, cache miss rate and TLB miss rate, front-end stall rate, branch prediction failure rate, and key performance counter event values obtained from performance counters.
[0035] To address the issue that existing technologies cannot perform performance evaluations based on real hardware conditions, please refer to... Figure 2 , Figure 2 The flowchart of step 100 provided in this embodiment of the invention is shown; in this embodiment, step 100, which obtains the performance calibration parameters of the processor under test under user configuration parameters based on the roof line model and generates the roof line diagram, includes steps 101 to 104.
[0036] Step 101: Obtain the physical topology binding information of the processor under test based on the user configuration parameters, and obtain the thread distribution characteristics of the processor under test under the user configuration parameters.
[0037] Step 102: Construct affinity constraints based on physical topology binding information.
[0038] Step 103: Based on preset rules, obtain the actual calculated roof value, theoretical calculated peak value, measured DRAM bandwidth value, and cache capacity boundaries corresponding to each level of cache for the processor under the affinity constraint.
[0039] Step 104: Construct a roofline diagram based on the actual calculated roof value, the theoretically calculated peak value, and the measured DRAM bandwidth value.
[0040] In this embodiment, user configuration parameters include at least matrix size, data type, hyper-threading switch, target instruction set version, iteration count, CPU core binding strategy for the test, memory allocation node, and iteration count, to construct a physical boundary that constrains subsequent tests. Data types include FP32, FP64, etc., and target instruction set versions include SSE, AVX2, AVX-512, NEON, SVE, etc. The iteration count sets the number of times the test load is repeated to ensure statistical validity; the CPU core binding strategy specifies the logical or physical core to which the test thread is bound; the memory allocation node sets the NUMA node to which memory allocation belongs; and the hyper-threading switch enables or disables hyper-threading technology for comparative testing.
[0041] In one possible implementation, the matrix scaling mode in this embodiment adopts an automatic detection mode, which can be understood as automatically generating a series of scales covering the entire range from the first-level cache (i.e., L1 cache) to the memory DRAM overflow through subsequent steps.
[0042] It should be noted that the matrix size mode in this embodiment can also be manually specified. For example, one or more specific matrix sizes can be directly input based on user configuration parameters for targeted evaluation.
[0043] After obtaining the above user configuration parameters, the processor under test can parse the above user configuration parameters to automatically identify the processor architecture, query the SIMD feature set supported by the current hardware, determine the maximum vector width, confirm whether it supports multiply-accumulate fusion instructions and mask instructions, initialize high-precision timers, initialize performance counter access interfaces, and initialize power consumption data reading paths, etc.
[0044] Furthermore, this embodiment also needs to obtain the processor physical topology information (i.e., the aforementioned physical topology binding information) corresponding to the current user configuration parameters. In this embodiment, the physical topology binding information includes at least the mapping relationship between physical cores and logical cores, the hyper-threading relationship correspondence table, the die to which each physical core belongs, the NUMA node division method, and the socket affiliation.
[0045] It should be noted that this embodiment binds the CPU core and memory node with the current user configuration parameters, which can eliminate cross-node access drift that may be introduced by the default scheduling of the operating system.
[0046] Furthermore, to facilitate users in obtaining the hardware characteristics of the current user configuration parameters, this embodiment can further confirm and output the thread distribution characteristics corresponding to the current binding result. The thread distribution characteristics in this embodiment include, but are not limited to: whether a hyper-threaded logical core is used, whether there is cross-die behavior, whether there is cross-NUMA access, and whether cross-Socket communication is involved.
[0047] In one possible implementation, if the user plans to evaluate the impact of hyper-threading, they can run each standardized test load multiple times under different BIOS settings and binding policies, and automatically correlate and compare the data from each run.
[0048] In summary, this application prioritizes determining the physical boundaries of each subsequent test item based on user configuration parameters and automatically generates a hardware topology scheme based on the current hardware, thereby improving the problem that existing technologies cannot perform performance evaluation based on real hardware conditions.
[0049] After obtaining the physical topology binding information, this embodiment can further construct affinity constraints, thereby obtaining the performance calibration parameters of the processor under test (corresponding hardware) under these affinity constraints, and thus obtaining the upper limit of the actual computing power and memory bandwidth of the current processor under test, improving the accuracy of bottleneck analysis.
[0050] In this embodiment, the affinity constraint can be understood as the CPU core set, memory node identifier, and topology behavior constraint, to ensure that each subsequent calibration test item can be executed within the physical topology range corresponding to the above physical topology binding information.
[0051] In this embodiment, the performance calibration parameters include, but are not limited to, the actual calculated roof value, the theoretical calculated peak value, the measured bandwidth value of DRAM, and the cache capacity boundaries corresponding to each level of cache.
[0052] Please refer to Figure 3 , Figure 3 A flowchart of step 103 provided in an embodiment of the present invention is shown; in this embodiment, step 103 includes steps 1031 to 1033.
[0053] Step 1031: Obtain the actual calculated roof value and theoretical calculated peak value of the processor under test under affinity constraints based on the matrix multiplication kernel with a preset working set size.
[0054] In one possible implementation, this embodiment can perform roof calibration based on a matrix multiplication kernel with an extremely small working set size that resides entirely in the L1 cache, obtaining the actual calculated roof value and the theoretical calculated peak value. The block parameters of this matrix multiplication kernel are rigorously designed to ensure that all operands are always in the L1 cache during computation, thereby completely isolating the influence of the memory access subsystem and completing the measurement of the ultimate throughput capacity of the vector computation unit within the processor core.
[0055] It should be noted that this embodiment does not limit the specific calculation process and formula for the actual calculated roof value and the theoretical calculated peak value. In one possible implementation, during the operation of the matrix multiplication microkernel, the theoretical calculated peak value based on the nominal highest turbo frequency and the actual calculated roof value based on the measured maximum single-core floating-point performance, SIMD instruction width, and the number of FMA instructions that can be completed per cycle can be obtained during the operation of the matrix multiplication microkernel. The theoretical calculated peak value represents the upper limit of performance under ideal heat dissipation and power consumption conditions; the actual calculated roof value represents the upper limit of computing performance that is actually achievable under the current operating environment.
[0056] Step 1032: Obtain the measured DRAM bandwidth value of the memory subsystem in the processor under test under affinity constraints based on the preset DRAM streaming read / write load.
[0057] This embodiment can directly obtain the measured bandwidth value of DRAM by running a large-size matrix of streaming read and write loads. In addition, to ensure the accuracy of the measurement results, the streaming read and write load in this embodiment adopts a sequential access mode and prefetch optimization to fully stimulate the throughput capacity of the memory subsystem.
[0058] Step 1033: Obtain the cache capacity boundaries corresponding to each level of cache according to the cache order based on the preset working set size set. The preset working set size set includes a set of matrix multiplication kernels with working set sizes increasing stepwise.
[0059] This embodiment can obtain the cache capacity boundaries corresponding to each level of cache by running a set of workloads with progressively increasing working set sizes.
[0060] In one possible implementation, this embodiment can start with a very small scale, allowing the data to be completely contained in the L1 cache; then gradually increase the matrix dimension, causing the working set to overflow the L1 cache (level 1 cache), L2 cache (level 2 cache), and L3 cache (level 3 cache) in sequence, until it falls completely into the memory DRAM, and accurately record the performance point at each scale increase, marking the effective capacity boundary of each level of cache.
[0061] In addition, this embodiment can also construct a working set-performance step diagram based on the effective capacity boundaries of each level of cache to visually display the performance cliff position when the capacity of each level of cache overflows.
[0062] After obtaining the above performance calibration parameters, this embodiment can further construct a roof line diagram based on the performance calibration parameters.
[0063] In this embodiment, step 104 is implemented as follows: A coordinate system is constructed with arithmetic strength as the x-axis and measured floating-point performance as the y-axis. Two horizontal lines are drawn within this coordinate system, representing the actual calculated peak value and the theoretical calculated peak value, respectively, serving as the actual calculated peak value and the theoretical calculated peak value. A diagonal line is also drawn within the coordinate system, with the measured DRAM bandwidth as the slope, serving as the bandwidth peak value.
[0064] In one possible implementation, the actual calculated roof value in this embodiment can be represented by a dashed line, and the theoretically calculated peak value can be represented by a solid line; and the vertical drop between the dashed and solid lines can be used to visually reflect the frequency reduction loss caused by power consumption or temperature limitations.
[0065] In addition, this embodiment also marks the cache capacity boundaries corresponding to the above-mentioned cache levels, especially the cache capacity boundary (or L3 cache overflow point) corresponding to L3 cache, as auxiliary reference lines.
[0066] After constructing the performance calibration parameters and roofline diagram, this embodiment further constructs multiple standardized test loads based on user configuration parameters and physical topology binding information, and determines the performance parameters and microarchitecture metrics corresponding to each standardized test load, obtaining a multidimensional test dataset to achieve performance bottleneck analysis. The multiple standardized test loads constructed in this embodiment can be used to quantify the vectorization capabilities of the current processor under test. Each standardized test load corresponds to a key measurement.
[0067] Please refer to Figure 4 , Figure 4 A flowchart of step 200 provided in an embodiment of the present invention is shown; in this embodiment, step 200 includes steps 201 to 203.
[0068] Step 201: Construct multiple standardized test loads based on user configuration parameters and performance calibration parameters.
[0069] Step 202: Obtain the performance parameters and microarchitecture metrics of each standardized test load in any running cycle.
[0070] Step 203: Encapsulate the performance parameters and microarchitecture metrics of each standardized test load into a multidimensional test dataset according to the timestamp and the affinity label corresponding to each standardized test load.
[0071] Taking cache ladder scan load as an example, this embodiment can be used to calibrate the CPU cache performance of the processor under test. In one possible implementation, this embodiment can maintain its own computational intensity and gradually increase the matrix dimension, so that the working set gradually grows from fully adapting to the L1 cache to far exceeding the L3 cache. At the same time, an independent test run is performed at each size step, and the performance achieved is recorded, thus forming a "working set size-performance" ladder curve. Among them, each step on the curve where the performance drops sharply corresponds to the capacity boundary of the current level of cache.
[0072] Taking DRAM memory access bandwidth stress load as an example, in this embodiment, the DRAM memory access bandwidth stress load can be used to calibrate the bandwidth performance of the processor under test, such as the actual sustainable bandwidth of DRAM. In one possible implementation, the DRAM memory access bandwidth stress load can adopt sequential streaming read and write operations of a large matrix, whose working set far exceeds the capacity of the last-level cache. It is necessary to ensure that each data access must penetrate the cache and reach the DRAM directly to fully stimulate the throughput capacity of the memory subsystem.
[0073] Taking the arithmetic intensity scan workload as an example, in this embodiment, the arithmetic intensity scan workload can be used to calibrate the performance of the processor under test as a function of computational intensity. In one possible implementation, this embodiment utilizes the flexible adjustability of matrix multiplication block parameters to generate test kernels with computational intensities ranging from extremely low (approximately 0.05 FLOP / Byte, highly bandwidth-sensitive) to extremely high (tens of FLOP / Byte, highly computationally sensitive) by changing the matrix block size. This embodiment can fully cover the entire range of hardware from the bandwidth-limited region to the computationally-limited region based on the arithmetic intensity scan workload, obtaining a continuous curve reflecting the performance change with computational intensity.
[0074] Taking a specific test load as an example, this test load is characterized by running under the binding conditions specified by the user configuration parameters to determine the actual impact of topology factors such as hyper-threading, cross-die, cross-NUMA, or cross-Socket on vector performance.
[0075] In one possible implementation, this embodiment can construct two typical scenarios (e.g., cache-resident scenario and DRAM access scenario) based on the die grouping information in the physical topology binding information, and further distinguish the data residency levels. In this embodiment, the cache-resident scenario strictly controls the working set within the last-level cache capacity, and runs under configurations where threads are concentrated in a single die and evenly distributed across multiple dies, respectively, to reflect the performance difference brought about by the cross-die cache coherence protocol. In this embodiment, the DRAM access scenario makes the working set far exceed the last-level cache capacity, and combined with NUMA topology, runs under the conditions of using only single-die associated local memory and cross-die distributed memory access, to reflect the impact of cross-die memory access on effective bandwidth.
[0076] It should be noted that the specific test load in this embodiment is only used to measure the performance data and hardware behavior information of the current processor under test under different configurations, and does not restrict the binding method or hyper-threading settings adopted. The above configuration can be decided by the user according to the actual application scenario.
[0077] Furthermore, if users plan to evaluate the impact of hyper-threading, they can run the current specific test load multiple times under different BIOS settings and binding policies, and automatically correlate and compare the data from each run. Specifically, for the specific test load targeting hyper-threading behavior, users need to configure the BIOS to have hyper-threading disabled and enabled, and run the test in both states. This automatically identifies test points with the same configuration in the two runs, matches and compares their data, and outputs comparisons of performance differences, changes in floating-point execution unit utilization, and pipeline stalls, generating specific test comparison charts.
[0078] In one possible implementation, the specific test comparison chart in this embodiment can be presented as a bar chart or line chart to show the performance differences and changes in key indicators (data from multiple configurations such as hyper-threading enabled and disabled, local and cross-die binding, etc.).
[0079] Furthermore, this embodiment can summarize the topology behavior (whether hyper-threading is used, whether it crosses dies, whether it crosses NUMA, etc.) and its corresponding performance under the current configuration in descriptive language on a specialized test comparison chart, providing a theoretical basis for users to choose the optimal strategy according to actual conditions. For test items that cannot be executed due to hardware topology limitations, they should also be clearly marked and the reasons explained.
[0080] Assuming that hyperthreading is enabled and disabled once each, or that local tests within the same die and cross-die tests are performed separately, this embodiment will automatically compare the performance, bandwidth utilization, execution unit utilization, and pipeline pause metrics of the same test load under different configurations, and present the differences in the form of structured tables and graphs.
[0081] For a single run, this embodiment will directly output a description of the topology behavior under the current binding conditions, such as "current configuration uses hyper-threaded logical cores", "current configuration is distributed across dies" or "current configuration has cross-NUMA memory access", and will include corresponding performance and counter metrics for users to analyze and judge for themselves.
[0082] It should be noted that this embodiment does not automatically make decisions about the merits of hyper-threading or cross-die, but is limited to providing objective data (i.e., the above-mentioned structured tables and graphs or topological behavior descriptions) for users to make decisions.
[0083] In summary, this embodiment can synchronously collect and record effective frequency and package power consumption during the complete runtime of each standardized test load, and synchronously collect a series of microarchitectural events through the performance counter via perf_event or PMU interface, including the number of floating-point instruction retirements, the number of cache misses at each level, SIMD instruction width distribution statistics, TLB misses, etc., to obtain a multi-dimensional test dataset.
[0084] It should be noted that in this embodiment, the multidimensional test dataset includes all collected data such as performance parameters and microarchitecture metrics, all of which are accompanied by timestamps and core affinity tags.
[0085] Furthermore, this embodiment can also construct a data analysis graph based on the above-mentioned multi-dimensional test dataset, including frequency-performance relationship curves and frequency-power consumption relationship curves, to further demonstrate the changing trends of performance and power consumption with effective frequency, and mark the inflection point of energy efficiency ratio.
[0086] After obtaining the multidimensional test dataset, this embodiment can use the Roofline model (the roofline diagram mentioned above) as the macro-classification entry point, and use the incremental data of the microarchitecture performance counter and control experiment as the core basis for fine-grained root cause determination.
[0087] Specifically, this embodiment can map data from a multidimensional test dataset to a roofline plot, resulting in multiple mapped test points. In one possible implementation, this embodiment maps each test point from the multidimensional test dataset to the roofline plot according to its arithmetic strength and measured floating-point performance. In another possible implementation, this embodiment can also use color mapping to reflect the average effective frequency of each test point during operation.
[0088] It should be noted that each test point in the above multidimensional test dataset is a set of data points used to characterize the performance parameters and microarchitecture metrics of any standardized test load in any running cycle.
[0089] This embodiment can perform macroscopic triage based on the coordinate position of each test point on the roof line map, and divide each test point into the corresponding evaluation branch.
[0090] In another possible implementation, after constructing the standardized test load described above, this embodiment can perform performance evaluation based on OpenBLAS GEMM matrix operations, and combine performance analysis tools such as perf and turbostat to measure the vectorization capability of the CPU cores and collect relevant indicators.
[0091] Please refer to Figure 5 , Figure 5 The flowchart of step 300 provided in the embodiment of the present invention is shown; in this embodiment, step 300, which divides multiple mapped test points into test sets corresponding to the performance bottleneck category according to a preset division rule, includes steps 301 to 304.
[0092] Step 301: Select the first target test point from multiple mapping test points to construct a computational bottleneck test set.
[0093] In this embodiment, the first target test point is used to characterize that the coordinates of the current mapped test point are within the preset boundary threshold range of the actual calculated roof line.
[0094] In one possible implementation, this embodiment can be based on the coordinate position of each test point on the roof line diagram. For example, test points whose coordinates are close to the actual calculated roof line (i.e., fall within the preset boundary threshold range of the actual calculated roof line) indicate that their performance is mainly limited by the core computing power. Based on this, this embodiment can use test points that meet the above conditions as the first target test points to construct a set of computing bottleneck test points for computing bottleneck evaluation.
[0095] Step 302: Select the second target test point from multiple mapped test points to construct the DRAM bandwidth bottleneck test set.
[0096] In this embodiment, the second target test point is used to characterize that the coordinates of the current mapped test point are within the preset boundary threshold range of the bandwidth roofline.
[0097] In one possible implementation, test points whose coordinates are close to the measured bandwidth value of DRAM (i.e., falling within the preset boundary threshold range of the bandwidth roofline) represent performance that is mainly limited by memory bandwidth. Based on this, this embodiment can use test points that meet the above conditions as the second target test points to construct a DRAM bandwidth bottleneck test set for DRAM bandwidth bottleneck evaluation.
[0098] Step 303: Select a third target test point from multiple mapped test points to construct a cache capacity overflow bottleneck test set.
[0099] In this embodiment, the third target test point is used to characterize the deviation condition that the coordinates of the current mapped test point are lower than the actual calculated roof line and have not reached the boundary constraint of the bandwidth roof line. At the same time, its working set size exceeds the target-level cache capacity boundary and the miss rate at the target level exceeds the first preset failure rate threshold.
[0100] In this embodiment, the third target test point can be used to characterize that its position (in the roofline diagram) is outside the preset boundary threshold range of the actual calculated roofline and does not fall within the effective bandwidth range; it can be understood that the position of the test point is significantly lower than the actual calculated roofline and has not reached the measured bandwidth value of DRAM; at the same time, the working set of the test point itself has exceeded the capacity of a certain level of cache and is accompanied by a significant increase in the cache miss rate of that level.
[0101] It should be noted that this embodiment does not limit the specific values of the first preset failure rate threshold, the deviation conditions of the actual calculated roof line, and the boundary constraints of the bandwidth roof line, and these values can be flexibly configured by the user.
[0102] Step 304: Select the fourth target test point from multiple mapped test points to construct a microarchitecture bottleneck test set.
[0103] In this embodiment, the fourth target test point is used to characterize the current mapped test point whose own coordinates are far from the actual computational roofline and bandwidth roofline, and whose cache miss rate at each level is lower than the second preset miss rate threshold. Typically, test points whose own coordinates are significantly lower than all rooflines and whose cache miss rate is not high indicate a clear microarchitectural bottleneck outside the computational unit. In this embodiment, test points that meet the above conditions can be selected from the roofline diagram as the fourth test point for microarchitectural bottleneck analysis.
[0104] In summary, compared with traditional benchmark tests that only provide macro-level execution time scores, this application provides a more refined evaluation of CPU core computing power, memory bandwidth, power consumption, and other dimensions. It also supports performance evaluation under hyper-threading and cross-die scenarios, thus improving the compatibility of this computing power performance evaluation method.
[0105] After the test sets are divided, this embodiment performs analysis in parallel based on the bottleneck determination rules of each test set.
[0106] Example 1 Taking the computational bottleneck test set as an example, the performance parameters in this embodiment also include the number of instructions executed per cycle and the percentage of idle cycles in the instruction scheduling queue.
[0107] Please refer to Figure 6 , Figure 6 A flowchart of step 400 provided in an embodiment of the present invention is shown; in this embodiment, step 400 includes steps 401A to 403A.
[0108] Step 401A: For each first target test point under the computing bottleneck test set, determine whether the effective frequency of the first target test point itself is lower than the nominal turbo frequency.
[0109] If so, then the current first target test point is determined to be the frequency loss bottleneck.
[0110] Step 402A: If it is not lower than the nominal turbo frequency, then determine whether the SIMD width utilization rate of the current first target test point exceeds the preset vector width rule based on the SIMD instruction width distribution of the current first target test point. If so, then determine that the current first target test point has insufficient vector width utilization.
[0111] Step 403A: If the preset vector width rule is met, determine whether the number of instructions executed by the current first target test point in each cycle meets the preset IPC rule. If yes, determine that the current first target test point is a pure computational bottleneck; if no, determine that the current first target test point is insufficient in instruction-level parallelism.
[0112] This embodiment can prioritize determining the frequency loss bottleneck based on the effective frequency of any first target test point in the computational bottleneck test set. If the effective frequency of the first target test point is significantly lower than the nominal turbo frequency, for example, lower than the lower limit of the preset error range of the nominal turbo frequency (e.g., set to be 0.05GHz lower than the range of the nominal turbo frequency), it is determined to be a frequency loss bottleneck. In addition, if the actual package power consumption has reached the theoretical power consumption limit (TDP limit) or the temperature sensor on the processor under test has triggered the throttling mechanism, this embodiment can also simultaneously mark the bottleneck cause and improvement suggestions at the first target test point when outputting the frequency loss bottleneck. Taking the first target test point as an example, it can output: The bottleneck cause of the first target test point is: frequency reduction caused by power consumption wall or temperature wall, and the improvement suggestion is: improve the system heat dissipation conditions or relax the processor power consumption limit to restore the frequency.
[0113] If the effective frequency of the current first target test point is not lower than the nominal turbo frequency, then the SIMD instruction width distribution of the current first target test point can be used to determine whether its own SIMD width utilization (i.e. the vector width distribution actually used by the current first target test point in analyzing floating-point instructions) exceeds the preset vector width rule. For example, if the proportion of narrow-width operations (such as 128-bit or 256-bit) is too high and the processor under test itself supports wider vector instructions (such as 512-bit or SVE variable width), then it is diagnosed as insufficient vector width utilization.
[0114] If the preset vector width rule is met, it can be further determined whether the number of instructions executed per cycle at the current first target test point meets the preset IPC rule, that is, check the number of instructions completed per cycle. If the value is significantly lower than the theoretical peak and the idle cycle of the instruction scheduling queue is relatively high, it indicates that the computing unit is waiting for the data dependency to be resolved, and the diagnosis is insufficient instruction-level parallelism. If all the above indicators are close to the theoretical limit, it is determined to be a pure computing bottleneck, indicating that the vector computing unit in the processor under test has been fully utilized.
[0115] Example 2 For example, regarding the DRAM bandwidth bottleneck, please refer to... Figure 7 , Figure 7 A flowchart of another step of step 400 provided in an embodiment of the present invention is shown; in this embodiment, step 400 includes step 401B.
[0116] Step 401B: For each second target test point under the DRAM bandwidth bottleneck test set, determine whether the DRAM bandwidth utilization rate of the second target test point meets the preset range of the DRAM measured bandwidth value. If it does, determine that the current second target test point belongs to the DRAM bandwidth bottleneck; otherwise, determine that the current second target test point does not belong to the DRAM bandwidth bottleneck.
[0117] In one possible implementation, this embodiment can determine whether the DRAM bandwidth utilization rate of the second target test point is close to the actual DRAM bandwidth value (i.e., meets the preset range of the actual DRAM bandwidth value). If so, it is diagnosed as a DRAM bandwidth bottleneck, indicating that the memory access rate of the current load has reached the hardware physical limit; otherwise, it has not reached the DRAM bandwidth bottleneck.
[0118] Similar to the previous embodiment, this embodiment can also simultaneously mark its own bottleneck improvement suggestions at the second target test point when outputting the DRAM bandwidth bottleneck. For example, each second target test point with a DRAM bandwidth bottleneck can be marked: This second target test point has a DRAM bandwidth bottleneck, and it is recommended to use higher frequency memory or increase the number of memory channels to obtain linear bandwidth gain.
[0119] Example 3 For example, regarding the cache capacity overflow bottleneck, please refer to... Figure 8 , Figure 8 A flowchart of another step of step 400 provided in an embodiment of the present invention is shown; in this embodiment, step 400 includes steps 401C to 403C.
[0120] Step 401C: For each third target test point under the cache capacity overflow bottleneck test set, if the working set of the third target test point itself exceeds the L1 cache capacity boundary and the L1 cache miss rate exceeds the preset burst judgment rule, then the current third target test point is determined to meet the L1 cache capacity overflow.
[0121] Step 402C, or if the working set of the third target test point exceeds the L2 cache capacity boundary and the L2 cache miss rate exceeds the preset burst judgment rule; then the current third target test point is determined to meet the L2 cache capacity overflow.
[0122] Step 403C, or if the working set of the third target test point exceeds the capacity boundary of the L3 cache, and the miss rate of the L3 cache exceeds the preset burst judgment rule; then it is determined that the current third target test point meets the L3 cache capacity overflow.
[0123] It should be noted that this embodiment does not limit the implementation of the above steps; they can be implemented in parallel or sequentially. In one possible implementation, to improve evaluation efficiency, this embodiment adopts a parallel execution method.
[0124] In one possible implementation, this embodiment can directly determine the size of the working set of each third target test point. For example, if the working set of the current third target test point exceeds the capacity of the first-level cache and the miss rate of the current first-level cache exceeds the preset percentage of the preset miss value, it is determined to meet the sudden increase determination rule, indicating that the miss rate of the current first-level cache has suddenly increased, and it is determined to be a first-level cache capacity overflow.
[0125] Similarly, if the working set of the current third target test point exceeds the capacity of the second-level cache, and if the miss rate of the current second-level cache exceeds the preset percentage of the preset miss value, it is determined to meet the sudden increase judgment rule, indicating that the miss rate of the current second-level cache has suddenly increased, and it is determined to be a second-level cache capacity overflow.
[0126] Similarly, if the working set of the current third target test point exceeds the capacity of the L3 cache, and if the miss rate of the current L3 cache exceeds the preset percentage of the preset miss value, it is determined to meet the burst judgment rule. If the miss rate of the current L3 cache increases significantly, it is determined to be an L3 cache capacity overflow.
[0127] It should be noted that this embodiment does not limit the specific values of the preset miss value and its preset percentage. The configuration of this parameter can be flexibly set by the user.
[0128] Furthermore, following the same approach as the previous embodiment, this embodiment can also simultaneously mark its own bottleneck improvement suggestions at the third target test point when outputting the corresponding cache capacity overflow bottleneck. Taking the third target test point that satisfies the third-level cache capacity overflow as an example, it can be marked: This third target test point has a third-level cache capacity overflow bottleneck, and it is recommended to adjust the matrix block size to within the third-level cache capacity in order to reuse the higher bandwidth cache level.
[0129] Example 4 For an example of microarchitectural bottlenecks, please refer to [link / reference]. Figure 9 , Figure 9 A flowchart of another step of step 400 provided in an embodiment of the present invention is shown; in this embodiment, step 400 includes steps 401D to 403D.
[0130] Step 401D: For each fourth target test point under the microarchitecture bottleneck test set, determine whether the TLB miss rate of the fourth target test point exceeds the preset miss threshold. If so, determine that the current fourth target test point belongs to the TLB-type microarchitecture bottleneck. Step 402D: Determine whether the front-end stall rate of the fourth target test point exceeds the preset stall rate. If so, determine that the current fourth target test point belongs to the front-end stall type microarchitectural bottleneck. Step 403D: Determine whether the branch prediction failure rate of the fourth target test point exceeds the preset failure rate. If so, determine that the current fourth target test point belongs to the branch prediction failure type of microarchitecture bottleneck.
[0131] It should be noted that this embodiment does not limit the implementation of the above steps; they can be implemented in parallel or sequentially. In one possible implementation, to improve evaluation efficiency, this embodiment adopts a parallel execution method.
[0132] This embodiment can evaluate each fourth target test point based on the above TLB miss rate, front-end stall rate, or branch prediction failure rate. If all indicators are normal, it means that no obvious bottleneck has been found based on the above analysis. At this time, a manual analysis mark should be output to remind the user to optimize.
[0133] Furthermore, following the same approach as the previous embodiment, this embodiment can also simultaneously annotate its own bottleneck improvement suggestions at the fourth target test point when outputting the corresponding microarchitectural bottleneck. Further details will not be elaborated here.
[0134] In summary, this embodiment can accurately decouple the kernel vector bottleneck and memory access bottleneck by separating peak detection and memory bandwidth detection and combining adjustable arithmetic intensity scanning; at the same time, it can automatically locate performance bottlenecks based on the roofline model automatic engine and microarchitecture counter.
[0135] Based on this, this embodiment can perform bottleneck analysis based on the above steps to obtain the bottleneck results and their causes corresponding to each test point. This embodiment can overlay the above analysis results onto the roofline diagram to obtain a multi-level roofline analysis diagram. The multi-level roofline analysis diagram includes the theoretically calculated roofline, the actual calculated roofline, the bandwidth roofline, and the L3 cache capacity overflow point, and clearly marks the bottleneck category to which each test point belongs, so as to form a bottleneck analysis report.
[0136] It should be noted that the bottleneck analysis report in this embodiment includes at least the aforementioned multidimensional test dataset, multi-level roofline analysis diagram, special test comparison chart, working set-performance step diagram, and data analysis diagram.
[0137] In summary, this application constructs multiple standardized test loads based on vector performance influencing factors such as hyper-threading, system architecture, CPU cache, DDR bandwidth, clock speed, and cross-die behavior to obtain a multi-dimensional test dataset corresponding to the processor under test. Based on the multi-dimensional test dataset, bottleneck performance analysis is automatically performed. Thus, based on the evaluation of CPU vector computing power, performance bottleneck analysis is automatically realized, effectively revealing the performance influencing factors such as processor core computing power, memory bandwidth, cross-die behavior, and hyper-threading.
[0138] The same idea applies as the previous embodiment; please refer to [the previous embodiment]. Figure 10 , Figure 10 A block diagram of a computing power performance evaluation device provided in an embodiment of the present invention is shown; the computing power performance evaluation device is applied to the computing power performance evaluation method of any of the first aspects described above, and the computing power performance evaluation device 10 includes: The performance calibration module 11 is used to obtain user configuration parameters, obtain performance calibration parameters of the processor under test under user configuration parameters based on the roofline model, and generate a roofline diagram; the performance calibration parameters are used to characterize the computing power and memory bandwidth of the processor under test.
[0139] The data acquisition module 12 is used to acquire performance parameters and microarchitecture metrics corresponding to multiple standardized test loads based on user configuration parameters, and obtain a multidimensional test dataset. The multiple standardized test loads include at least cache ladder scan load, DRAM memory access bandwidth pressure load, arithmetic strength scan load, and special test loads. The special test loads are used to characterize cross-die behavior and / or test loads corresponding to hyper-threading.
[0140] The bottleneck segmentation module 13 is used to map each performance parameter in the multidimensional test dataset to the roofline diagram to obtain multiple mapped test points, and to divide the multiple mapped test points into test sets corresponding to the performance bottleneck categories according to the preset segmentation rules; each test set corresponds to a bottleneck determination rule.
[0141] The bottleneck analysis output module 14 is used to perform bottleneck analysis on each test set based on its own corresponding bottleneck judgment rules and obtain a bottleneck analysis report.
[0142] The specific implementation scheme and technical effects of the computing power performance evaluation device in this embodiment can be referred to the computing power performance evaluation method provided in the previous embodiment. Its basic principle and the resulting technical effects are the same as those in the above embodiments. For the sake of brevity, any parts not mentioned in this embodiment can be referred to the corresponding content in the above embodiments.
[0143] The same idea applies as the previous embodiment; please refer to [the previous embodiment]. Figure 11 , Figure 11 A block diagram of a computing power performance evaluation system provided in an embodiment of the present invention is shown. This computing power performance evaluation system is applied to the computing power performance evaluation method of any of the first aspects described above. The computing power performance evaluation system 20 includes a user configuration layer 21, a scheduling and execution engine layer 22, a platform abstraction layer 23, a system monitoring layer 24, and an output layer 25. It obtains user configuration parameters, acquires performance calibration parameters of the processor under test under the user configuration parameters based on the roofline model, and generates a roofline diagram. Based on the user configuration parameters, it obtains performance parameters and microarchitecture index parameters corresponding to multiple standardized test loads to obtain a multidimensional test dataset. Subsequently, it maps each data point in the multidimensional test dataset to the roofline diagram to obtain multiple mapped test points, and divides these multiple mapped test points into test sets corresponding to performance bottleneck categories according to preset partitioning rules. Each test set performs bottleneck analysis based on its corresponding bottleneck determination rules to obtain a bottleneck analysis report.
[0144] In this embodiment, the user configuration layer 21 is used to obtain user configuration parameters. This embodiment supports two configuration modes: automatic detection and manual specification. Users can obtain specific configuration items through the user configuration layer, such as matrix size, number of iterations, CPU cores, memory nodes, instruction type, and hyper-threading switch.
[0145] The scheduling and execution engine layer 22 is used to generate and execute standardized workloads based on user-configured parameters, while simultaneously calling the unified interface provided by the platform abstraction layer. In this embodiment, the scheduling and execution engine layer may include an x86 SIMD instruction module, an ARM SIMD instruction module, and a benchmark kernel library. The x86 SIMD instruction module implements the computing kernel based on the x86 platform SIMD instruction set (AVX, AVX2, AVX-512); the ARM SIMD instruction module implements the computing kernel based on the ARM platform SIMD instruction set (NEON, SVE); and the benchmark kernel library can generate three types of workloads: small matrix (size adapted to L1 cache) core computing workloads (i.e., matrix multiplication kernels with a very small working set size, residing entirely in the L1 cache), large matrix (size exceeding L3 cache) streaming access and memory bandwidth test workloads (e.g., preset DRAM streaming read / write workloads), and mixed test workloads with configurable arithmetic strength.
[0146] Platform abstraction layer 23, as the underlying implementation architecture encapsulating hardware platform-related aspects, provides a unified API interface to the upper layers, independent of instruction sets, timing, system calls, and memory operations. Specifically, in one possible implementation, conditional compilation supports both x86 and ARM backends, and uniformly encapsulates high-precision timing functions, operating system calls, and cross-platform memory allocation interfaces.
[0147] The system monitoring layer 24 is used to collect system operating status and microarchitectural events in real time, providing raw data for performance analysis and bottleneck diagnosis. Collection dimensions include top-down data collection, memory bandwidth, CPU frequency, and power consumption.
[0148] Output layer 25 generates a final analysis report based on performance data generated by the scheduling execution engine layer and runtime metrics collected by the system monitoring layer. The output includes a roofline model and a visualization report.
[0149] It should be noted that the specific implementation scheme and technical effects of the computing power performance evaluation system in this embodiment can refer to the computing power performance evaluation method provided in the previous embodiment. Its basic principle and the resulting technical effects are the same as those in the above embodiments. For the sake of brevity, any parts not mentioned in this embodiment can be referred to the corresponding content in the above embodiments.
[0150] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative; for example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0151] In addition, the functional modules in the various embodiments of the present invention can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0152] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0153] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for evaluating computing performance in vector computing scenarios, characterized in that, The computing power performance evaluation method includes: Obtain user configuration parameters, and based on the roofline model, obtain the performance calibration parameters of the processor under test under the user configuration parameters to generate a roofline diagram; the performance calibration parameters are used to characterize the computing power and memory bandwidth of the processor under test. Based on the user configuration parameters, performance parameters and microarchitecture metrics corresponding to multiple standardized test loads are obtained to obtain a multidimensional test dataset. The multiple standardized test loads include at least cache ladder scan load, DRAM memory access bandwidth pressure load, arithmetic strength scan load, and specialized test loads. The specialized test loads are used to characterize cross-die behavior and / or test loads corresponding to hyper-threading. Each test point in the multidimensional test dataset is mapped to the roofline diagram to obtain multiple mapped test points. These multiple mapped test points are then divided into test sets corresponding to the performance bottleneck categories according to a preset partitioning rule. Each test set corresponds to a bottleneck determination rule. Each test set performs bottleneck analysis based on its own corresponding bottleneck determination rules, and obtains a bottleneck analysis report.
2. The computing power performance evaluation method according to claim 1, characterized in that, When the performance calibration parameters include the actual calculated roof value, the theoretical calculated peak value, the measured DRAM bandwidth value, and the cache capacity boundaries corresponding to each level of cache, the step of obtaining the performance calibration parameters of the processor under test under the user-configured parameters based on the roofline model and generating the roofline diagram includes: Based on the user configuration parameters, the physical topology binding information of the processor under test is obtained, and the thread distribution characteristics of the processor under test under the user configuration parameters are obtained. Affinity constraints are constructed based on the physical topology binding information; Based on preset rules, the actual calculated roof value, theoretical calculated peak value, DRAM measured bandwidth value, and cache capacity boundaries corresponding to each level of cache of the processor under test are obtained under the affinity constraint conditions. The roofline diagram is constructed based on the actual calculated roof value, the theoretical calculated peak value, and the measured DRAM bandwidth value.
3. The computing power performance evaluation method according to claim 2, characterized in that, The steps for constructing a roofline coordinate system based on the actual calculated roof value, the theoretical calculated peak value, and the measured DRAM bandwidth value, and generating the roofline map, include: Construct a coordinate system with arithmetic strength as the x-axis and measured floating-point performance as the y-axis; Two horizontal lines are drawn in the coordinate system based on the actual calculated roof value and the theoretical calculated peak value, respectively serving as the actual calculated roof line and the theoretical calculated roof line. Draw a diagonal line in the coordinate system with the measured bandwidth value of the DRAM as the slope, and use it as the bandwidth roofline.
4. The computing power performance evaluation method according to claim 1, characterized in that, The step of obtaining performance parameters and microarchitecture metrics corresponding to multiple standardized test loads based on the user configuration parameters to obtain a multidimensional test dataset includes: Multiple standardized test loads are constructed based on the user configuration parameters and the performance calibration parameters; Obtain the performance parameters and microarchitecture metrics of each standardized test load in any running cycle. The performance parameters include effective frequency, package power consumption, measured floating-point performance, SIMD instruction width distribution, and the miss rate corresponding to each level of cache data. The microarchitecture metrics include TLB miss rate, front-end stall rate, and branch prediction failure rate. The performance parameters and microarchitecture metrics of each standardized test load are encapsulated into the multidimensional test dataset according to the timestamp and the affinity label corresponding to each standardized test load.
5. The computing power performance evaluation method according to claim 1, characterized in that, When the roofline diagram includes the actual calculated roofline, the theoretical calculated roofline, the bandwidth roofline, and the cache capacity boundaries corresponding to each level of cache, and the performance bottleneck categories include computational bottlenecks, cache capacity overflow bottlenecks, DRAM bandwidth bottlenecks, and microarchitecture bottlenecks, the step of dividing the multiple mapped test points into test sets corresponding to the performance bottleneck categories according to preset division rules includes: A first target test point is selected from the plurality of mapped test points to construct a computational bottleneck test set; the first target test point is used to characterize that the coordinates of the current mapped test point are within the preset boundary threshold range of the actual calculated roof line. A second target test point is selected from the plurality of mapped test points to construct a DRAM bandwidth bottleneck test set; the second target test point is used to characterize that the coordinates of the current mapped test point are within the preset boundary threshold range of the bandwidth roofline. A third target test point is selected from the multiple mapping test points to construct a cache capacity overflow bottleneck test set; the third target test point is used to characterize the deviation condition that the current mapping test point's own coordinates are lower than the actual calculated roof line and have not reached the boundary constraint of the bandwidth roof line, while its own working set size exceeds the target level cache capacity boundary and the miss rate at the target level exceeds the first preset failure rate threshold. A fourth target test point is selected from the multiple mapping test points to construct a microarchitecture bottleneck test set; the fourth target test point is used to characterize that the coordinates of the current mapping test point are far away from the actual calculation roof line and the bandwidth roof line, and the miss rate of its own cache at all levels is lower than the second preset miss rate threshold.
6. The computing power performance evaluation method according to claim 1 or 5, characterized in that, When the test set includes a bottleneck test set, the steps for each test set to perform bottleneck analysis based on its corresponding bottleneck determination rule and obtain a bottleneck analysis report include: For each first target test point under the aforementioned computational bottleneck test set, determine whether the effective frequency of the first target test point itself is lower than the nominal turbo frequency; If so, then the current first target test point is determined to be the frequency loss bottleneck; If the frequency is not lower than the nominal turbo frequency, then the SIMD instruction width distribution of the current first target test point is used to determine whether the SIMD width utilization rate exceeds the preset vector width rule. If so, the current first target test point is determined to be underutilized in terms of vector width. If the preset vector width rule is met, then it is determined whether the number of instructions executed by the current first target test point in each cycle meets the preset IPC rule. If yes, then the current first target test point is determined to be a pure computational bottleneck; if no, then the current first target test point is determined to be insufficient instruction-level parallelism.
7. The computing power performance evaluation method according to claim 1 or 5, characterized in that, When the performance calibration parameters include cache capacity boundaries corresponding to each level of cache, and the test set includes a cache capacity overflow bottleneck test set, the steps for each test set to perform bottleneck analysis based on its corresponding bottleneck determination rules and obtain a bottleneck analysis report include: For each third target test point under the cache capacity overflow bottleneck test set, if it is determined that the working set of the third target test point itself exceeds the first-level cache capacity boundary, and the miss rate of the first-level cache exceeds the preset burst judgment rule, then the current third target test point is determined to meet the first-level cache capacity overflow. If it is determined that the working set of the third target test point exceeds the L2 cache capacity boundary, and the miss rate of the L2 cache exceeds the preset burst judgment rule, then it is determined that the current third target test point meets the L2 cache capacity overflow condition. If it is determined that the working set of the third target test point exceeds the capacity boundary of the L3 cache, and the miss rate of the L3 cache exceeds the preset burst judgment rule, then the current third target test point is determined to meet the L3 cache capacity overflow.
8. The computing power performance evaluation method according to claim 1 or 5, characterized in that, When the performance calibration parameters include the measured bandwidth value of DRAM, and the test set includes a DRAM bandwidth bottleneck test set, the steps for each test set to perform bottleneck analysis based on its corresponding bottleneck determination rules and obtain a bottleneck analysis report include: For each second target test point under the DRAM bandwidth bottleneck test set, it is determined whether the DRAM bandwidth utilization rate of the second target test point meets the preset range of the measured DRAM bandwidth value. If it does, the current second target test point is determined to be a DRAM bandwidth bottleneck; otherwise, the current second target test point is determined not to be a DRAM bandwidth bottleneck.
9. The computing power performance evaluation method according to claim 1 or 5, characterized in that, When the microarchitecture metrics include TLB miss rate, frontend stall rate, and branch prediction failure rate, and the test set includes a microarchitecture bottleneck test set, the steps for each test set to perform bottleneck analysis based on its corresponding bottleneck determination rules and obtain a bottleneck analysis report include: For each fourth target test point under the microarchitecture bottleneck test set, determine whether the TLB miss rate of the fourth target test point exceeds the preset miss threshold. If so, determine that the current fourth target test point belongs to the TLB-type microarchitecture bottleneck. Alternatively, determine whether the front-end stall rate of the fourth target test point exceeds the preset stall rate. If so, determine that the current fourth target test point belongs to the front-end stall type microarchitectural bottleneck. Alternatively, determine whether the branch prediction failure rate of the fourth target test point exceeds the preset failure rate. If so, determine that the current fourth target test point belongs to the branch prediction failure type of microarchitectural bottleneck.
10. A computing power performance evaluation device, characterized in that, The computing power performance evaluation device, applied to the computing power performance evaluation method according to any one of claims 1 to 9, comprises: The performance calibration module is used to obtain user configuration parameters, obtain performance calibration parameters of the processor under test under the user configuration parameters based on the roofline model, and generate a roofline diagram; the performance calibration parameters are used to characterize the computing power and memory bandwidth of the processor under test; The data acquisition module is used to acquire performance parameters and microarchitecture metrics corresponding to multiple standardized test loads based on the user configuration parameters, thereby obtaining a multidimensional test dataset. The multiple standardized test loads include at least cache ladder scan load, DRAM memory access bandwidth pressure load, arithmetic strength scan load, and specialized test loads. The specialized test loads are used to characterize cross-die behavior and / or test loads corresponding to hyper-threading. The bottleneck segmentation module is used to map each test point in the multidimensional test dataset to the roofline diagram to obtain multiple mapped test points, and to divide the multiple mapped test points into test sets corresponding to the performance bottleneck categories according to preset segmentation rules; each test set corresponds to a bottleneck determination rule; The bottleneck analysis output module is used to perform bottleneck analysis on each test set based on its own corresponding bottleneck determination rules and generate a bottleneck analysis report.