Graphic processor performance analysis method and system, electronic equipment and computer storage medium
By obtaining the workload signal and key time information of the graphics processor functional module, analyzing and judging the bottleneck module, the problem of low efficiency in the existing technology is solved, and efficient performance analysis in the early design stage is achieved.
Patent Information
- Application Number
- CN202510533751.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-08
AI Technical Summary
Existing graphics processor performance analysis methods are inefficient, unable to accurately identify performance bottlenecks in early design stages, and require built-in performance counters to increase chip area and power consumption, and cannot monitor all performance indicators simultaneously.
By obtaining the workload signals of each functional module of the graphics processor, recording key time information, using key time information to determine the bottleneck functional module, and outputting performance analysis reports.
Quickly identify graphics processor performance bottlenecks, improve analysis efficiency, reduce design modification time, and reduce chip area and power consumption.
Smart Images

Figure CN120276954A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of processor performance testing, and in particular to a graphics processor performance analysis method and system, electronic equipment and computer storage medium. Background Art
[0002] Graphics Processing Unit (GPU) is a coprocessor used to process images and graphics operations, and is widely used in personal computers, workstations, smart phones, tablets and other devices. In order to ensure the performance of the graphics processor, it is necessary to perform performance testing on it during the design phase to confirm whether the designed graphics processor meets the set performance goals.
[0003] Currently, GPU evaluation tools such as GFXBench, 3DMark, and BaseMark are usually used to analyze the performance of graphics processors. The process usually involves counting the workload signals of each functional module of the graphics processor, and then comparing the statistical values with the expected values to obtain a performance score. Finally, the performance of the graphics processor is judged based on the performance score.
[0004] However, since there are usually multiple functional modules that affect each other in a test case, it is necessary to check the workload signal of each functional module to find the bottleneck affecting the performance of the GPU, which is a very time-consuming process. Moreover, although the existing GPU performance analysis method can tell the location of the GPU performance bottleneck on the pipeline, it cannot accurately tell the functional module, initial location and root cause that causes the performance bottleneck.
[0005] In addition, the existing graphics processor performance analysis method must wait until the entire graphics processor chip can run actual applications. At this time, the design of the entire graphics processor chip has been basically completed, and most of the functional verification at the module level and chip level has also been completed. If a bottleneck is found during performance analysis and the design is modified, the verification work that has been done needs to be redone, which undoubtedly increases the tape-out time of the graphics processor chip.
[0006] In addition, in order to adapt to the existing performance analysis methods, the graphics processor needs to have built-in scalar performance counters and vector performance counters, which will increase the area and power consumption of the graphics processor chip; moreover, all scalar performance counters and vector performance counters cannot count all monitored performance indicators at the same time, but can only count in time-sharing, so multiple performance analysis simulations are required to confirm the bottleneck. Summary of the invention
[0007] The object of the present invention is to provide a method and system for analyzing the performance of a graphics processing unit, an electronic device and a computer storage medium, so as to solve the problem of low efficiency of the existing method for analyzing the performance of a graphics processing unit.
[0008] To solve the above technical problems, the present invention provides a method for analyzing the performance of a graphics processing unit, including: Running a test case; Obtaining the workload signals of each functional module in the graphics processing unit; Counting the workload signals of each functional module and recording the key time information of each workload signal; Using the key time information to analyze and judge the actual performance of the functional module to confirm the bottleneck functional module; Outputting a performance analysis report.
[0009] Optionally, in the method for analyzing the performance of a graphics processing unit, the key time information includes valid time points, invalid time points, valid durations, and invalid durations.
[0010] Optionally, in the method for analyzing the performance of a graphics processing unit, the method of using the key time information to analyze and judge the actual performance of the functional module to confirm the bottleneck functional module includes: Obtaining the first invalid time point of each workload signal; Sorting the workload signals according to the time sequence; Identifying the functional module corresponding to the workload signal ranked first as the bottleneck functional module.
[0011] Optionally, in the method for analyzing the performance of a graphics processing unit, the method of using the key time information to analyze and judge the actual performance of the functional module to confirm the bottleneck functional module further includes: Obtaining the time point when each workload signal first exceeds a preset invalid threshold; Sorting the workload signals according to the time sequence; Identifying the functional module corresponding to the workload signal ranked first as the bottleneck functional module.
[0012] Optionally, in the method for analyzing the performance of a graphics processing unit, the method of using the key time information to analyze and judge the actual performance of the functional module to confirm the bottleneck functional module further includes: Obtaining the configuration file corresponding to the test case; Calculating the performance target value of each functional module by using the configuration file; Calculating the performance score of the functional module according to the actual performance value and the performance target value of the functional module; Compare the performance scores calculated by each functional module with the preset expected scores; If the performance score of a certain functional module is less than the preset expected score, find the functional module corresponding to the workload signal where the earliest invalid time point is located, and identify this functional module as the candidate bottleneck functional module; Sort the workload signals of the candidate bottleneck functional module according to the chronological order of the time points when they first exceed the preset invalid threshold, find the functional module corresponding to the workload signal where the earliest time point of first exceeding the preset invalid threshold is located, and identify this functional module as the bottleneck functional module.
[0013] To solve the above technical problems, the present invention also provides a graphics processor performance analysis system for implementing the graphics processor performance analysis method described in any one of the above, and the graphics processor performance analysis system includes: A recording module for obtaining the workload signals of each functional module in the graphics processor, counting the workload signals of each functional module, and recording the key time information of each workload signal; A processing module for analyzing and judging the actual performance of the functional module according to the data information statistically obtained by the recording module to confirm the bottleneck functional module; A reporting module for outputting a performance analysis report.
[0014] Optionally, in the graphics processor performance analysis system, the recording module is integrated in the performance analysis simulation environment, and the design language used is the same as the hardware description language used in the graphics processor design.
[0015] Optionally, in the graphics processor performance analysis system, the recording module includes a data input unit, at least one signal sampling unit, and a recording output unit; the input module is used to obtain the workload signals, clock signals, control signals, and recording configuration files of each functional module in the graphics processor; the signal sampling unit corresponds to each functional module of the graphics processor to obtain the workload signals of each functional module in the graphics processor, count the workload signals of each functional module according to the clock signal, the control signal, and the configuration file, and record the key time information of each workload signal; the recording output unit is used to count the workload signals, the count values of the workload signals, and the key time information obtained by all the signal sampling units, and send the statistically obtained data information to the processing module.
[0016] Optionally, in the graphics processor performance analysis system, the processing module includes a receiving unit, an arithmetic unit, a reordering unit, and a result output unit; the receiving unit is configured to receive the data information and the processing configuration file statistically recorded by the recording module; the arithmetic unit is configured to use the processing configuration file to calculate the performance target value and the performance score of each functional module according to the received data information; the reordering module is configured to sort the workload signals according to the chronological order of the first invalid time points of the respective workload signals, and / or, configured to sort the workload signals according to the chronological order of the time points when the respective workload signals first exceed a preset invalid threshold, and / or, configured to sort the workload signals according to the chronological order of each invalid time point of the respective workload signals; the result output unit is configured to confirm the bottleneck functional module according to the sorting result of the reordering module and output the same.
[0017] To solve the above technical problems, the present invention further provides an electronic device, including a memory, a processor, and an executable program stored on the memory and capable of being run by the processor; when the processor runs the executable program, it executes the graphics processor performance analysis method described in any one of the above.
[0018] To solve the above technical problems, the present invention further provides a computer storage medium, which stores an executable program; when the executable program is executed, it implements the graphics processor performance analysis method described in any one of the above.
[0019] The graphics processor performance analysis method, system, electronic device, and computer storage medium provided by the present invention include: running a test case; obtaining workload signals of each functional module in the graphics processor; counting the workload signals of each functional module and recording the key time information of each workload signal; using the key time information to analyze and judge the actual performance of the functional module to confirm the bottleneck functional module; outputting a performance analysis report. By sampling and analyzing the workload signals of each functional module in the graphics processor, the change in the data throughput rate of each functional module can be quickly detected using the key time information, so as to quickly confirm the bottleneck functional module, improve the efficiency of graphics processor performance analysis, and solve the problem of low efficiency of existing graphics processor performance analysis methods. Description of the Drawings
[0020] Figure 1 It is a flowchart of the graphics processor performance analysis method provided in this embodiment; Figure 2 It is a schematic diagram of the data flow in a simple pipeline model composed of 3 functional modules provided in this embodiment; Figure 3 It is a structural block diagram of the graphics processor performance analysis system provided in this embodiment; Figure 4 is the structural block diagram of the recording module provided in this embodiment; Figure 5 is the timing diagram of the counting process of the signal sampling unit provided in this embodiment; Figure 6 is the timing diagram of the process of using the sliding window algorithm to count the time points and durations exceeding the invalid threshold provided in this embodiment; Figure 7 is the structural schematic diagram of the processing module provided in this embodiment. Detailed implementation manners
[0021] The following further elaborates on the graphics processor performance analysis method and system, electronic device, and computer storage medium proposed by the present invention in conjunction with the accompanying drawings and specific embodiments. It should be noted that the accompanying drawings are all in a very simplified form and use non-precise scales, only for conveniently and clearly assisting in explaining the purpose of the embodiments of the present invention. In addition, the structures shown in the accompanying drawings are often part of the actual structures. In particular, the accompanying drawings need to show different emphases and sometimes use different scales.
[0022] It should be noted that the "first", "second", etc. in the description, claims, and accompanying drawings of the present invention are used to distinguish similar objects for describing the embodiments of the present invention, rather than for describing a specific order or sequence. It should be understood that such structures can be interchanged under appropriate circumstances. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily limit to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0023] Currently, when designing a graphics processor, generally, the architecture design is first carried out to determine the design specifications and performance indicators of each functional module, and then each functional module is designed according to the established design specifications and performance indicators. After each designed functional module passes the functional simulation test, it is integrated into the entire graphics processor system for performance analysis and debugging.
[0024] In order to accurately detect the performance bottleneck of the graphics processor and perform design optimization at the early design stage of the graphics processor chip, that is, at the design stage of each functional module, this embodiment provides a graphics processor performance analysis method, as Figure 1 shown, including: S1, running test cases; S2, obtaining the workload signals of each functional module in the graphics processor; S3. Count the workload signals of each functional module and record the key time information of each workload signal; S4. Use the key time information to analyze and judge the actual performance of the functional module to confirm the bottleneck functional module; S5. Output a performance analysis report.
[0025] The graphics processor performance analysis method provided in this embodiment can quickly detect the change of data throughput rate of each functional module by sampling and analyzing the workload signals of each functional module in the graphics processor, and use the key time information to quickly confirm the bottleneck functional module, improve the efficiency of graphics processor performance analysis, and solve the problem of low efficiency of the existing graphics processor performance analysis method.
[0026] Specifically, in this embodiment, in step S1, run the test case.
[0027] In practical applications, multiple test cases can be designed according to the performance analysis requirements of the graphics processor, so that the performance of each functional module in the graphics processor can be comprehensively tested through different test cases.
[0028] The test case also includes a corresponding configuration file, which is set with the number of functional modules, the enable signal of the functional module, the number of counts to draw, the unit time, the invalid threshold, the definition of each clock, the output data line width of each functional module, the data line utilization rate in different working modes, the workload target value of each functional module under this test case, etc. The configuration file can be read in with a file reading function, and the corresponding values are assigned to relevant variables.
[0029] Preferably, in order to improve the simulation speed of performance analysis, in this embodiment, the typical settings of each functional module are not enabled. Therefore, when running the test case, it is necessary to use the configuration file to configure the enable signal of the relevant functional module to enable each used functional module.
[0030] Further, in this embodiment, in step S2, obtain the workload signals of each functional module in the graphics processor.
[0031] In practical applications, when running the test case, each functional module in the graphics processor will start corresponding work according to the settings of the test case, thereby generating workload signals. By sampling the workload signals of all functional modules, the workload signals of each functional module in the graphics processor can be obtained.
[0032] And, in this embodiment, in step S3, count the workload signals of each functional module and record the key time information of each workload signal.
[0033] In practical applications, when running test cases, the workload signals of each functional module are counted, and the key time information of each workload signal is recorded. Specifically, within the counting period, the workload signals are sampled and counted, where the counting period can be flexibly controlled according to the control signal and the configuration file and changed to one or more draws. And the key time information includes valid time points, invalid time points, valid durations, and invalid durations.
[0034] In this embodiment, the workload signal of each functional module is valid (high level) when the functional module is working at full load, indicating that the data throughput rate of the functional module is equal to the value defined by the design specification; the workload signal of each functional module is invalid (low level) when the functional module is not working at full load, indicating that the data throughput rate of the functional module is less than the value defined by the design specification, that is, there are bubbles in the pipeline of the functional module during operation.
[0035] Further, in this embodiment, in step S4, the method of using the key time information to analyze and judge the actual performance of the functional module to confirm the bottleneck functional module includes: S41-1, obtaining the first invalid time point of each workload signal; S41-2, sorting the workload signals according to the chronological order; S41-3, identifying the functional module corresponding to the workload signal ranked first as the bottleneck functional module.
[0036] In this way, by sorting the first invalid time points in the workload signals of each functional module under the current test case and identifying the functional module with the earliest invalid time point as the bottleneck functional module, the functional module with the bottleneck can be quickly located, which is beneficial for the design engineer to analyze whether the reason for the functional module to become a bottleneck is a hardware code design problem or an architecture design problem, and then the design can be modified according to the problem reason, thus accelerating the design and development speed of the graphics processor.
[0037] Preferably, considering that some functional modules may have a situation of non-full-load operation (some bubbles will be generated) when they start working, and this situation is unavoidable in design. Therefore, in order to improve the accuracy of confirming the bottleneck functional module, in this embodiment, the sliding window algorithm is used to filter these bubbles. Specifically, in step S4, the method of using the key time information to analyze and judge the actual performance of the functional module to confirm the bottleneck functional module further includes: S42-1, obtaining the time point when each workload signal first exceeds the preset invalid threshold; S42-2, sorting the workload signals according to the chronological order; S42-3. Identify the functional module corresponding to the workload signal ranked first as the bottleneck functional module.
[0038] In this way, by presetting an invalid threshold, sorting the time points of the first workload signal of each functional module that exceeds the preset invalid threshold under the current test case, and identifying the functional module with the earliest time point exceeding the preset invalid threshold as the bottleneck functional module, not only can the functional module with a bottleneck be quickly located, but also misjudgment of the bottleneck functional module can be avoided, effectively improving the performance analysis efficiency of the graphics processor.
[0039] Preferably, in order to more quickly and accurately confirm the bottleneck functional module, in this embodiment, in step S4, the method of analyzing and judging the actual performance of the functional module by using key time information to confirm the bottleneck functional module further includes: S43-1. Obtain the configuration file corresponding to the test case.
[0040] Specifically, in this embodiment, the configuration file includes the definitions of each clock, the bit width of the output data line of each module, the data line utilization rate in different working modes, etc.
[0041] S43-2. Calculate the performance target value of each functional module by using the configuration file.
[0042] In this embodiment, the performance target value refers to the number of clocks required for each functional module to calculate all the input data and output it.
[0043] Taking a simple pipeline model containing 3 functional modules as an example, as Figure 2 shown, the working clock of functional module A is Clock_A, and the frequency is f A ; the working clock of functional module B is Clock_B, and the frequency is f B ; the working clock of functional module C is Clock_C, and the frequency is f C . The data line width of the interface between functional module A and functional module B is X bits (excluding control lines), the data line width of the interface between functional module B and functional module C is Y bits (excluding control lines), and the data line width of the output of functional module C is Z bits (excluding control lines). For different working modes, the data lines between each functional module may not be fully utilized. Assuming that in the same working mode, the data line utilization rate between functional module A and functional module B is U%, the data line utilization rate between functional module B and functional module C is V%, and the data line utilization rate of the output of functional module C is W%.
[0044] The design goal of the architecture for the data throughput rate between functional modules is: f A ×X×U% ≤ fB X × Y × V% ≤ f C × Z × W%, so that it can ensure that the data is input from functional module A until it is output from functional module C, and the entire pipeline will not be blocked.
[0045] According to each test case and using the data throughput rate design goal, the performance target values of each functional module under each test case can be calculated. Assume that in a test case, Figure 2 the pipeline model shown needs to process a set of data with m rows and n columns, and the bit width of each data is p bits. Then the performance target values of functional module A, functional module B, and functional module C can be expressed as: Target_a = m × n × p / (X × U%) Target_b = m × n × p / (Y × V%) Target_c = m × n × p / (Z × W%) Of course, in other embodiments, those skilled in the art can obtain the calculation methods of the performance target values in other embodiments according to the calculation methods provided in this embodiment, and this application will not elaborate further.
[0046] S43-3. Calculate the performance score of the functional module according to the actual performance value and the performance target value of the functional module.
[0047] Specifically, in this embodiment, dividing the performance target value of the functional module by the actual performance value can obtain the performance score of the functional module under the current test case.
[0048] Since each functional module may have multiple working modes, it is necessary to perform performance analysis and simulation on each working mode of the functional module to obtain the performance scores of each functional module in each working mode.
[0049] S43-4. Compare the performance score calculated for each functional module with the preset expected score; if the performance score of a certain functional module is less than the preset expected score, it indicates that the performance of this functional module is unqualified. At this time, it is necessary to find the functional module corresponding to the workload signal at the earliest invalid time point and identify this functional module as the candidate bottleneck functional module.
[0050] As mentioned above, in order to improve the accuracy of identifying the bottleneck functional module, in this embodiment, the workload signals of the candidate bottleneck functional modules are further sorted according to the chronological order of the time points when they first exceed the preset invalid threshold, and the functional module corresponding to the workload signal at the earliest time point when it first exceeds the preset invalid threshold is found and identified as the bottleneck functional module.
[0051] Further, in this embodiment, in step S5, a performance analysis report is output. The performance analysis report may include the count values of the workload signals of each functional module, key time information, performance target values, performance scores, preset invalid thresholds, expected scores, sorted data obtained by sorting according to the chronological order, and the identified bottleneck functional modules, etc.
[0052] In practical applications, the content reflected in the performance analysis report and the display method of the content can be adjusted according to actual needs, and the present application does not limit this.
[0053] The graphics processor performance analysis method provided in this embodiment can be applied to the early design and debugging stage of each functional module of the graphics processor, quickly locate the functional module with performance bottlenecks, and thus solve the bottleneck problem by analyzing the reasons and modifying the corresponding RTL (Register Transfer Level) code or architecture, which speeds up the design and development speed of the graphics processor.
[0054] The graphics processor performance analysis method provided in this embodiment can also be applied to any stage from the overall debugging of the RTL code of the graphics processor to the chip tape-out of the graphics processor. After modifying the RTL code, the waveform can be directly run without collecting the waveform, and according to the comparison result between the performance score and the preset expected score, it can be quickly checked whether the modification of the RTL code affects the performance of the graphics processor. Since waveforms are not collected during this process, the speed of performance analysis can be greatly increased, and the design and development time of the graphics processor can be shortened.
[0055] The graphics processor performance analysis method provided in this embodiment can determine which period of waveform to collect for debugging the problems according to the output performance analysis report, thereby effectively reducing the storage space of the waveform and facilitating the simulation operation of larger test cases.
[0056] This embodiment also provides a graphics processor performance analysis system for implementing the above-mentioned graphics processor performance analysis method, as Figure 3 shown, the graphics processor performance analysis system includes: A recording module, configured to obtain the workload signals of each functional module in the graphics processor, count the workload signals of each functional module, and record the key time information of each workload signal; A processing module, configured to analyze and judge the actual performance of the functional module according to the data information statistically obtained by the recording module to confirm the bottleneck functional module; A reporting module, configured to output a performance analysis report.
[0057] Specifically, in this embodiment, since the recording module needs to record simulation signals, it must be integrated into the performance analysis simulation environment. For example, it is included in the testbench, and the design language used by the recording module is the same as the hardware description language (HDL) used for the graphics processor design.
[0058] For the processing module, which is used to calculate and process the data generated by the recording module, it can be integrated into the simulation environment or placed outside the simulation environment. When the processing module is placed outside the simulation environment, all the simulation data can be processed after the simulation ends. In this way, in the simulation environment, when a test case runs to completion, the next test case can be directly run without waiting for the data processing result of the processing module, thereby saving the simulation time for performance analysis and improving the efficiency of performance analysis.
[0059] In practical applications, when building the simulation environment, the recording module and the processing module can be integrated into the simulation environment. In this way, when performing the functional simulation of the graphics processor, the performance simulation environment can be debugged simultaneously. Thus, when the functional simulation of each functional module of the graphics processor ends, the performance analysis simulation can be directly started, effectively improving the overall efficiency of the graphics processor simulation test.
[0060] Further, in this embodiment, as Figure 4 shown, the recording module includes a data input unit, at least one signal sampling unit, and a recording output unit.
[0061] Among them, the input module is used to obtain the workload signals, clock signals, control signals, and recording configuration files of each functional module in the graphics processor, and send the relevant signals to the signal sampling unit. The recording configuration file obtained by the input module contains the functional modules and the number of functional modules required for the current test case, the corresponding enable signals, the number of counts to be drawn, the unit time, the invalid threshold, etc. The input module can use the function of reading files to read the recording configuration file and assign the corresponding values to relevant variables.
[0062] In practical applications, the signals of each functional module can be connected to the data input unit in the form of flylines. Taking Verilog HDL language as an example, the workload signal of functional module A in the graphics processor is workload_a, the clock is clock_a, and adding the hierarchy of functional module A in the entire simulation environment, the obtained signal names are: Module_A_hierarchy.workload_a, Module_A_hierarchy.clock_a. According to the flyline connection method, it can be expressed as: assignworkload_A = Module_A_hierarchy.workload_a assignclock_A = Module_A_hierarchy.clock_a Among them, workload_A and clock_A are signals of the data input unit.
[0063] Moreover, the signal sampling unit corresponds one-to-one with the functional modules of the graphics processor, so as to obtain the workload signals of each functional module in the graphics processor through the data input unit, count the workload signals of each functional module according to the clock signal, the control signal and the configuration file, and record the key time information of each workload signal.
[0064] Taking the signal of functional module A with a counting period of one drawing as an example, as Figure 5 shown, the first high level in the drawing control signal indicates the start of a drawing, and the second high level indicates the end of the drawing; Count_total represents the total working duration of functional module A, that is, the actual performance value of functional module A; Count_valid represents the actual effective working duration of functional module A, that is, the performance target value of functional module A; Count_invalid represents the duration of each invalidation of the workload signal of functional module A. Simulation time 1 is the time point when functional module A starts working in the current drawing; Simulation time 2 and simulation time 3 are the time points when each invalidation of the workload signal is recorded, where simulation time 2 is the time point when the workload signal of functional module A is first invalidated in the current drawing; Simulation time 4 is the time point when functional module A ends working in the current drawing.
[0065] When the signal sampling unit performs statistics, it records the value of Count_invalid after simulation time 2 and simulation time 2 ( Figure 5 which is 1 here) together; records the value of Count_invalid after simulation time 3 and simulation time 3 ( Figure 5 which is 8 here) together; thus obtaining the time point and duration of each invalidation of the workload signal of functional module A.
[0066] In addition, in this embodiment, the signal sampling unit also records the time point and corresponding duration when the workload signal exceeds the preset invalid threshold within a unit time. Assuming that the unit time is m clocks and the invalid threshold is n clocks, m cycle counters and m workload signal invalid time counters are used for counting. If the invalid time of the workload signal within m clocks is greater than or equal to n clocks, then record the time point when the workload signal is first invalidated within this period, and the total invalid duration within this period (counted in clock numbers).
[0067] Taking the case where the unit time is 3 clocks and the invalid threshold is 2 clocks as an example, as Figure 6 shown, Count_window1, Count_window2, and Count_window3 correspond to three loop counters, all of which count cyclically with a period of 3 clocks. The start times of these three loop counters differ by one clock cycle in sequence, so that any consecutive 3 clock cycles can be covered, realizing the function of the sliding window algorithm. Count_invalid1, Count_invalid2, and Count_invalid3 are three workload signal invalid duration counters, which count the invalid duration of the workload signal during the counting periods of Count_window1, Count_window2, and Count_window3 respectively.
[0068] Simulation time 1, simulation time 2, and simulation time 3 record the time points of each invalid (low level) of the workload signal. When Count_window2 = 2, Count_invalid2 = 1, and the workload signal is 0, that is, when the invalid duration of the workload signal is detected to be equal to the set invalid threshold 2 within window 1. Record the values of simulation time 3 and (Count_invalid2 + 1), which is 2, respectively, as the time point and invalid duration when the invalid threshold is exceeded for the first time. The same situation is also detected in window 2, which will not be elaborated in this application.
[0069] And, the recording and output unit is used to count the workload signals, count values of the workload signals, and key time information acquired by all the signal sampling units, and send the counted data information to the processing module.
[0070] In practical applications, all the data information acquired by the recording and output unit can also be output to a log file for easy storage and access of the data.
[0071] Furthermore, in this embodiment, as Figure 7 shown, the processing module includes a receiving unit, an operation unit, a reordering unit, and a result output unit.
[0072] Among them, the receiving unit is used to receive the data information and processing configuration file statistically calculated by the recording module. The processing configuration file here includes the definitions of each clock, the output data line widths of each functional module, the data line utilization rates of each functional module in different working modes, etc.
[0073] Moreover, the operation unit is used to calculate the performance target value and performance score of each functional module according to the received data information by using the processing configuration file. In this embodiment, the calculation of the performance target value can refer to the content of step S43-2 above, and the calculation of the performance score can refer to the content of step S43-3 above.
[0074] In a specific embodiment, for the calculation of the performance score, refer to Figure 5 , where Count_valid represents the performance target value of functional module A, and Count_total represents the actual performance value of functional module A. Let Perf_A be the performance score of functional module A in this simulation, and the calculation formula is as follows: Perf_A = (Count_valid_A / Count_total_A)×100% Ideally, the value of Perf_A is 100%. However, in practical applications, due to various limitations, such as considerations based on chip area and power consumption, it is necessary to sacrifice some performance, so the expected score of performance will also be less than 100%.
[0075] Let the expected performance score of functional module A be Perf_a. If Perf_A = Perf_a, it can be considered that the performance test of functional module A in this simulation is qualified; if Perf_A < Perf_a, it is considered that the performance test of functional module A in this simulation is unqualified.
[0076] In practical applications, the performance target values and performance scores of each functional module calculated by the operation unit can be output to a log file for easy storage and access of data.
[0077] Moreover, the reordering module is used to sort the workload signals according to the chronological order of the first invalid time points of each workload signal, and / or is used to sort the workload signals according to the chronological order of the time points when each workload signal first exceeds the preset invalid threshold, and / or is used to sort the workload signals according to the chronological order of each invalid time point of each workload signal. By sorting the invalid time points in chronological order through the reordering module, the time point when invalid first occurs can be quickly found, and then the corresponding bottleneck functional module can be quickly found.
[0078] Moreover, the result output unit is used to confirm the bottleneck functional module according to the sorting result of the reordering module and output it.
[0079] The graphics processor performance analysis system provided in this embodiment, based on the characteristics of simple control, emphasis on parallel computing, and large data throughput rate in the graphics processor design, monitors the change of data throughput rate of each functional module to find the functional module with performance bottleneck. Among them, the change of the workload signal of each functional module being valid and invalid reflects the change of the data throughput rate of this functional module. Therefore, if the workload signal of a functional module is invalid, it indicates that the data throughput rate of this functional module at the current moment has decreased compared to the design value. Moreover, the bubbles of the current workload signal generally conduct downstream to the downstream functional modules. Therefore, in performance analysis, quickly finding the functional module where the bubbles first occur, this functional module is very likely to be the functional module with performance bottleneck.
[0080] In addition, this embodiment also provides an electronic device, including a memory, a processor, and an executable program stored in the memory and capable of being run by the processor; when the processor runs the executable program, it executes the graphics processor performance analysis method as described above.
[0081] In addition, this embodiment also provides a computer storage medium, the computer storage medium stores an executable program; when the executable program is executed, it implements the graphics processor performance analysis method as described above.
[0082] It should be noted that the various embodiments in this specification are described in a progressive manner. The key point of each embodiment is to describe the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. In addition, the different parts among the various embodiments can also be combined and used with each other. The present invention makes no limitation on this.
[0083] The graphics processor performance analysis method, system, electronic device, and computer storage medium provided in this embodiment include: running a test case; obtaining the workload signals of each functional module in the graphics processor; counting the workload signals of each functional module and recording the key time information of each workload signal; using the key time information to analyze and judge the actual performance of the functional module to confirm the bottleneck functional module; outputting a performance analysis report. By sampling and analyzing the workload signals of each functional module in the graphics processor, the change of the data throughput rate of each functional module can be quickly detected using the key time information, so as to quickly confirm the bottleneck functional module, improve the efficiency of graphics processor performance analysis, and solve the problem of low efficiency of the existing graphics processor performance analysis method.
[0084] The above description is only a description of the preferred embodiments of the present invention, and does not limit the scope of the present invention in any way. Any changes and modifications made by those of ordinary skill in the art of the present invention based on the above disclosure shall fall within the protection scope of the claims.
Claims
1. A method for analyzing the performance of a graphics processor, characterized in that, Including: Running test cases; Obtaining workload signals of each functional module in the graphics processing unit; Counting the workload signals of each functional module and recording the key time information of each workload signal; Analyzing and judging the actual performance of the functional module by using the key time information to confirm the bottleneck functional module; Outputting a performance analysis report.
2. The method for analyzing the performance of a graphics processor according to claim 1, wherein The key time information includes valid time points, invalid time points, valid durations, and invalid durations.
3. The method for analyzing the performance of a graphics processor according to claim 1, wherein The method of analyzing and judging the actual performance of the functional module by using the key time information to confirm the bottleneck functional module includes: Obtaining the first invalid time point of each workload signal; Sorting the workload signals according to the chronological order; Identifying the functional module corresponding to the workload signal ranked first as the bottleneck functional module.
4. The graphic processor performance analysis method according to claim 1, characterized in that, The method of analyzing and judging the actual performance of the functional module by using the key time information to confirm the bottleneck functional module further includes: Obtaining the time point when each workload signal first exceeds the preset invalid threshold; Sorting the workload signals according to the chronological order; Identifying the functional module corresponding to the workload signal ranked first as the bottleneck functional module.
5. The method for analyzing the performance of a graphics processor according to claim 1, wherein The method of analyzing and judging the actual performance of the functional module by using the key time information to confirm the bottleneck functional module further includes: Obtaining the configuration file corresponding to the test case; Calculating the performance target value of each functional module by using the configuration file; Calculating the performance score of the functional module according to the actual performance value and the performance target value of the functional module; Comparing the performance score calculated for each functional module with the preset expected score; If the performance score of a certain functional module is less than the preset expected score, find the functional module corresponding to the workload signal where the earliest invalid time point is located, and identify this functional module as the candidate bottleneck functional module; Sort the workload signals of the candidate bottleneck functional module according to the chronological order of the time points when they first exceed the preset invalid threshold, find the functional module corresponding to the workload signal where the earliest time point when it first exceeds the preset invalid threshold is located, and identify this functional module as the bottleneck functional module.
6. A graphics processor performance analysis system for implementing the graphics processor performance analysis method according to any one of claims 1 to 5, characterized in that, The graphics processing unit performance analysis system includes: A recording module, configured to obtain the workload signals of each functional module in the graphics processing unit, count the workload signals of each functional module, and record the key time information of each workload signal; A processing module, configured to analyze and judge the actual performance of the functional module according to the data information statistically obtained by the recording module to confirm the bottleneck functional module; A reporting module, configured to output a performance analysis report.
7. The graphics processor performance analysis system according to claim 6, wherein The recording module is integrated in the performance analysis simulation environment, and the design language used is the same as the hardware description language used for the graphics processing unit design.
8. The graphics processor performance analysis system according to claim 6, characterized in that The recording module includes a data input unit, at least one signal sampling unit, and a recording output unit; the input module is used to obtain the workload signals, clock signals, control signals, and recording configuration files of each functional module in the graphics processor; the signal sampling units correspond one-to-one with the functional modules of the graphics processor to obtain the workload signals of each functional module in the graphics processor, count the workload signals of each functional module according to the clock signal, the control signal, and the configuration file, and record the key time information of each workload signal; the recording output unit is used to count the workload signals, the count values of the workload signals, and the key time information obtained by all the signal sampling units, and send the counted data information to the processing module.
9. The graphics processor performance analysis system according to claim 6, wherein The processing module includes a receiving unit, an arithmetic unit, a reordering unit, and a result output unit; the receiving unit is used to receive the data information and the processing configuration file counted by the recording module; the arithmetic unit is used to calculate the performance target value and performance score of each functional module according to the received data information by using the processing configuration file; the reordering module is used to sort the workload signals according to the chronological order of the first invalid time points of each workload signal, and / or, is used to sort the workload signals according to the chronological order of the time points when each workload signal first exceeds the preset invalid threshold, and / or, is used to sort the workload signals according to the chronological order of each invalid time point of each workload signal; the result output unit is used to confirm the bottleneck functional module according to the sorting result of the reordering module and output it.
10. An electronic device, characterized in that, It includes a memory, a processor, and an executable program stored in the memory and capable of being run by the processor; when the processor runs the executable program, it executes the graphics processor performance analysis method according to any one of claims 1 to 5.
11. A computer storage medium, characterized in that, The computer storage medium stores an executable program; when the executable program is executed, it implements the graphics processor performance analysis method according to any one of claims 1 to 5.