Attention kernel performance evaluation method and device, equipment and storage medium
By constructing an attention mechanism module for multiple operating environments and using Nsight Systems tools to track and collect data, the problem of lack of comprehensiveness and accuracy in the performance evaluation of attention mechanisms was solved, and accurate performance evaluation under different environments was achieved.
Patent Information
- Application Number
- CN202510888624.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-11-07
Smart Images

Figure CN120911513A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, and particularly relates to a method and device for evaluating performance of attention kernel based on Nsight Systems, equipment and storage medium. BACKGROUND
[0002] In the field of artificial intelligence today, attention mechanism (Attention) becomes the core part of accelerating calculation. In order to understand the running performance of the attention mechanism, the current main way is to manually check the running situation of the model on the graphics processing unit (GPU) by the developer, and to record the running time to evaluate, but such implementation is not comprehensive and lacks comparison, especially when running in different environments, comparison cannot be realized, resulting in inaccurate evaluation results. SUMMARY
[0003] Therefore, the purpose of the present disclosure is to provide a method and device for evaluating performance of attention kernel based on Nsight Systems, equipment and storage medium, to solve the problem that the current softmax exponential approximation calculation calls more floating-point multiply-add instructions, resulting in a large operation load of GPU.
[0004] In a first aspect, the present disclosure provides a method for evaluating performance of attention kernel based on Nsight Systems, comprising:
[0005] obtaining a target attention algorithm to be evaluated and at least one old version algorithm of the target attention algorithm;
[0006] based on the target attention algorithm and the at least one old version algorithm, constructing an attention mechanism module containing an exponential operation module and at least two running environments;
[0007] running the attention mechanism module in each running environment according to a preset running configuration, and tracking and collecting key running data of the attention mechanism module in the running process by using Nsight Systems;
[0008] taking the key running data of the at least one old version algorithm as a benchmark, analyzing the performance indicators of the target attention algorithm, and obtaining an evaluation result.
[0009] In a feasible implementation, the target attention algorithm is a floating-point multiply-add operation fusion algorithm, and the at least one old version algorithm includes a standard exponential function algorithm and a fast exponential approximation algorithm.
[0010] The method for evaluating performance of attention kernel based on Nsight Systems comprises the following steps:
[0011] The floating-point multiplication and addition operation fusion algorithm, the standard exponential function algorithm and the fast exponential approximation algorithm are embedded into an attention mechanism core kernel to obtain an attention mechanism module;
[0012] A second calculation precision is calculated based on a first calculation precision of the target attention algorithm and a preset error benchmark;
[0013] A hardware system is built according to the running requirements of the algorithms, and parameters of the hardware system are adjusted based on the first calculation precision and the second calculation precision to obtain two running environments.
[0014] In an implementable embodiment, the running of the attention mechanism module in each of the running environments according to a preset running configuration comprises:
[0015] According to the running requirements of the target attention algorithm and the running requirements at the time of evaluation, the running times of each algorithm, the input tensor dimensions of each algorithm and the performance indicators are determined;
[0016] Based on the input tensor dimensions of each algorithm, a tensor matrix of floating-point multiplication and addition operations input to the attention mechanism module is constructed;
[0017] Based on the running times and the tensor matrix, the attention mechanism module is run in each of the running environments.
[0018] In an implementable embodiment, the tracking and collection of key running data of the attention mechanism module in the running process by using Nsight Systems comprises:
[0019] Based on the performance indicators, the field identifiers of the key data are determined;
[0020] The field identifiers are tracked and collected by using Nsight Systems according to a preset sampling interval to obtain the key running data of each algorithm.
[0021] In an implementable embodiment, the performance indicators of the target attention algorithm are analyzed based on the key running data of the at least one old version algorithm to obtain an evaluation result, which comprises:
[0022] One of the at least one old version algorithm is selected as a standard algorithm, and key information in the key running data of each algorithm is extracted based on the field identifiers corresponding to the performance indicators;
[0023] Based on the extracted key information of each algorithm, performance evaluation among algorithms and performance evaluation among running environments are performed to obtain the evaluation result.
[0024] In an implementation, the performance evaluation between algorithms and the performance evaluation between running environments are performed based on the extracted key information of each algorithm, and evaluation results are obtained, including:
[0025] The execution parameters of each algorithm are calculated based on the extracted key information of each algorithm, wherein the execution parameters include average execution time, average resource utilization, and average calculation accuracy.
[0026] The standardized delay is calculated based on the average execution time of the standard algorithm and the average execution time of the target attention algorithm.
[0027] The performance evaluation between running environments is performed based on the standardized delay, the average resource utilization, and the average calculation accuracy, and evaluation results are obtained.
[0028] In an implementation, after the performance indicators of the target attention algorithm are analyzed based on the key running data of the at least one old version algorithm, the evaluation results are obtained, and the method further includes:
[0029] The evaluation results are converted into a performance comparison report containing each running environment, each calculation accuracy, and each resource utilization based on a visualization tool, and the comparison report includes a visualization chart.
[0030] In a second aspect, the embodiments of the present disclosure provide an attention kernel performance evaluation device based on Nsight Systems, and the device includes:
[0031] An acquisition module is configured to acquire a target attention algorithm to be evaluated and at least one old version algorithm of the target attention algorithm.
[0032] A construction module is configured to construct an attention mechanism module containing an exponential operation module and at least two running environments based on the target attention algorithm and the at least one old version algorithm.
[0033] A collection module is configured to run the attention mechanism module in each running environment according to a preset running configuration, and track and collect key running data of the attention mechanism module in a running process by using Nsight Systems.
[0034] An evaluation module is configured to analyze performance indicators of the target attention algorithm based on key running data of the at least one old version algorithm, and obtain evaluation results.
[0035] In a third aspect, the embodiments of the present disclosure provide an electronic device, including a processor and a memory, the memory storing machine executable instructions capable of being executed by the processor, and the processor executes the machine executable instructions to implement the Nsight Systems-based attention kernel performance evaluation method provided above.
[0036] In a fourth aspect, the embodiments of the present disclosure provide a computer readable storage medium, the computer readable storage medium storing computer executable instructions, and when the computer executable instructions are invoked and executed by a processor, the computer executable instructions cause the processor to implement the Nsight Systems-based attention kernel performance evaluation method provided above.
[0037] The embodiments of the present disclosure bring the following beneficial effects:
[0038] The Nsight Systems-based attention kernel performance evaluation method, device, equipment and storage medium provided above, obtain a target attention algorithm to be evaluated and at least one old version algorithm of the target attention algorithm; based on the target attention algorithm and the at least one old version algorithm, construct an attention mechanism module containing an exponential operation module and at least two running environments; run the attention mechanism module in each running environment according to a preset running configuration, and track and collect key running data of the attention mechanism module in the running process by using Nsight Systems; take the key running data of the at least one old version algorithm as a benchmark to analyze the performance index of the target attention algorithm, and obtain an evaluation result.
[0039] In the method, the target attention algorithm and the corresponding at least one old version algorithm are integrated into the attention mechanism module, and multiple running environments are set to execute the algorithm, and key running data is tracked and collected by using Nsight Systems, and one of them is taken as a benchmark to evaluate the target attention algorithm, so as to realize the running and comparative evaluation of multiple algorithms on one kernel, and solve the problem that the performance comparative evaluation of different running environments cannot be realized and the evaluation accuracy is low in the prior art.
[0040] Other features and advantages of the present disclosure will be set forth in the following description, and in part will become apparent from the description, or will be learned by practice of the present disclosure. The objects and other advantages of the present disclosure will be realized and achieved by the structures particularly pointed out in the specification, claims, and drawings.
[0041] In order to make the above objects, features and advantages of the present disclosure more apparent, the following preferred embodiments are specifically described with reference to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0042] In order to make the technical solutions in the specific embodiments of the present disclosure or the prior art clearer, the accompanying drawings needed in the specific embodiments or prior art description will be briefly introduced. Obviously, the accompanying drawings in the following description are some embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without any creative work on the premise of the accompanying drawings.
[0043] Figure 1 An embodiment schematic diagram of the attention kernel performance evaluation method based on Nsight Systems provided by the embodiments of the present disclosure;
[0044] Figure 2 Another embodiment schematic diagram of the attention kernel performance evaluation method based on Nsight Systems provided by the embodiments of the present disclosure;
[0045] Figure 3 A first embodiment schematic diagram of the attention kernel performance evaluation device based on Nsight Systems provided by the embodiments of the present disclosure;
[0046] Figure 4 A second embodiment schematic diagram of the attention kernel performance evaluation device based on Nsight Systems provided by the embodiments of the present disclosure;
[0047] Figure 5 A schematic diagram of an electronic device provided by the embodiments of the present disclosure. DETAILED DESCRIPTION
[0048] In order to make the technical solutions in the specific embodiments of the present disclosure or the prior art clearer, the accompanying drawings needed in the specific embodiments or prior art description will be briefly introduced. Obviously, the accompanying drawings in the following description are some embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without any creative work on the premise of the accompanying drawings.
[0049] The attention kernel performance evaluation method based on Nsight Systems in one of the embodiments of the present disclosure can run on a terminal device or a server. The terminal device can be a local terminal device. When the attention kernel performance evaluation method based on Nsight Systems runs on the server, the method can be implemented and executed based on a cloud interaction system, wherein the cloud interaction system includes a server and a client device.
[0050] For ease of understanding, the specific process of the present embodiment will be described below. Please refer to Figure 1In this embodiment, an embodiment of the attention kernel performance evaluation method based on Nsight Systems includes the following steps:
[0051] 101. Obtain a target attention algorithm to be evaluated and at least one old version algorithm of the target attention algorithm.
[0052] It should be noted that the target attention algorithm is the latest exponential algorithm applied to the GPU in the vehicle system, such as the fusion algorithm of FMA operation; the at least one old version algorithm refers to the old version algorithm in the latest exponential algorithm or an alternative algorithm with the same function.
[0053] 102. Based on the target attention algorithm and the at least one old version algorithm, construct an attention mechanism module containing an exponential operation module and at least two running environments.
[0054] Specifically, based on the algorithm logic of the attention mechanism, an exponential operation module is added in a deep learning model, and then the target attention algorithm and the at least one old version algorithm are integrated in the exponential operation module, thereby constructing the attention mechanism module. The running environment is constructed based on the hardware environment currently used by the target attention algorithm, and then a running environment with higher or lower accuracy is expanded based on the running environment of the target attention algorithm.
[0055] 103. Run the attention mechanism module in each running environment according to the preset running configuration, and use Nsight Systems to track and collect key running data of the attention mechanism module during running.
[0056] The running environment is scheduled by running instructions, and the corresponding algorithm in the attention mechanism module is started based on the algorithm identifier in the running instructions, and is run based on the input data to obtain running data. The tracking field of Nsight Systems is configured to track and extract the corresponding key running data from the obtained running data.
[0057] 104. Take the key running data of the at least one old version algorithm as a benchmark to analyze the performance indicators of the target attention algorithm and obtain an evaluation result.
[0058] Select one of the at least one old version algorithm as a standard using a random algorithm or a preset standard selection rule, and then select the key running data of the standard as a data standard. Specifically, the data standard needs to be analyzed according to the key running data of the selected standard to calculate the corresponding indicator standard, and then the key running data of the target attention algorithm is analyzed to compare the analyzed indicators with the corresponding indicator standard to obtain the evaluation result.
[0059] It should be noted that the comparison here includes comparison in the same running environment and comparison across running environments, and then the final performance index is calculated based on the results of the two comparisons using a weight coefficient to obtain the evaluation result.
[0060] By executing the above-mentioned scheme, by utilizing the target attention algorithm and at least one old version algorithm, an attention mechanism module containing an exponential operation module and at least two running environments are constructed, and then controlled to run to track and collect key running data of each algorithm and perform performance analysis. Since this method integrates multiple algorithms and multiple running environments, it realizes multi-environment test evaluation, and solves the problems of existing evaluation schemes that cannot realize performance comparison and evaluation across different running environments and have low evaluation accuracy.
[0061] Referring to Figure 2 In another embodiment of the attention kernel performance evaluation method based on Nsight Systems in the embodiment, a general process is used to evaluate the speed and resource usage of different algorithms when running on a GPU, and finally a clear comparison result is obtained. The method specifically includes the following steps:
[0062] 201, obtaining a target attention algorithm to be evaluated and at least one old version algorithm of the target attention algorithm;
[0063] 202, based on the target attention algorithm and the at least one old version algorithm, constructing an attention mechanism module containing an exponential operation module and at least two running environments;
[0064] In the embodiment, the target attention algorithm is a floating-point multiplication and addition operation fusion algorithm, and the at least one old version algorithm includes a standard exponential function algorithm and a fast exponential approximation algorithm; based on this, the steps of constructing the attention mechanism module and the at least two running environments include:
[0065] First, the floating-point multiplication and addition operation fusion algorithm, the standard exponential function algorithm and the fast exponential approximation algorithm are embedded into the attention mechanism core kernel to obtain an attention mechanism module.
[0066] It can be understood that the floating-point multiplication and addition operation fusion algorithm is a self-defined exponential approximation function MagicExp, the at least one old version algorithm is a standard exponential function StandardExp and a fast approximation exponential function FastExp, an exponential operation module containing multiple exponential function implementations is constructed, the exponential functions include, for example, a standard exponential function StandardExp, a fast approximation exponential function FastExp and a self-defined exponential approximation function MagicExp, and the exponential functions are integrated into a kernel module of the attention mechanism to obtain an attention mechanism module.
[0067] Then, a second calculation precision is calculated based on a first calculation precision of the target attention algorithm and a preset error benchmark; a hardware system is built according to a running requirement of the algorithm, and parameters of the hardware system are adjusted based on the first calculation precision and the second calculation precision, so as to obtain two running environments.
[0068] It should be noted that the running environment includes hardware environment configuration and software environment configuration.
[0069] The hardware environment configuration includes two, which are a high-performance desktop GPU: NVIDIA RTX 3080 and an embedded AI platform: NVIDIA Jetson Orin.
[0070] The software environment configuration mainly includes calculation precision configuration, which also includes two, which are single-precision floating point (FP32) and half-precision floating point (FP16).
[0071] In addition, in order to quickly collect key running data, a tool compatible with the running environment configuration is also needed, which specifically includes using Nsight Systems as the main analysis tool and using the CUDA Profiling tool package for precision control and kernel function parameter debugging.
[0072] 203. According to the running requirements of the target attention algorithm and the running requirements during evaluation, the running times of each algorithm, the input tensor dimensions of each algorithm, and the performance indicators are determined.
[0073] 204. Based on the input tensor dimensions of each algorithm, a tensor matrix of floating-point multiplication and addition operations from the input to the attention mechanism module is constructed.
[0074] 205. Based on the running times and the tensor matrix, the attention mechanism module is run in each running environment.
[0075] In an embodiment, each algorithm is set to run 10 rounds; and a fixed input is run in each round, such as an input tensor matrix Q = [B, H, L, D], wherein B is the batch size, H is the attention head number, L is the sequence length, and D is the embedding dimension, for example, [16, 8, 128, 64]. The dimension elements in the matrix are converted using the same batch size, sequence length, and head size; after conversion, dummy writing is used for output to obtain a bright matrix. This configuration mode unifies the input data and batch processing configuration of each algorithm running, so as to ensure the consistency and comparability of the test data, and to avoid IO interference performance.
[0076] 206. Based on the performance indicators, the field identifiers of the key data are determined.
[0077] The field identifies CUDA kernel execution time, TensorCore usage rate, and SM activity rate, etc.
[0078] 207. Track and collect the field identified according to a preset sampling interval by using Nsight Systems to obtain key running data of each algorithm.
[0079] Nsight Systems is called to collect the following information at an interval of 100 microseconds: kernel function execution time and TensorCore active time period and usage rate.
[0080] In another embodiment, no less than 10 rounds of repeated tests are performed in the running of each algorithm, the input tensor is fixed in each round of running, and write-out operations are avoided to reduce IO interference and improve test stability; the mean value of the key indicators in each round of test is recorded, including kernel execution time and TensorCore average utilization.
[0081] Specifically, the NVTX label is opened for recording, and Nsight Systems can accurately track the start time, duration, and end time of each exponential operation kernel; and the TensorCore usage in the GPU is counted, including the following indicators:
[0082] TensorCore Activity;
[0083] Tensor Utilization (%Active);
[0084] Memory Throughput (to assist in determining the bottleneck position).
[0085] In another embodiment, after the key running data is collected, the collected data is also standardized, the average execution time of the old version algorithm under the FP32 condition is taken as the benchmark, and the time of the remaining algorithms is converted into standardized delay values in proportion; a structured data table is formed to record the algorithm name, platform, precision, running time, TensorCore utilization rate, and standardized delay.
[0086] 208. Taking the key running data of at least one old version algorithm as a benchmark, analyze the performance indicators of the target attention algorithm to obtain an evaluation result.
[0087] In this step, one of the at least one old version algorithm is selected as a standard algorithm, and the key information in the key running data of each algorithm is extracted based on the field identified corresponding to the performance indicators; performance evaluation between algorithms and performance evaluation between running environments are performed based on the extracted key information of each algorithm to obtain an evaluation result.
[0088] The performance evaluation between algorithms and the performance evaluation between running environments are performed based on the extracted key information of each algorithm, and evaluation results are obtained, including: calculating the execution parameters of each algorithm based on the extracted key information of each algorithm, wherein the execution parameters include: average execution time, average resource utilization and average calculation accuracy; calculating the standardized delay based on the average execution time of the standard algorithm and the average execution time of the target attention algorithm; performing performance evaluation between running environments based on the standardized delay, the average resource utilization and the average calculation accuracy, and obtaining evaluation results.
[0089] In actual application, the average execution time of each kernel is obtained by averaging the results of 10 runs; the standardized kernel delay is calculated: the FP32 execution time of StandardExp is taken as the standard "1", and the execution times of the remaining algorithms are normalized with this value; the average TensorCore utilization rate corresponding to each algorithm is recorded as a measure of hardware resource use efficiency.
[0090] Further, in order to facilitate the viewing of the evaluation results, after the performance indicators of the target attention algorithm are analyzed based on the key running data of the at least one old version algorithm to obtain the evaluation results, it further includes: converting the evaluation results into a performance comparison report containing each running environment, each calculation accuracy and each resource utilization based on a visualization tool, and the comparison report includes a visualization chart.
[0091] Specifically, the execution time trend chart and the TensorCore utilization rate horizontal comparison chart are drawn based on the evaluation results; the visualization chart can be a column chart, a line chart or a heat map, and is output as a PDF or JSON format file, etc.
[0092] The report contains the following contents: test platform and environment configuration description; comparison of average execution time of each algorithm, standardized delay chart, TensorCore utilization rate trend chart, adaptability conclusion of algorithms under different hardware and accuracy, recommended scene (for example, Jetson Orin recommends using MagicExp FP16 version).
[0093] Further, for the report output format using file, it can be in the form of PDF or PPT display; when using the chart, the following formats are recommended: bar chart comparison of execution time, line chart or heat map comparison of TensorCore utilization rate and table comparison of algorithm difference summary.
[0094] The attention kernel performance evaluation method based on Nsight Systems solves the problems that the existing evaluation scheme cannot realize performance comparison and evaluation of different running environments and has low evaluation accuracy.
[0095] Corresponding to the method embodiment, referring to Figure 3 The attention kernel performance evaluation device based on Nsight Systems includes:
[0096] The obtaining module 310 is configured to obtain a target attention algorithm to be evaluated and at least one old version algorithm of the target attention algorithm.
[0097] The constructing module 320 is configured to construct an attention mechanism module including an exponential operation module and at least two running environments based on the target attention algorithm and the at least one old version algorithm.
[0098] The collecting module 330 is configured to run the attention mechanism module in each running environment according to a preset running configuration, and track and collect key running data of the attention mechanism module in a running process by using Nsight Systems.
[0099] The evaluation module 340 is configured to take the key running data of the at least one old version algorithm as a benchmark, analyze a performance index of the target attention algorithm, and obtain an evaluation result.
[0100] The attention kernel performance evaluation device based on Nsight Systems solves the problems that the existing evaluation scheme cannot realize performance comparison and evaluation of different running environments and has low evaluation accuracy.
[0101] For details, refer to Figure 4 Another embodiment of the attention kernel performance evaluation device based on Nsight Systems includes:
[0102] The acquisition module 310 is configured to acquire a target attention algorithm to be evaluated and at least one old version algorithm of the target attention algorithm.
[0103] The construction module 320 is configured to construct an attention mechanism module including an exponential operation module and at least two running environments based on the target attention algorithm and the at least one old version algorithm.
[0104] The collection module 330 is configured to run the attention mechanism module in each of the running environments according to a preset running configuration, and track and collect key running data of the attention mechanism module in a running process by using Nsight Systems tracking.
[0105] The evaluation module 340 is configured to analyze a performance index of the target attention algorithm based on the key running data of the at least one old version algorithm, and obtain an evaluation result.
[0106] Optionally, the target attention algorithm is a floating-point multiply-add operation fusion algorithm, and the at least one old version algorithm includes a standard exponential function algorithm and a fast exponential approximation algorithm; and the construction module 320 includes:
[0107] The embedding unit 321 is configured to embed the floating-point multiply-add operation fusion algorithm, the standard exponential function algorithm and the fast exponential approximation algorithm into an attention mechanism core kernel to obtain an attention mechanism module.
[0108] The calculation unit 322 is configured to calculate a second calculation precision based on a first calculation precision of the target attention algorithm and a preset error benchmark, and to build a hardware system according to a running requirement of an algorithm, and adjust parameters of the hardware system based on the first calculation precision and the second calculation precision, to obtain two running environments.
[0109] Optionally, the collection module 330 includes:
[0110] The determination unit 331 is configured to determine a running time of each algorithm, an input tensor dimension of each algorithm and a performance index according to a running requirement of the target attention algorithm and a running requirement at the time of evaluation.
[0111] The construction unit 332 is configured to construct a tensor matrix of a floating-point multiply-add operation input to the attention mechanism module based on the input tensor dimension of each algorithm.
[0112] The running unit 333 is configured to run the attention mechanism module in each of the running environments based on the running time and the tensor matrix.
[0113] Optionally, the collection module 330 further includes a collection unit 334, which is configured to:
[0114] determine a field identifier of the key data based on the performance indicator;
[0115] track and collect the field identifier according to a preset sampling interval by using Nsight Systems to obtain key running data of each algorithm.
[0116] Optionally, the evaluation module 340 includes:
[0117] an extraction unit 341 configured to select one of the at least one old version algorithm as a standard algorithm, and extract key information in the key running data of each algorithm based on the field identifier corresponding to the performance indicator;
[0118] an evaluation unit 342 configured to perform performance evaluation between algorithms and performance evaluation between running environments based on the extracted key information of each algorithm to obtain an evaluation result.
[0119] Optionally, the evaluation unit 342 is specifically configured to:
[0120] calculate execution parameters of each algorithm based on the extracted key information of each algorithm, wherein the execution parameters include average execution time, average resource utilization, and average calculation accuracy;
[0121] calculate a standardized delay based on the average execution time of the standard algorithm and the average execution time of the target attention algorithm;
[0122] perform performance evaluation between running environments based on the standardized delay, the average resource utilization, and the average calculation accuracy to obtain the evaluation result.
[0123] Optionally, the apparatus further includes a conversion module 350 configured to:
[0124] convert the evaluation result into a performance comparison report including each running environment, each calculation accuracy, and each resource utilization based on a visualization tool, wherein the comparison report includes a visualization chart.
[0125] The embodiment also provides an electronic device including a processor and a memory, the memory storing machine executable instructions capable of being executed by the processor, and the processor executes the machine executable instructions to implement the Nsight Systems-based attention kernel performance evaluation method provided by the above embodiment. The electronic device can be a server or a terminal device.
[0126] Referring to Figure 5As shown, the electronic device includes a processor 500 and a memory 501 storing machine executable instructions executable by the processor 500 to implement the method for evaluating performance of an attention kernel based on Nsight Systems according to the above embodiments.
[0127] Further, Figure 5 As shown, the electronic device further includes a bus 502 and a communication interface 503, and the processor 500, the communication interface 503 and the memory 501 are connected through the bus 502.
[0128] The memory 501 can include a high-speed random access memory (RAM) and can also include a non-volatile memory such as at least one disk memory. The communication between the system network element and at least one other network element is realized through at least one communication interface 503 (which can be wired or wireless), and the Internet, a wide area network, a local area network, a metropolitan area network, etc. can be used. The bus 502 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one bidirectional arrow is used in the figure, but it does not mean that there is only one bus or only one type of bus.
[0129] The processor 500 can be an integrated circuit chip having a processing capability of signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of hardware in the processor 500 or the instruction in the form of software. The processor 500 described above can be a general processor, including a central processing unit (CPU), a network processor (NP), etc.; can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. Each method, step and logic block diagram disclosed in the embodiment can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiment can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read only memory, a programmable read only memory or an electrically erasable programmable memory, a register or other mature storage medium in the art. The storage medium is located in the memory 501, and the processor 500 reads the information in the memory 501, and combines the hardware to complete the steps of the Nsight Systems based attention kernel performance evaluation method provided in the above embodiment.
[0130] The embodiment also provides a computer readable storage medium, which stores computer executable instructions. When the computer executable instructions are called and executed by a processor, the computer executable instructions cause the processor to implement the Nsight Systems based attention kernel performance evaluation method provided in the above embodiment.
[0131] The Nsight Systems based attention kernel performance evaluation method, device, electronic equipment and computer program product of the storage medium provided in the embodiment include a computer readable storage medium storing program codes. The instructions included in the program codes can be used to execute the method described in the foregoing method embodiment. The specific implementation can be referred to the method embodiment, and will not be described here.
[0132] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system and device described above can refer to the corresponding process in the foregoing method embodiment, and will not be described here.
[0133] In addition, in the description of the embodiments, unless specifically defined and limited, the terms "mounting", "connected", "connection" should be interpreted broadly, for example, can be fixed connection, can also be detachable connection, or integrally connected; can be mechanical connection, can also be electrical connection; can be directly connected, can also be indirectly connected through an intermediate medium, can be internal communication of two elements. For those skilled in the art, the specific meaning of the above terms in the present disclosure can be understood according to the specific circumstances.
[0134] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present disclosure or the part of the prior art that essentially contributes to the present disclosure or the part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the embodiments of the present disclosure. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0135] In the description of the present disclosure, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present disclosure and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present disclosure. In addition, the terms "first", "second", "third" are only for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0136] Finally, it should be noted that: the above embodiments are only specific embodiments of the present disclosure, used to illustrate the technical solutions of the present disclosure, and are not limitations thereof, and the protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can make modifications or easily think of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed by the present disclosure, or make equivalent replacements to some of the technical features; and these modifications, changes or replacements do not cause the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments, and should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A method for evaluating performance of an attention kernel based on Nsight Systems, characterized in that, The method comprises: obtaining a target attention algorithm to be evaluated and at least one old version algorithm of the target attention algorithm; based on the target attention algorithm and the at least one old version algorithm, constructing an attention mechanism module containing an exponential operation module and at least two running environments; running the attention mechanism module in each running environment according to a preset running configuration, and tracking and collecting key running data of the attention mechanism module in the running process by using Nsight Systems tracking; taking the key running data of the at least one old version algorithm as a benchmark, analyzing the performance indicators of the target attention algorithm, and obtaining an evaluation result.
2. The method of claim 1, wherein, The target attention algorithm is a floating-point multiply-add operation fusion algorithm, and the at least one old version algorithm includes a standard exponential function algorithm and a fast exponential approximation algorithm; The method comprises: embedding the floating-point multiply-add operation fusion algorithm, the standard exponential function algorithm and the fast exponential approximation algorithm into an attention mechanism core kernel to obtain an attention mechanism module; calculating a second calculation accuracy based on a first calculation accuracy of the target attention algorithm and a preset error benchmark; building a hardware system according to the running requirements of the algorithm, and adjusting the parameters of the hardware system based on the first calculation accuracy and the second calculation accuracy to obtain two running environments.
3. The method of claim 1, wherein, The method comprises: determining the running times of each algorithm, the input tensor dimensions of each algorithm and the performance indicators according to the running requirements of the target attention algorithm and the running requirements at the time of evaluation; based on the input tensor dimensions of each algorithm, constructing a tensor matrix of floating-point multiply-add operations input to the attention mechanism module; based on the running times and the tensor matrix, running the attention mechanism module in each running environment.
4. The method of claim 3, wherein, The method comprises: determining the field identification of key data based on the performance indicators; tracking and collecting the field identification according to a preset sampling interval by using Nsight Systems to obtain the key running data of each algorithm.
5. The method of claim 1, wherein, The method comprises: selecting one of the at least one old version algorithm as a standard algorithm, and extracting key information in the key running data of each algorithm based on the field identification corresponding to the performance indicators; based on the extracted key information of each algorithm, performing performance evaluation between algorithms and between running environments to obtain an evaluation result.
6. The method of claim 5, wherein the Nsight Systems-based attention kernel performance evaluation method is based on a method of claim 1. The method comprises: The execution parameters of each algorithm are calculated based on the extracted key information of each algorithm, wherein the execution parameters include average execution time, average resource utilization, and average calculation accuracy; The standardized delay is calculated based on the average execution time of the standard algorithm and the average execution time of the target attention algorithm; The performance evaluation between the running environments is performed based on the standardized delay, the average resource utilization, and the average calculation accuracy, and an evaluation result is obtained.
7. The method of claim 1-6, wherein, After the performance indicators of the target attention algorithm are analyzed based on the key running data of the at least one old version algorithm, an evaluation result is obtained, the method further includes: The evaluation result is converted into a performance comparison report containing each running environment, each calculation accuracy, and each resource utilization based on a visualization tool, and the comparison report includes a visualization chart.
8. An Nsight Systems based attention kernel performance evaluation apparatus, characterized by, The device includes: An acquisition module configured to acquire a target attention algorithm to be evaluated and at least one old version algorithm of the target attention algorithm; A construction module configured to construct an attention mechanism module including an exponential operation module and at least two running environments based on the target attention algorithm and the at least one old version algorithm; A collection module configured to run the attention mechanism module in each running environment according to a preset running configuration, and track and collect key running data of the attention mechanism module in a running process by using Nsight Systems tracking; An evaluation module configured to analyze performance indicators of the target attention algorithm based on key running data of the at least one old version algorithm, and obtain an evaluation result.
9. An electronic device, comprising: The processor executes the machine executable instructions to implement the Nsight Systems-based attention kernel performance evaluation method of any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer executable instructions, and when the computer executable instructions are called and executed by the processor, the computer executable instructions cause the processor to implement the Nsight Systems-based attention kernel performance evaluation method of any one of claims 1-7.