NPU performance optimization method and system
By adopting dynamic bandwidth constraints in NPU performance optimization, splitting the calculation graph, sampling the actual bandwidth and verifying the performance differences, the problem of NPU performance degradation in the prior art is solved, and more efficient bandwidth utilization and performance optimization are achieved.
Patent Information
- Application Number
- CN202510311528.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art adopts static bandwidth constraints when deploying AI applications on NPUs, failing to fully utilize bandwidth resources on SoCs, resulting in a degradation of NPU performance and failing to effectively improve NPU performance when available bandwidth increases.
Dynamic bandwidth constraints are used to divide the computational graph into multiple subgraphs through the compilation stage, and performance evaluation is performed in combination with the specified bandwidth constraints to generate NPU execution files; in the bandwidth sampling stage, the sampling obtains actual bandwidth constraints; in the performance verification stage, the difference between estimated performance and operating performance is verified, and the bandwidth constraints are adjusted until the difference is less than or equal to the threshold.
Make full use of bandwidth resources on the SoC, optimize NPU performance, improve NPU performance, and ensure that the compilation effect and generation efficiency of AI applications are improved without affecting other computing devices on the SoC.
Smart Images

Figure CN120106160A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and system for optimizing NPU performance. Background Art
[0002] NPU (Neural Processing Unit) is a processor designed specifically for artificial intelligence (AI) and machine learning tasks, and is particularly good at accelerating neural network calculations, such as neural network reasoning, image recognition, speech recognition, and natural language processing. NPU provides higher computing efficiency with lower power consumption through simplified control flow and optimized hardware structure, and is particularly suitable for processing application scenarios with high real-time requirements such as vision, voice, and autonomous driving.
[0003] Against the backdrop of the rapid development and increasing application of AI technology, NPU plays a vital role in the end-side SoC (System on Chip, which simply means integrating CPU, GPU, NPU, memory, wifi chip, Bluetooth chip, etc. into a single chip called SoC). It usually works with other computing units such as CPU and GPU to form a heterogeneous computing platform to balance performance, power consumption and cost. In SoC, resource constraints are common, especially on cost-sensitive end-side devices, which limits the hardware resources that SoC can use, and memory bandwidth is usually much smaller than that of servers or cloud devices. How to coordinate with many different types of computing resources to make good use of limited bandwidth is a core issue in NPU performance optimization.
[0004] Currently, static bandwidth constraints are used when deploying AI applications on NPUs, that is, it is assumed that the available bandwidth resources are constant during the entire operation period. It does not take into account that higher-priority applications on the SoC preempt the bandwidth, resulting in a reduction in the bandwidth provided to the NPU and a degradation of the NPU performance. It also does not consider the opportunity for NPU performance improvement when the available bandwidth increases. Summary of the invention
[0005] The purpose of the present invention is to provide a method and system for optimizing NPU performance, which uses dynamic bandwidth constraints to replace the static bandwidth constraints of the prior art, fully utilizes the bandwidth resources on the SoC, optimizes the NPU performance, and improves the performance of the NPU without affecting other computing devices on the SoC.
[0006] In order to achieve the above object, the present invention provides a NPU performance optimization method, which includes a compilation stage, a bandwidth sampling stage and a performance verification stage;
[0007] The compilation stage includes: dividing the computation graph into multiple subgraphs, completing the performance evaluation of the subgraphs in combination with the specified bandwidth constraints, obtaining the estimated performance of the subgraphs, and generating NPU execution files;
[0008] The bandwidth sampling stage includes: running an execution file, the execution file includes the NPU execution file, sampling to obtain an actual bandwidth constraint;
[0009] The performance verification phase includes: obtaining the operating performance of the NPU in the bandwidth sampling phase, verifying the difference between the estimated performance and the operating performance, if the difference is less than or equal to a threshold, the optimization is completed, if the difference is greater than the threshold, the actual bandwidth constraint is used as the specified bandwidth constraint to re-compile the phase, bandwidth sampling phase and performance verification phase until the difference is less than or equal to the threshold.
[0010] Optionally, the compilation stage includes a segmentation stage, a performance evaluation stage and a code generation stage, wherein the computation graph is segmented in the segmentation stage; the performance evaluation is completed in the performance evaluation stage to obtain the estimated performance; an NPU execution file is generated in the code generation stage; the segmentation stage and the performance evaluation stage are iteratively executed until the segmentation of the entire computation graph is completed, and then the code generation stage is executed.
[0011] Optionally, in the performance evaluation stage, the optimal subgraph is selected according to the results of the performance evaluation, and it is determined whether the end node of the optimal subgraph is consistent with the end node of the computational graph. If not, the end node of the subgraph is used as the starting point of the next subgraph to execute the splitting stage. If they are consistent, it indicates that the splitting of the entire computational graph is completed, and the code generation stage begins.
[0012] Optionally, the step of completing the performance evaluation of the subgraph includes: obtaining available bandwidth at each time granularity in combination with the specified bandwidth constraint and the start time of the subgraph, and performing performance evaluation on the subgraph in combination with the available bandwidth.
[0013] Optionally, the execution file also includes execution files of other computing resources in addition to the NPU execution file, and the execution files of other computing resources can be customized and generated according to actual SoC computing requirements.
[0014] Optionally, the actual bandwidth constraint in the bandwidth sampling phase is updated as follows: the total bandwidth on the SoC is subtracted from the bandwidth used by the other computing resources sampled at the specified granularity when the SoC is actually running.
[0015] Optionally, the NPU execution file is generated according to the segmentation information and the operator library during the compilation stage.
[0016] Optionally, when the compilation phase is executed for the first time, the specified bandwidth constraint is a preset value, and when the compilation phase is executed subsequently, the actual bandwidth constraint is used as the specified bandwidth constraint.
[0017] Based on another aspect of the present invention, the present invention also provides an NPU performance optimization system, including: a compilation module, a bandwidth sampling module and a performance verification module;
[0018] The compiling module is used to divide the computation graph into multiple subgraphs, complete the performance evaluation of the subgraphs in combination with the specified bandwidth constraints, obtain the estimated performance of the subgraphs, and generate NPU execution files;
[0019] The bandwidth sampling module is used to run an execution file, the execution file includes the NPU execution file, and sample to obtain the actual bandwidth constraint;
[0020] The performance verification module is used to obtain the operating performance of the NPU in the bandwidth sampling stage, verify the difference between the estimated performance and the operating performance, if the difference is less than or equal to a threshold, the optimization is completed, if the difference is greater than the threshold, the actual bandwidth constraint is used as the specified bandwidth constraint to re-compile the stage, bandwidth sampling stage and performance verification stage until the difference is less than or equal to the threshold.
[0021] Optionally, the estimated performance includes IO performance and computing performance.
[0022] As configured above, the actual bandwidth constraint is sampled in the bandwidth sampling stage. In the performance verification stage, if the difference between the estimated performance and the running performance is greater than the threshold, the actual bandwidth constraint is used as the specified bandwidth constraint to re-compile the stage, bandwidth sampling stage and performance verification stage until the difference is less than or equal to the threshold, that is, except for the initial compilation stage, the subsequent compilation stages are all based on the estimated performance obtained according to the actual bandwidth constraint. In summary, the present invention uses dynamic bandwidth constraints to replace the static bandwidth constraints of the prior art, fully utilizes the bandwidth resources on the SoC, optimizes the NPU performance, can optimize the AI application compilation effect according to the actual bandwidth availability, generate a more efficiency-friendly version, and improve the performance of the NPU without affecting other computing devices on the SoC. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Those skilled in the art should understand that the drawings are provided for a better understanding of the present invention and do not constitute any limitation on the scope of the present invention.
[0024] Figure 1 A schematic diagram of a method for optimizing NPU performance according to an embodiment of the present invention;
[0025] Figure 2 FIG. 4 is a schematic diagram of an NPU performance optimization system according to an embodiment of the present invention. DETAILED DESCRIPTION
[0026] In this document, unless otherwise specified, the terms "upper", "lower", "left", "right", "inside", "outside", "front", "back", "top", "bottom", etc. are used to indicate directions or positional relationships based on the drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction and operation. Therefore, they cannot be understood as limiting the present invention.
[0027] The specific implementation of the present invention will be described in more detail below in conjunction with the schematic diagram. The advantages and features of the present invention will become clearer based on the following description. It should be noted that the drawings are all in a very simplified form and are not in exact proportions, and are only used to facilitate and clearly assist in explaining the purpose of the embodiments of the present invention.
[0028] Figure 1 is a schematic diagram of an NPU performance optimization method according to an embodiment of the present invention, please refer to Figure 1 An embodiment of the present invention provides a NPU performance optimization method, including a compilation stage, a bandwidth sampling stage and a performance verification stage.
[0029] The compilation phase includes: dividing the computation graph into multiple subgraphs, completing the performance evaluation of the subgraphs in combination with the specified bandwidth constraints, obtaining the estimated performance of the subgraphs, and generating the NPU execution file. The computation graph of this embodiment takes the AI application computation graph as an example.
[0030] Furthermore, the compilation phase includes a segmentation phase, a performance evaluation phase, and a code generation phase. In the segmentation phase, the computational graph is segmented; in the performance evaluation phase, the performance evaluation is completed to obtain the estimated performance; in the code generation phase, the NPU execution file is generated; the segmentation phase and the performance evaluation phase are iteratively executed until the segmentation of the entire computational graph is completed, and then the code generation phase is executed. In the performance evaluation phase, the optimal subgraph is selected according to the results of the performance evaluation, and it is determined whether the end node of the optimal subgraph is consistent with the end node of the computational graph. If not, the end node of the subgraph is used as the starting point of the next subgraph to execute the segmentation phase to perform the next segmentation, and the performance evaluation phase is also performed after the segmentation phase, and this cycle is repeated until the end node of the optimal subgraph is consistent with the end node of the computational graph. If they are consistent, it means that the segmentation of the entire computational graph is completed, and the code generation phase begins.
[0031] The step of completing the performance evaluation of the subgraph includes: obtaining the available bandwidth at each time granularity in combination with the specified bandwidth constraint and the start time of each subgraph, and performing performance evaluation on the subgraph in combination with the available bandwidth.
[0032] Specifically, in the segmentation phase, the AI application calculation graph is segmented into different subgraphs according to specified rules, and the start time of each subgraph is recorded. The subgraphs can contain the same nodes. In the performance evaluation phase, the specified bandwidth constraints and the start time of each subgraph are combined to obtain the accurate available bandwidth at each time granularity, complete the performance evaluation of the subgraph, obtain the estimated performance of the subgraph, and select the subgraph with the best performance. For example, the estimated performance may include but is not limited to IO performance (the time and efficiency of NPU data access through DRAM), computing performance, and overall performance, etc., which are determined by the user. It can be understood that any performance involved in this embodiment is represented by various performance indicators (specific values, such as time and efficiency).
[0033] In order to obtain more accurate subgraph performance evaluation, the parallel characteristics of the NPU hardware and the execution model can be combined to calculate the start time of each IO behavior in the corresponding subgraph, so as to determine the available bandwidth at each time granularity in the IO process and evaluate the performance in segments. After completing the performance evaluation of all subgraphs, the optimal subgraph can be selected according to customized rules, such as the highest MAC utilization (multiply-accumulate operations, which characterizes the efficiency of hardware computing units) and the lowest proportion of redundant data. The end node of the optimal subgraph is used as the starting point of the new subgraph, and the estimated overall running time of the optimal subgraph is updated to the time cache. The IO performance, computing performance, and overall performance of the optimal subgraph are recorded.
[0034] Furthermore, during the compilation phase, an NPU execution file is generated based on the segmentation information, operator library, and auxiliary components. That is, during the code generation phase of the compilation phase, the code to be executed on the corresponding NPU is generated based on all the optimal subgraph information, the built-in operator library (such as the built-in operator library in the SDK that comes with the NPU), and auxiliary components.
[0035] The bandwidth sampling phase includes: running the execution file, the execution file includes the NPU execution file, and sampling to obtain the actual bandwidth constraint. Furthermore, the execution file also includes the execution files of other computing resources in addition to the NPU execution file, and the execution files of other computing resources can be customized and generated according to the actual SoC computing requirements. Preferably, the actual bandwidth constraint in the bandwidth sampling phase is updated as follows: the total bandwidth on the SoC is subtracted from the bandwidth used by other computing resources when the SoC is actually running.
[0036] Specifically, the NPU and other computing resources are run on the SoC according to actual needs, the bandwidth usage of other computing devices (computing resources) except the NPU is sampled at a specified time granularity, and the bandwidth occupied by other computing devices (computing resources) obtained by sampling is subtracted from the total bandwidth available on the SoC to obtain the actual bandwidth constraint.
[0037] The performance verification phase includes: obtaining the running performance of the NPU in the bandwidth sampling phase, and the running performance includes the corresponding subgraph performance recorded in the bandwidth sampling phase, that is, the corresponding subgraph performance when running the executable file, such as IO performance, computing performance and overall performance. It can be understood that when running the executable file, the npu will count the actual running performance; verify the difference between the estimated performance and the running performance. If the difference is less than or equal to the threshold, the optimization is completed. If the difference is greater than the threshold, the actual bandwidth constraint is used as the specified bandwidth constraint to re-compile the phase, bandwidth sampling phase and performance verification phase until the difference is less than or equal to the threshold. In this way, the compilation effect of the AI application can be optimized according to the actual bandwidth availability, and a more efficiency-friendly version can be generated. It can be understood that when the compilation phase is executed for the first time, the bandwidth constraint is specified as the preset value, and the actual bandwidth constraint is used as the specified bandwidth constraint when the compilation phase is re-executed subsequently. The above differences are quantified into various performance indicators, which are specific values.
[0038] Based on another aspect of the present invention, the present invention also provides an NPU performance optimization system, including: a compilation module, a bandwidth sampling module and a performance verification module.
[0039] The compilation module is used to divide the computational graph into multiple subgraphs, complete the performance evaluation of the subgraphs in combination with the specified bandwidth constraints, obtain the estimated performance of the subgraphs, and generate NPU execution files. In this embodiment, the compilation module can be, for example, a compiler, and the compilation stage is completed on the compiler. Furthermore, the compilation module includes a segmentation module, a performance evaluation module, and a code generation module, and the compilation module supports alternating iterative execution of the segmentation module and the performance evaluation module.
[0040] The bandwidth sampling module is used to run the execution file, which includes the NPU execution file, and obtain the actual bandwidth constraint through sampling.
[0041] The performance verification module is used to obtain the operating performance of the NPU in the bandwidth sampling phase and verify the difference between the estimated performance and the operating performance. If the difference is less than or equal to the threshold, the optimization is completed. If the difference is greater than the threshold, the actual bandwidth constraint is used as the specified bandwidth constraint to re-compile the phase, bandwidth sampling phase and performance verification phase until the difference is less than or equal to the threshold.
[0042] As configured above, the actual bandwidth constraint is sampled in the bandwidth sampling stage. In the performance verification stage, if the difference between the estimated performance and the running performance is greater than the threshold, the actual bandwidth constraint is used as the specified bandwidth constraint to re-compile the stage, bandwidth sampling stage and performance verification stage until the difference is less than or equal to the threshold, that is, except for the initial compilation stage, the subsequent compilation stages are all based on the estimated performance obtained according to the actual bandwidth constraint. In summary, the present invention uses dynamic bandwidth constraints to replace the static bandwidth constraints of the prior art, fully utilizes the bandwidth resources on the SoC, optimizes the NPU performance, can optimize the AI application compilation effect according to the actual bandwidth availability, generate a more efficiency-friendly version, and improve the performance of the NPU without affecting other computing devices on the SoC.
[0043] The following is a specific embodiment:
[0044] Please refer to Figure 1 and Figure 2 , Figure 2 Indicates that the compilation module and performance verification module are integrated on the Host, and the SoC side includes the bandwidth sampling module, memory management module, NPU and other computing resources. The connection method is as follows Figure 2 As shown in Figure 2, the actual bandwidth constraint is obtained by real-time sampling of the memory usage of the NPU and other computing resources.
[0045] (1) Initialize the bandwidth constraint according to a preset fixed value, that is, specify the bandwidth constraint as a preset value at this time.
[0046] (2) Create a mapping relationship between a subgraph starting point and time and cache it. The time cache is initialized to 0, and the subgraph starting point is initialized to the starting point of the input AI application calculation graph.
[0047] (3) The subgraphs are segmented according to the starting point of the subgraph and the input AI application calculation graph. The segmentation rules and constraints can be customized to ensure that all segmented subgraphs start from the same subgraph starting point. There can be repeated points between subgraphs, and they can include or be included. The start time of all subgraphs in this round is obtained from the time cache.
[0048] (4) Evaluate the estimated performance of the subgraph: Obtain the accurate available bandwidth at each time granularity based on the specified bandwidth constraint and the start time of each subgraph. For a more accurate evaluation, the start time of each IO behavior in the corresponding subgraph can be calculated in combination with the parallel characteristics of the NPU hardware and the execution model, thereby determining the available bandwidth at each time granularity in the IO process, and evaluating the performance in segments. After completing the performance evaluation of all subgraphs, the optimal subgraph can be selected according to customized rules, such as the highest MAC utilization rate, the lowest proportion of redundant data, etc. The end node of the optimal subgraph is used as the new subgraph starting point, and the estimated overall running time of the optimal subgraph is updated to the time cache. Record the IO performance, computing performance, and overall performance of the optimal subgraph.
[0049] (5) Determine whether the end node of the optimal subgraph is the same as the end node of the AI application computation graph. If not, jump to step (3); if the same, proceed to step (6).
[0050] (6) Generate code to be executed on the corresponding NPU based on all optimal subgraph information and the built-in operator library and auxiliary components.
[0051] (7) Run the NPU and other computing resources on the SoC according to actual needs, sample the bandwidth usage of other computing devices except the NPU at the specified time granularity, and subtract the sampled bandwidth from the total bandwidth available on the SoC to obtain the actual bandwidth constraint. Collect the NPU's operating information, namely the operating performance, including the IO performance, computing performance, and overall performance of each subgraph.
[0052] (8) Verify the difference between the estimated performance of the subgraph recorded in step (4) and the running performance of the corresponding subgraph collected in step (7) according to the customized rules, such as the proportion of IO data volume, the proportion of IO and computing parallelism, IO, computing and overall performance errors, actual bandwidth curve and bandwidth constraints, etc.
[0053] (9) Determine whether the performance meets expectations based on the result of step (8) and the corresponding threshold. If it meets expectations, complete the NPU performance optimization. If it does not meet expectations, proceed to step (10).
[0054] (10) Update the actual bandwidth constraint obtained in step (7) to the specified bandwidth constraint, and jump to step (2), and repeat the cycle.
[0055] It should be noted that references to "one embodiment", "an embodiment", "a specific embodiment", "some embodiments", etc. in the specification only indicate that the described embodiment may include a particular feature, structure or characteristic. Moreover, such phrases do not necessarily refer to the same embodiment. In addition, when a particular feature, structure or characteristic is described in conjunction with an embodiment, whether or not explicitly described, it is within the knowledge of a person skilled in the relevant art to implement such feature, structure or characteristic in conjunction with other embodiments.
[0056] It should be noted that the various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description.
[0057] It should also be noted that, although the present invention has been disclosed as a preferred embodiment, the above embodiment is not intended to limit the present invention. For any technician familiar with the art, without departing from the scope of the technical solution of the present invention, the technical content disclosed above can be used to make many possible changes and modifications to the technical solution of the present invention, or modified into equivalent embodiments of equivalent changes. Therefore, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the content of the technical solution of the present invention still falls within the scope of protection of the technical solution of the present invention.
[0058] It should also be understood that, unless otherwise specified or indicated, the terms "first", "second", "third", etc. in the specification are merely used to distinguish between the various components, elements, steps, etc. in the specification, and are not used to indicate the logical relationship or sequential relationship between the various components, elements, steps, etc.
[0059] It should also be recognized that the terms described herein are only used to describe specific embodiments and are not intended to limit the scope of the invention. It should be noted that the singular forms "a" and "an" used herein and in the appended claims include plural references unless the context clearly indicates otherwise. For example, a reference to "a step" or "a device" means a reference to one or more steps or devices, and may include secondary steps and secondary devices. All conjunctions used should be understood in the broadest sense. And, the word "or" should be understood to have the definition of a logical "or", rather than a logical "exclusive or", unless the context clearly indicates otherwise. In addition, the implementation of the method and / or device in the embodiments of the present invention may include performing the selected task manually, automatically, or in combination.
Claims
1. A method for optimizing NPU performance, characterized in that: It includes compilation phase, bandwidth sampling phase and performance verification phase; The compilation stage includes: dividing the computation graph into multiple subgraphs, completing the performance evaluation of the subgraphs in combination with the specified bandwidth constraints, obtaining the estimated performance of the subgraphs, and generating NPU execution files; The bandwidth sampling stage includes: running an execution file, the execution file includes the NPU execution file, sampling to obtain an actual bandwidth constraint; The performance verification phase includes: obtaining the operating performance of the NPU in the bandwidth sampling phase, verifying the difference between the estimated performance and the operating performance, if the difference is less than or equal to a threshold, the optimization is completed, if the difference is greater than the threshold, the actual bandwidth constraint is used as the specified bandwidth constraint to re-compile the phase, bandwidth sampling phase and performance verification phase until the difference is less than or equal to the threshold.
2. The NPU performance optimization method according to claim 1, characterized in that: The compilation stage includes a segmentation stage, a performance evaluation stage and a code generation stage. The computation graph is segmented in the segmentation stage; the performance evaluation is completed in the performance evaluation stage to obtain the estimated performance; an NPU execution file is generated in the code generation stage; the segmentation stage and the performance evaluation stage are iteratively executed until the segmentation of the entire computation graph is completed, and then the code generation stage is executed.
3. The NPU performance optimization method according to claim 2, characterized in that: In the performance evaluation stage, the optimal subgraph is selected according to the result of the performance evaluation, and it is determined whether the end node of the optimal subgraph is consistent with the end node of the computational graph. If not, the end node of the subgraph is used as the starting point of the next subgraph to execute the segmentation stage. If they are consistent, it indicates that the segmentation of the entire computational graph is completed, and the code generation stage begins.
4. The NPU performance optimization method according to claim 1, characterized in that: The step of completing the performance evaluation of the subgraph includes: obtaining the available bandwidth at each time granularity in combination with the specified bandwidth constraint and the start time of the subgraph, and performing performance evaluation on the subgraph in combination with the available bandwidth.
5. The NPU performance optimization method according to claim 1, characterized in that: The execution file also includes execution files of other computing resources except the NPU execution file. The execution files of other computing resources can be customized and generated according to actual SoC computing requirements.
6. The NPU performance optimization method according to claim 5, characterized in that: The actual bandwidth constraint in the bandwidth sampling phase is updated as follows: the total bandwidth on the SoC is subtracted from the bandwidth used by the other computing resources sampled at the specified granularity when the SoC is actually running.
7. The NPU performance optimization method according to claim 1, characterized in that: In the compilation stage, the NPU execution file is generated according to the segmentation information and the operator library.
8. The NPU performance optimization method according to claim 1, characterized in that: When the compilation phase is executed for the first time, the specified bandwidth constraint is a preset value, and when the compilation phase is executed subsequently, the actual bandwidth constraint is used as the specified bandwidth constraint.
9. A NPU performance optimization system, characterized in that: include: Compilation module, bandwidth sampling module and performance verification module; The compiling module is used to divide the computation graph into multiple subgraphs, complete the performance evaluation of the subgraphs in combination with the specified bandwidth constraints, obtain the estimated performance of the subgraphs, and generate NPU execution files; The bandwidth sampling module is used to run an execution file, the execution file includes the NPU execution file, and sample to obtain the actual bandwidth constraint; The performance verification module is used to obtain the operating performance of the NPU in the bandwidth sampling stage, verify the difference between the estimated performance and the operating performance, if the difference is less than or equal to a threshold, the optimization is completed, if the difference is greater than the threshold, the actual bandwidth constraint is used as the specified bandwidth constraint to re-compile the stage, bandwidth sampling stage and performance verification stage until the difference is less than or equal to the threshold.
10. The NPU performance optimization system according to claim 9, characterized in that: The estimated performance includes IO performance and computing performance.