Adaptive precision adjustment method in NPU field
The NPU adaptive precision adjustment method addresses the challenge of dynamic precision adjustment in NPU hardware by employing task analysis and mixed precision conflict resolution, enhancing resource utilization and extending battery life in edge devices.
Patent Information
- Application Number
- CN202510498856.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-15
AI Technical Summary
Existing NPU hardware cannot flexibly adjust calculation accuracy according to task requirements, resulting in waste of performance and resources, and low-precision calculations may cause quantization errors to affect model prediction performance.
Adaptive accuracy adjustment method is adopted, and data distribution and operator calculation complexity are monitored in real time through the task analysis module, a multi-dimensional scoring model is built, a hybrid accuracy conflict digestion mechanism is implemented, and the calculation accuracy is dynamically adjusted, and resource allocation is optimized through the hardware interface module to achieve high-precision and low-precision switching of critical paths.
It realizes the global optimal balance of accuracy and performance in deep learning inference tasks, improves the resource utilization of NPU, extends the battery life of edge devices, and improves the computing reliability of key operators.
Smart Images

Figure CN120315892A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of neural processing units, and specifically relates to a method for adaptive precision adjustment in the field of NPUs. Background Art
[0002] During the inference process of deep learning models, due to different task types, there are significant differences in the requirements for computing precision and performance. High-precision tasks (such as medical image analysis) require FP32 floating-point operations to ensure the accuracy of results, while real-time tasks (such as autonomous driving perception) are more concerned about performance and can use low-precision operations such as FP16 or INT8. Currently, NPU hardware usually adopts fixed computing precision.
[0003] However, traditional methods cannot be flexibly adjusted according to task requirements, resulting in waste of performance and resources. In addition, the quantization errors that may occur in low-precision calculations will also affect the prediction performance of the model.
[0004] Therefore, a method that can dynamically adjust computing precision is needed, and at the same time, a quantization-aware mechanism is combined to ensure the best balance between precision and performance during the inference process. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for adaptive precision adjustment in the field of NPUs in order to solve the above-mentioned problems.
[0006] The technical solution adopted by the present invention is as follows: A method for adaptive precision adjustment in the field of NPUs, the method comprising the following steps:
[0007] S1: The task analysis module obtains the input data stream of the current inference task through the hardware interface module, monitors the data distribution characteristics and operator calculation complexity in real time, and establishes a task precision requirement baseline;
[0008] S2: The task analysis module combines the task type and the hardware resource status to construct a multi-dimensional scoring model including computational intensity, error tolerance, and latency sensitivity, and generates a precision priority label for each operator;
[0009] S3: The task analysis module collaborates with the quantization-aware optimization module to introduce a mixed-precision conflict resolution mechanism: based on the quantization-aware model, predict the error propagation paths of different precision combinations, perform global precision reallocation on the operator groups with conflicts, ensure that the critical path preferentially uses high precision, and the non-critical path automatically reduces to low precision. At the same time, pre-load the computing resource allocation scheme through the hardware interface module to reduce the switching delay;
[0010] S4: According to the policy, the dynamic precision switching module sends instructions to the hardware interface module to trigger the reconstruction of the mixed-precision computing unit of the NPU; switches the weight buffer of the convolutional layer to the INT8 mode, retains FP16 for the activation function, and realizes zero additional overhead for precision switching through hardware-level register remapping and pipeline optimization;
[0011] S5: The quantization-aware optimization module monitors the output distribution of the intermediate layer in real time during the inference process and adopts a progressive scale factor adjustment algorithm: based on the sliding window to statistically calculate the feature mean and variance, dynamically correct the scale factor and zero-point offset, and directly inject the parameters into the quantization acceleration engine of the NPU through the hardware interface module to achieve the gradual convergence of the error with iterations;
[0012] S6: For the highly sensitive operators marked by the task analysis module, the dynamic precision switching module temporarily inserts FP32 calculation cores, and at the same time the quantization-aware optimization module freezes its quantization parameters, and the hardware interface module allocates exclusive cache channels to ensure the high reliability of the key results;
[0013] S7: The hardware interface module collects the power consumption and latency data of the NPU and feeds them back to the task analysis module to trigger the precision policy iteration: if the system detects that the battery power is lower than the threshold, it forces non-critical tasks to be downgraded to the INT8 mode, and enhances the error compensation through the quantization-aware optimization module to maintain the overall precision without deterioration;
[0014] S8: The quantization-aware optimization module performs parameter adjustment in stages: loads the pre-calibrated global quantization table before inference; fine-tunes the local parameters based on the real-time error gradient during inference; updates the global table through a lightweight perception model after inference to form a closed-loop optimization link;
[0015] S9: The task analysis module stores the optimization policy of the current task in the persistent buffer area of the hardware interface module for direct invocation by similar tasks, reduces the repeated analysis overhead, and ensures the dynamic update of the policy and hardware compatibility through the version comparison mechanism.
[0016] In a preferred embodiment, in step S1, the task analysis module accesses the data bus of the NPU through the hardware interface module to capture the real-time features of the input data stream in a streaming processing manner; adopts a multi-channel sampling technique to perform Gaussian kernel density estimation on the dynamic range of the input tensor, calculates its kurtosis and skewness to characterize the asymmetry of the data distribution; at the same time, performs binary masking analysis on the activation values through a sparse encoder to quantify the sparsity index; for the computing complexity of the operator, the module integrates a runtime performance analyzer to track the instruction-level parallelism and memory bandwidth occupancy rate, and constructs a Markov chain prediction model of the computing load in combination with the historical execution logs.
[0017] In a preferred embodiment, in step S2, the task analysis module constructs a decision-making model based on the analytic hierarchy process, decomposing computational intensity, error tolerance, and latency sensitivity into a three-level index tree; the computational intensity is obtained through the normalized weighted value of the operator floating-point operation number and the memory access ratio; the error tolerance uses backpropagation sensitivity analysis to calculate the condition number of the Jacobian matrix of the output layer gradient with respect to the quantization error; the latency sensitivity combines the proportion of the pipeline idle cycle feedback by the NPU hardware interface to dynamically adjust the weight coefficient; finally, an accuracy priority label is generated through fuzzy logic comprehensive evaluation.
[0018] In a preferred embodiment, in step S3, a mixed-precision conflict resolution mechanism is used to solve the problem of out-of-control error propagation caused by mixed-precision combinations;
[0019] The quantization-aware optimization module constructs an error propagation chain model based on the network topology structure, defines the transfer relationship of quantization errors between operators, and the specific calculation formula is:
[0020]
[0021] Where: E total is the overall quantization error of the task, e i is the local error introduced by the i-th operator due to precision reduction; α ij is the error transfer coefficient, indicating the influence weight of the error of the i-th operator on the downstream j-th operator, which is determined by the network dependency relationship 10 ≤ α ij ≤ 1)
[0022] The task analysis module screens out operator groups with precision conflicts according to the following formula:
[0023]
[0024] Where Gk represents a series operator group with a tight dependency relationship, and θconflict represents the conflict threshold;
[0025] The hardware interface module preloads the computing resource allocation scheme to minimize the precision switching delay, and the calculation formula is:
[0026]
[0027] Where T switch represents the total switching delay;
[0028] C new(i) , C old(i) represents the cache occupancy of the i-th operator under the new / old precision;
[0029] B cache represents the NPU cache bandwidth
[0030] δhardware Represents the hardware acceleration factor.
[0031] In a preferred embodiment, in step S4, the dynamic precision switching module issues microcode instructions through the hardware interface module to trigger the dynamic reconstruction of the NPU computing unit; for the mixed precision mode, the time-sharing multiplexing technology is adopted to realize the coexistence of multi-precision computing cores.
[0032] In a preferred embodiment, in step S5, the quantization-aware optimization module deploys a lightweight monitoring agent during the inference process to continuously collect the statistical histograms of the outputs of each layer; the moving average and standard deviation of the activation values are calculated using a sliding window mechanism, and the distribution difference before and after quantization is evaluated based on the KL divergence; the scale factor update adopts the incremental Newton iteration method, and the calculation formula is:
[0033]
[0034] where η is the adaptive learning rate, D KL is the distribution divergence of the current window; the zero-point offset is dynamically solved through a linear programming model that minimizes the quantization reconstruction error; the parameter update instruction is directly written into the quantization register set of the NPU through the hardware interface module to achieve single-cycle endogenesis.
[0035] In a preferred embodiment, in step S6, for high-sensitivity operators, the dynamic precision switching module starts the precision redundancy mechanism: a dedicated high-precision computing unit is reserved in the NPU, and an exclusive memory protection domain is configured through the hardware interface module to prohibit the memory access of low-precision tasks; the quantization-aware optimization module implements double-buffer isolation for the input and output tensors of such operators, where the main buffer is used for FP32 calculations, and the shadow buffer synchronously updates the low-precision quantization results for error comparison.
[0036] In a preferred embodiment, in step S7, the hardware interface module integrates a power consumption sensor to collect the voltage-frequency curve and temperature gradient data of each computing cluster of the NPU in real time, and constructs a three-dimensional energy efficiency model; when the system power is lower than the threshold, the task analysis module starts a greedy algorithm: traverses the precision degradation combinations of non-critical operators, and selects the scheme that maximally improves the energy efficiency ratio and minimizes the precision loss; the quantization-aware optimization module synchronously injects a noise compensation filter to perform frequency-domain band-limited filtering on the output of the downscaled precision operator to suppress high-frequency quantization noise.
[0037] In a preferred embodiment, in step S8, in the pre-inference stage, the quantization-aware optimization module loads the global codebook generated by offline calibration, and uses vector quantization technology to cluster the weights into 256 prototype vectors; in the in-inference stage, the prototype vectors are linearly interpolated and adjusted based on the real-time activation distribution, and the interpolation weights are dynamically determined by the cache hit rate of the hardware interface module; in the post-inference stage, the parameter adjustment amount during the inference process is backpropagated to the global codebook through an online distillation mechanism, and the codebook vectors are updated using stochastic gradient descent with momentum; the data in the three stages are transferred with zero-copy through the DMA engine of the hardware interface module, forming an end-to-end self-evolving link for quantization parameters.
[0038] In a preferred embodiment, in step S9, the task analysis module encodes the optimization strategy into a binary bitmask and a parameter matrix, stores them in the persistent storage area of the NPU after being encrypted by AES-256 of the hardware interface module; a differential hashing algorithm is used to generate a strategy fingerprint, and a Bloom filter is used to quickly match similar tasks; when a new task is triggered, the hardware interface module executes policy loading and task analysis in parallel. If the strategy fingerprint matching degree exceeds 95%, the cached policy is directly enabled, otherwise incremental policy fusion is started: the feature vector of the new task is cosine-similarity matched with the historical policy library, and a hybrid policy is generated by dynamic weighting; this mechanism shortens the cold start time from 230 ms to 65 ms.
[0039] In summary, due to the adoption of the above technical solutions, the beneficial effects of the present invention are:
[0040] 1. In the present invention, through the mixed-precision conflict resolution mechanism, the global optimal balance between precision and performance in deep learning inference tasks is achieved. When traditional methods adjust the computing precision, they often make independent decisions only based on the characteristics of local operators, resulting in the overall precision collapse caused by error accumulation in the operator groups with serial dependencies. This method constructs an error propagation chain through a quantization-aware model, dynamically predicts the influence path of different precision combinations on the task output, and uses the critical path first algorithm to perform differential precision allocation on the conflicting operator groups. In the object detection task, high precision calculations are forcibly retained in the critical paths such as the classification head, while the background filtering layer is reduced to a low-precision mode, so that the decrease in detection precision is controlled within 0.3%, and the inference speed is increased by 2.1 times. This mechanism fundamentally solves the problem of global sub-optimality caused by local optimization in traditional methods and ensures the computational reliability of high-value operators.
[0041] 2. In the present invention, through the pre-allocation strategy of hardware collaboration and dynamic resource scheduling, the switching overhead and energy consumption cost of mixed-precision computing are significantly reduced. Traditional precision switching requires frequent reconstruction of computing units and data transfer, resulting in a delay of up to 10 microseconds. However, this method uses register remapping and pipeline prediction execution technology to compress the switching delay to less than 200 nanoseconds. At the same time, combined with the energy efficiency adaptive feedback mechanism, the system can dynamically adjust the precision strategy of non-critical tasks according to the battery state, extending the battery life by 58% in actual measurements on edge devices, and the degradation of task precision is less than 1%. This optimized mode of deep hardware collaboration not only improves the resource utilization rate of the NPU but also expands the applicable boundary of the adaptive precision adjustment technology in low-power scenarios, providing a highly robust solution for scenarios with strict real-time requirements such as autonomous driving and mobile inference. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a schematic diagram of the process principle of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0044] Embodiment:
[0045] Referring to Figure 1 , a method for adaptive precision adjustment in the field of NPU, the method includes the following steps:
[0046] S1: The task analysis module obtains the input data stream of the current inference task through the hardware interface module, and real-time monitors the data distribution characteristics (such as dynamic range, sparsity) and the calculation complexity of the operator, and establishes a baseline for task precision requirements.
[0047] S2: The task analysis module combines the task type (such as classification, detection) and the hardware resource status to construct a multi-dimensional scoring model including computational intensity, error tolerance, and latency sensitivity, and generates a precision priority label (FP32 / FP16 / INT8) for each operator.
[0048] S3: The task analysis module collaborates with the quantization-aware optimization module and introduces a mixed-precision conflict resolution mechanism: based on the quantization-aware model, predict the error propagation path of different precision combinations, and perform global precision re-allocation on the operator groups with conflicts (such as series-dependent operators) to ensure that the critical path preferentially uses high precision, and the non-critical path automatically reduces to low precision. At the same time, pre-load the computing resource allocation scheme through the hardware interface module to reduce the switching delay.
[0049] S4: The dynamic precision switching module sends instructions to the hardware interface module according to the policy, triggering the reconstruction of the mixed-precision computing unit of the NPU. For example, switch the weight buffer of the convolutional layer to the INT8 mode, and keep the activation function in FP16. Through hardware-level register remapping and pipeline optimization, zero additional overhead for precision switching is achieved.
[0050] S5: The quantization-aware optimization module monitors the output distribution of the intermediate layer in real time during the inference process, and adopts a progressive scale factor adjustment algorithm: based on the sliding window to statistically calculate the feature mean and variance, dynamically correct the scale factor and zero-point offset, and directly inject the parameters into the quantization acceleration engine of the NPU through the hardware interface module, so that the error gradually converges with iterations.
[0051] S6: For the highly sensitive operators marked by the task analysis module (such as the edge detection layer in medical image segmentation), the dynamic precision switching module temporarily inserts FP32 calculation cores, and at the same time the quantization-aware optimization module freezes its quantization parameters, and the hardware interface module allocates exclusive cache channels to ensure the high reliability of key results.
[0052] S7: The hardware interface module collects the power consumption and latency data of the NPU, and feeds it back to the task analysis module, triggering the precision policy iteration: if the system detects that the battery power is lower than the threshold, then force non-critical tasks to be reduced to the INT8 mode, and enhance the error compensation through the quantization-aware optimization module to maintain the overall precision without deterioration.
[0053] S8: The quantization-aware optimization module performs parameter adjustment in stages: load the pre-calibrated global quantization table before inference; fine-tune local parameters based on real-time error gradients during inference; update the global table through a lightweight perception model after inference to form a closed-loop optimization link.
[0054] S9: The task analysis module stores the optimization strategy (such as the precision allocation table, quantization parameters) of the current task in the persistent buffer area of the hardware interface module for direct invocation by similar tasks, reducing the repeated analysis overhead, and ensuring the dynamic update of the strategy and hardware compatibility through the version comparison mechanism.
[0055] In step S1, the task analysis module accesses the data bus of the NPU through the hardware interface module, and captures the real-time features of the input data stream in a streaming processing manner. Adopting multi-channel sampling technology, perform Gaussian kernel density estimation on the dynamic range of the input tensor, and calculate its kurtosis and skewness to characterize the asymmetry of the data distribution; at the same time, perform binary masking analysis on the activation values through a sparse encoder to quantify the sparsity index. For the computational complexity of operators, the module integrates a runtime performance analyzer to track the instruction-level parallelism (ILP) and memory bandwidth occupancy rate, and constructs a Markov chain prediction model of the computational load in combination with historical execution logs. The module stores the feature vectors through a time series database, providing millisecond-level updated dynamic support for the precision requirement baseline.
[0056] In step S2, the task analysis module constructs a decision-making model based on the Analytic Hierarchy Process (AHP), and decomposes the computational intensity, error tolerance, and latency sensitivity into a three-level index tree. The computational intensity is obtained through the normalized weighted value of the operator floating-point operation count (FLOPs) and the memory access ratio (OPS / Mem); the error tolerance uses backpropagation sensitivity analysis to calculate the condition number of the Jacobian matrix of the output layer gradient with respect to the quantization error; the latency sensitivity combines the ratio of the pipeline idle cycle feedback by the NPU hardware interface to dynamically adjust the weight coefficient. Finally, the precision priority label is generated through fuzzy logic comprehensive evaluation, where the operator corresponding to the FP32 label needs to meet the error sensitivity threshold (such as gradient condition number > 1e4) and the computational intensity is lower than 60% of the hardware peak computing power.
[0057] In step S3, the mixed-precision conflict resolution mechanism is used to solve the problem of out-of-control error propagation caused by mixed-precision combinations;
[0058] Based on the network topology structure, the quantization-aware optimization module constructs an error propagation chain model and defines the transfer relationship of quantization errors between operators. The specific calculation formula is:
[0059]
[0060] where: E total is the overall quantization error of the task, and e i is the local error introduced by the i-th operator due to precision reduction; α ij is the error transfer coefficient, indicating the influence weight of the error of the i-th operator on the downstream j-th operator, which is determined by the network dependency relationship (10 ≤ α ij ≤ 1)
[0061] The task analysis module screens the operator groups with precision conflicts according to the following formula:
[0062]
[0063] where Gk represents a series operator group with a tight dependency relationship, and θconflict represents the conflict threshold;
[0064] The hardware interface module preloads the computing resource allocation scheme to minimize the precision switching delay. The calculation formula is:
[0065]
[0066] where T switch represents the total switching delay;
[0067] C new(i) , C old(i) represent the cache occupancy of the i-th operator under the new / old precision;
[0068] B cache Represents the NPU cache bandwidth (unit: GB / s)
[0069] δ hardware Represents the hardware acceleration factor (determined by the register remapping efficiency, typical value 0.2 - 0.5);
[0070] In the step S4, the dynamic precision switching module issues microcode instructions through the hardware interface module to trigger the dynamic reconstruction of the NPU computing unit. For the mixed precision mode, the time-division multiplexing technology is adopted to realize the coexistence of multi-precision computing cores: for example, independent partitions are divided in the Tensor Core, which are respectively configured as INT8 integer multiplier - accumulators and FP16 floating - point multiplier - accumulators, and the computing path switching is controlled by gated clocks. The hardware interface module uses the address offset remapping technology to align the storage addresses of different precision data blocks to the same physical cache area, eliminating the data transfer overhead. For pipeline conflicts, the speculative execution mechanism is adopted to pre - load the instruction set of the next precision mode, and the pipeline bubbles are eliminated through the dynamic branch predictor, reducing the precision switching latency from 10 μs in the traditional scheme to within 200 ns.
[0071] In the step S5, the quantization - aware optimization module deploys lightweight monitoring agents during the inference process to continuously collect the statistical histograms of the outputs of each layer. The moving average and standard deviation of the activation values are calculated using the sliding window mechanism (the window size is adjustable, default 32 frames), and the distribution difference before and after quantization is evaluated based on the KL divergence. The scale factor update adopts the incremental Newton iteration method, and the calculation formula is:
[0072]
[0073] where η is the adaptive learning rate, and D KL is the distribution divergence of the current window. The zero - point offset is dynamically solved through the linear programming model that minimizes the quantization reconstruction error. The parameter update instructions are directly written into the quantization register group of the NPU through the hardware interface module, and take effect within a single cycle.
[0074] In the step S6, for highly sensitive operators, the dynamic precision switching module activates the precision redundancy mechanism: reserve a dedicated high - precision computing unit (such as FP32 MAC array) in the NPU, configure an exclusive memory protection domain through the hardware interface module, and prohibit the memory access of low - precision tasks. The quantization - aware optimization module implements double - buffer isolation for the input and output tensors of such operators. The main buffer is used for FP32 calculations, and the shadow buffer synchronously updates the low - precision quantization results for error comparison. At the same time, the module introduces the residual compensation technology, and superimposes the calculation difference between FP32 and FP16 as a compensation term to the downstream operators to ensure that the error propagation rate of the critical path is lower than 0.05% / layer.
[0075] In step S7, the hardware interface module set successfully consumes sensors to collect the voltage-frequency curve and temperature gradient data of each computing cluster of the NPU in real time, and constructs a three-dimensional energy efficiency model (unit: TOPS / (W·℃)). When the system power is lower than the threshold, the task analysis module starts the greedy algorithm: traversing the precision degradation combinations of non-critical operators, and selecting the scheme that maximally improves the energy efficiency ratio and minimizes the precision loss. The quantization-aware optimization module synchronously injects a noise compensation filter to perform frequency-domain band-limited filtering on the output of the downscaled precision operator to suppress high-frequency quantization noise. Experiments show that this mechanism can extend the battery life by 58% on edge devices, and the PSNR index only drops by 0.8 dB.
[0076] In step S8, in the pre-inference stage, the quantization-aware optimization module loads the global codebook generated by offline calibration, and uses vector quantization technology to cluster the weights into 256 prototype vectors; in the in-inference stage, based on the real-time activation distribution, linear interpolation adjustment is performed on the prototype vectors, and the interpolation weights are dynamically determined by the cache hit rate of the hardware interface module; in the post-inference stage, the parameter adjustment amount in the inference process is backpropagated to the global codebook through the online distillation mechanism, and the codebook vectors are updated using the stochastic gradient descent method with momentum. The data in the three stages are transferred without copy through the DMA engine of the hardware interface module, forming an end-to-end self-evolving link of quantization parameters.
[0077] In step S9, the task analysis module encodes the optimization strategy into a binary bitmask and a parameter matrix, which are stored in the persistent storage area of the NPU after being encrypted by AES-256 of the hardware interface module. The differential hashing algorithm is used to generate the strategy fingerprint, and the same type of tasks are quickly matched through the Bloom filter. When a new task is triggered, the hardware interface module executes the strategy loading and task analysis in parallel. If the strategy fingerprint matching degree exceeds 95%, the cached strategy is directly enabled, otherwise incremental strategy fusion is started: the feature vector of the new task is matched with the historical strategy library by cosine similarity, and a hybrid strategy is dynamically weighted. This mechanism shortens the cold start time from 230 ms to 65 ms.
[0078] It can be known from the above that:
[0079] In the present invention, through the mixed-precision conflict resolution mechanism, the global optimal balance between precision and performance in deep learning inference tasks is achieved. When traditional methods adjust the computing precision, they often make independent decisions only based on the characteristics of local operators, resulting in the overall precision collapse caused by error accumulation in operator groups with serial dependencies. This method constructs an error propagation chain through a quantization-aware model, dynamically predicts the influence paths of different precision combinations on the task output, and uses the critical path first algorithm to implement differential precision allocation for conflicting operator groups. In the object detection task, high precision calculations are forcibly retained for critical paths such as classification heads, while the background filtering layer is reduced to a low-precision mode, so that the decrease in detection precision is controlled within 0.3%, and the inference speed is increased by 2.1 times. This mechanism fundamentally solves the problem of global sub-optimality caused by local optimization in traditional methods and ensures the computational reliability of high-value operators.
[0080] In the present invention, through the pre-allocation strategy and dynamic resource scheduling of hardware cooperation, the switching overhead and energy consumption cost of mixed-precision calculations are significantly reduced. Traditional precision switching requires frequent reconstruction of computing units and data transfer, resulting in a delay of up to 10 microseconds. This method uses register remapping and pipeline prediction execution technology to compress the switching delay to less than 200 nanoseconds. At the same time, combined with the energy efficiency adaptive feedback mechanism, the system can dynamically adjust the precision strategy of non-critical tasks according to the battery state, extending the battery life by 58% in actual measurements of edge devices, and the degradation of task precision is less than 1%. This optimized mode of deep hardware cooperation not only improves the resource utilization rate of the NPU, but also expands the applicable boundary of adaptive precision adjustment technology in low-power scenarios, providing a highly robust solution for scenarios with strict real-time requirements such as autonomous driving and mobile inference.
[0081] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the phrase "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.
[0082] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for adaptive precision adjustment in the field of NPUs, characterized in that: The method includes the following steps: S1: The task analysis module obtains the input data stream of the current inference task through the hardware interface module, monitors the data distribution characteristics and operator calculation complexity in real time, and establishes a task accuracy requirement baseline; S2: The task analysis module combines the task type and the hardware resource status to construct a multi-dimensional scoring model including computational intensity, error tolerance, and latency sensitivity, and generates accuracy priority labels for each operator; S3: The task analysis module collaborates with the quantization-aware optimization module and introduces a mixed-precision conflict resolution mechanism: based on the quantization-aware model, predict the error propagation path of different precision combinations, perform global accuracy reallocation on the operator groups with conflicts, ensure that the critical path preferentially uses high precision, and the non-critical path is automatically reduced to low precision. At the same time, pre-load the computing resource allocation scheme through the hardware interface module to reduce the switching delay; S4: According to the strategy, the dynamic precision switching module sends an instruction to the hardware interface module to trigger the reconstruction of the mixed-precision computing unit of the NPU; switch the weight buffer of the convolutional layer to the INT8 mode, and retain the FP16 for the activation function. Through hardware-level register remapping and pipeline optimization, achieve zero additional overhead for precision switching; S5: The quantization-aware optimization module monitors the output distribution of the intermediate layer in real time during the inference process and adopts a progressive scale factor adjustment algorithm: based on the sliding window to statistically calculate the feature mean and variance, dynamically correct the scale factor and zero-point offset, and directly inject the parameters into the quantization acceleration engine of the NPU through the hardware interface module to achieve the gradual convergence of the error with iterations; S6: For the highly sensitive operators marked by the task analysis module, the dynamic precision switching module temporarily inserts FP32 computing cores, and at the same time the quantization-aware optimization module freezes its quantization parameters, and the hardware interface module allocates exclusive cache channels to ensure the high reliability of the key results; S7: The hardware interface module collects the power consumption and latency data of the NPU and feeds them back to the task analysis module to trigger the accuracy strategy iteration: if the system detects that the battery power is lower than the threshold, force the non-critical tasks to be reduced to the INT8 mode, and enhance the error compensation through the quantization-aware optimization module to maintain the overall accuracy without deterioration; S8: The quantization-aware optimization module performs parameter adjustment in stages: load the pre-calibrated global quantization table before inference; fine-tune the local parameters based on the real-time error gradient during inference; update the global table through the lightweight perception model after inference to form a closed-loop optimization link; S9: The task analysis module stores the optimization strategy of the current task in the persistent buffer area of the hardware interface module for direct invocation by similar tasks, reduces the repeated analysis overhead, and ensures the dynamic update of the strategy and hardware compatibility through the version comparison mechanism.
2. The method for adaptive precision adjustment in the field of NPU according to claim 1, wherein: In step S1, the task analysis module accesses the data bus of the NPU through the hardware interface module to capture the real-time characteristics of the input data stream in a streaming processing manner; adopts a multi-channel sampling technique to perform Gaussian kernel density estimation on the dynamic range of the input tensor, and calculates its kurtosis and skewness to characterize the asymmetry of the data distribution; At the same time, perform binary mask analysis on the activation values through a sparse encoder to quantify the sparsity index; For the computational complexity of the operator, the module integrates a runtime performance analyzer, tracks the instruction-level parallelism and memory bandwidth occupancy rate, and constructs a Markov chain prediction model of the computational load in combination with the historical execution logs.
3. A method for adaptive precision adjustment in the field of NPU according to claim 1, characterized in that: In the step S2, the task analysis module constructs a decision-making model based on the analytic hierarchy process, and decomposes the computational intensity, error tolerance, and latency sensitivity into a three-level index tree; the computational intensity is obtained by the normalized weighted value of the ratio of the operator's floating-point operation count to the memory access; The error tolerance adopts the backpropagation sensitivity analysis to calculate the condition number of the Jacobian matrix of the output layer gradient with respect to the quantization error; the latency sensitivity combines the ratio of the pipeline idle cycle percentage fed back by the NPU hardware interface to dynamically adjust the weight coefficient; finally, the accuracy priority label is generated through fuzzy logic comprehensive evaluation.
4. A method for adaptive precision adjustment in the field of NPUs as described in claim 1, characterized in that: In the step S3, through the mixed-precision conflict resolution mechanism, the problem of out-of-control error propagation caused by the mixed-precision combination is solved; The quantization-aware optimization module constructs an error propagation chain model based on the network topology structure, defines the transfer relationship of the quantization error between operators, and the specific calculation formula is: Among them: E total is the overall quantization error of the task, and e i is the local error introduced by the i-th operator due to precision reduction; α ij is the error transfer coefficient, indicating the influence weight of the error of the i-th operator on the downstream j-th operator, which is determined by the network dependency relationship (10 ≤ α ij ≤ 1) The task analysis module screens the operator groups with accuracy conflicts according to the following formula: Where Gk represents the series operator group with close dependence, and θconflict represents the conflict threshold; The hardware interface module preloads the computing resource allocation scheme to minimize the accuracy switching delay, and the calculation formula is: where T switch represents the total switching delay; C new(i) , C old(i) represents the cache occupancy of the i-th operator under the new / old precision; B cache Indicates the NPU cache bandwidth δ hardware represents the hardware acceleration factor.
5. The method for field adaptive precision adjustment of an NPU according to claim 1, wherein: In the step S4, the dynamic accuracy switching module issues microcode instructions through the hardware interface module to trigger the dynamic reconstruction of the NPU computing unit; for the mixed-precision mode, the time-sharing multiplexing technology is adopted to realize the coexistence of multiple-precision computing cores.
6. The method for adaptive precision adjustment in the field of NPU according to claim 1, wherein: In the step S5, the quantization-aware optimization module deploys a lightweight monitoring agent during the inference process to continuously collect the statistical histograms of the outputs of each layer; the sliding window mechanism is used to calculate the moving average and standard deviation of the activation values, and the distribution difference before and after quantization is evaluated based on the KL divergence; the scale factor update adopts the incremental Newton iteration method, and the calculation formula is: where η is the adaptive learning rate, and D KL is the distribution divergence of the current window; the zero-point offset is dynamically solved by a linear programming model that minimizes the quantization reconstruction error; the parameter update instruction is directly written into the quantization register bank of the NPU through the hardware interface module to achieve single-cycle endogeneity.
7. A method for adaptive precision adjustment in the field of NPU as described in claim 1, characterized in that: In the step S6, for the highly sensitive operators, the dynamic accuracy switching module starts the accuracy redundancy mechanism: reserve a dedicated high-precision computing unit in the NPU, configure an exclusive memory protection domain through the hardware interface module, and prohibit the memory access of low-precision tasks; the quantization-aware optimization module implements double-buffer isolation for the input and output tensors of such operators, and the main buffer is used for FP32 calculation, and the shadow buffer synchronously updates the low-precision quantization results for error comparison.
8. A method for field adaptive precision adjustment of an NPU according to claim 1, characterized in that: In the step S7, the hardware interface module integrates power consumption sensors to collect the voltage-frequency curve and temperature gradient data of each computing cluster of the NPU in real time, and constructs a three-dimensional energy efficiency model; when the system power is lower than the threshold, the task analysis module starts the greedy algorithm: traverse the accuracy degradation combinations of non-critical operators, and select the scheme that maximally improves the energy efficiency ratio and minimizes the accuracy loss; the quantization-aware optimization module synchronously injects a noise compensation filter to perform frequency-domain band-limited filtering on the output of the downscaled operators to suppress high-frequency quantization noise.
9. A method for adaptive precision adjustment in the field of NPU as described in claim 1, characterized in that: In the step S8, in the pre-inference stage, the quantization-aware optimization module loads the global codebook generated by offline calibration and clusters the weights into 256 prototype vectors using vector quantization technology; in the in-inference stage, the prototype vectors are linearly interpolated and adjusted based on the real-time activation distribution, and the interpolation weights are dynamically determined by the cache hit rate of the hardware interface module; in the post-inference stage, the parameter adjustment amount during the inference process is backpropagated to the global codebook through the online distillation mechanism, and the codebook vectors are updated using the stochastic gradient descent method with momentum; the data in the three stages are transferred through the DMA engine of the hardware interface module with zero-copy to form an end-to-end self-evolving link of quantization parameters.
10. A method for field adaptive precision adjustment in an NPU, as described in claim 1, characterized in that: In the step S9, the task analysis module encodes the optimization strategy into a binary bitmask and a parameter matrix, stores them in the persistent storage area of the NPU after being encrypted by AES-256 of the hardware interface module; the strategy fingerprint is generated using the difference hashing algorithm, and the same type of tasks are quickly matched through the Bloom filter; when a new task is triggered, the hardware interface module executes the strategy loading and task analysis in parallel. If the strategy fingerprint matching degree exceeds 95%, the cached strategy is directly enabled, otherwise the incremental strategy fusion is started.
Citation Information
Cited By
Block chain data processing method and security guarantee system
CN121193425A
Rewriting model construction method based on mixing precision strategy
CN122086634A
NPU performance evaluation method, device and equipment under SoC multi-channel and medium
CN122432009A
Methods, apparatus, equipment and media for evaluating NPU performance under SoC multi-path
CN122432009B