Neural network application feature perception approximate calculation unit design method and system

By applying the approximate computing unit design method of feature perception in neural networks, the PPO algorithm is used to automatically select the approximate computing unit, which solves the problem that traditional processors are difficult to meet the computing needs of large-scale neural networks, and realizes high-efficiency approximate computing optimization.

CN120354913APending Publication Date: 2025-07-22袁泉
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510399018.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently optimize the computing needs of neural networks without reducing accuracy. In particular, the complexity and differences in computing requirements of different neural network models make manual analysis time-consuming and labor-intensive, and traditional processors are difficult to meet huge computing needs.

Method used

The approximate calculation unit design method for neural network application feature perception is adopted. By building an approximate calculation unit library, the PPO algorithm model is used to automatically select the optimal approximate calculation unit based on the feature vector of the neural network, and combined with dynamic bit width adjustment and overflow processing, approximate calculation optimization is achieved.

Benefits of technology

It improves the energy efficiency of neural network computing, reduces error accumulation, improves search efficiency and optimization efficiency, and reduces power consumption and delay.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354913A_ABST
    Figure CN120354913A_ABST
Patent Text Reader

Abstract

The invention discloses an approximate calculation unit design method for neural network application feature perception. The method comprises the following steps: constructing an approximate calculation unit library; obtaining a corresponding feature vector according to the quantized data after the execution of each layer of topological structure of the target neural network; and inputting the obtained feature vectors into a trained PPO algorithm model, and obtaining an approximate calculation unit for execution of each layer of topological structure of the target neural network from an approximate calculation unit library. According to the method, the optimal / better approximate calculation unit can be generated under certain precision constraint according to the target neural network calculation structure and the application scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of neural network processors, and more specifically to a design method and system for an approximate computing unit for neural network application feature perception. Background Art

[0002] In recent years, artificial intelligence technology has developed rapidly. Deep learning models based on artificial neural networks (ANN) have shown strong performance in tasks such as prediction, recognition, classification, and adversarial generation. Behind this is the advancement of training methodology, the improvement of hardware computing power, and the expansion of data scale. Over the past few decades, the scale of ANN network parameters has grown from megabytes (MB) to gigabytes (GB). For example, chatGPT-3 has about 175 billion parameters. Traditional processors are difficult to meet such huge computing needs, and a more efficient computing architecture is urgently needed.

[0003] In the post-Moore era, it is increasingly difficult to improve chip manufacturing processes, and advanced processes are expensive. The focus of neural network computing optimization has gradually shifted to exploring new computing methods. Research and experiments have shown that in current important applications of neural networks such as pattern recognition and computer vision, nearly 83% of the calculations are fault-tolerant calculations, which makes it feasible to use approximate calculations instead of precise calculations, thereby greatly improving energy efficiency under certain accuracy constraints.

[0004] At present, most neural network structures (such as convolutional neural networks) have a certain degree of computational tolerance, and the efficiency can be improved under certain precision constraints through approximate computing methods. However, different neural network models have different computing requirements, and the network structure is becoming increasingly complex, including unconventional computing parts such as dropout and residual networks. The optimization space for approximate computing is huge, manual analysis is time-consuming and labor-intensive, and it is often impossible to obtain the optimal design based on experience alone.

[0005] This technical solution aims to develop an automated design platform for approximate computing units that perceive the characteristics of neural network applications, which can generate optimal / relatively optimal approximate computing units under certain precision constraints based on the target neural network computing structure and application scenarios. It can efficiently optimize the approximate computing of neural network operations and improve the energy efficiency of artificial neural network applications under certain precision constraints. Summary of the invention

[0006] In view of this, the present invention provides a method and system for designing an approximate computing unit based on neural network application feature perception, so as to generate an optimal approximate computing unit.

[0007] In order to achieve the above object, the present invention adopts the following technical solution:

[0008] A design method for an approximate computing unit with feature perception in neural network applications, comprising the following steps:

[0009] Construct an approximate computing unit library;

[0010] Obtain corresponding feature vectors according to the quantized data after the execution of the topological structures of each layer of the target neural network;

[0011] Input the obtained feature vectors into the trained PPO algorithm model, and obtain approximate computing units for the execution of the topological structures of each layer of the target neural network from the approximate computing unit library.

[0012] Further, in the step of constructing the approximate computing unit library:

[0013] The approximate computing unit library includes: approximate adder types, bit widths, and dynamic saturation thresholds.

[0014] Further, the approximate adder types include:

[0015] LOA adder, ACA adder, ETA fault-tolerant adder, CSA carry-skip adder, and SCSA carry prediction selection adder.

[0016] Further, among the LOA adder, ACA adder, ETA fault-tolerant adder, CSA carry-skip adder, and SCSA carry prediction selection adder, any one adder includes:

[0017] Dynamic bit width adjustment circuit and overflow handling unit;

[0018] Among them, the dynamic bit width adjustment circuit is implemented by a multiplexer and a configurable carry chain structure, and supports 4 / 8 / 12 / 16-bit mode switching;

[0019] The overflow handling unit is implemented by a sign extension circuit, used to execute dynamic saturation logic, and can automatically configure the overflow truncation threshold according to the dynamic range of the current layer parameters of the neural network.

[0020] Further, obtaining corresponding feature vectors according to the quantized data after the execution of the topological structures of each layer of the target neural network specifically includes:

[0021] Statistical zero-value ratio and parameter dynamic range after the execution of the features of each layer of the neural network;

[0022] Combine the numerical zero-value ratio and parameter dynamic range into a feature vector.

[0023] Further, the PPO algorithm model is trained through the following steps:

[0024] Collect historical data:

[0025] Extract the feature vectors of each layer of the known neural network as the state of the agent;

[0026] Take the approximate adder type, bit width, and dynamic saturation threshold in the approximate computing unit library as the actions of the agent;

[0027] Construct the reward function R of the agent: R = α·A1 + β·A2 + γ·A3;

[0028] Among them, α, β, and γ all represent the set corresponding weight coefficients; A1 represents the accuracy loss after using the approximate computing unit to replace the neural network structure; A2 represents the power consumption of the approximate computing unit; A3 represents the delay of the approximate computing unit;

[0029] Based on the reward function, perform iterative optimization to obtain the optimal action corresponding to the state.

[0030] Furthermore, it further includes: adjusting the weight coefficients of the reward function according to the state value of the agent.

[0031] The present invention also discloses an approximate computing unit design system for neural network application feature perception, including:

[0032] An approximate computing unit library for storing approximate computing units, including approximate adder type, bit width, and dynamic saturation threshold;

[0033] A feature vector acquisition unit for obtaining corresponding feature vectors according to the quantization data after the execution of the topological structure of each layer of the target neural network;

[0034] A PPO algorithm model for receiving the feature vector and using the trained PPO algorithm model to obtain the approximate computing unit for the execution of each layer of the topological structure of the target neural network from the approximate computing unit library.

[0035] Preferably, in the above system, the PPO algorithm model is trained through the following steps:

[0036] Collect historical data:

[0037] Extract the feature vectors of each layer of the known neural network as the state of the agent;

[0038] Take the approximate adder type, bit width, and dynamic saturation threshold in the approximate computing unit library as the actions of the agent;

[0039] Construct the reward function R of the agent: R = α·A1 + β·A2 + γ·A3;

[0040] Among them, α, β, and γ all represent the set corresponding weight coefficients; A1 represents the accuracy loss after using the approximate calculation unit to replace the neural network structure; A2 represents the power consumption of the approximate calculation unit; A3 represents the delay of the approximate calculation unit.

[0041] Based on the above reward function, iterative optimization is performed to obtain the optimal action corresponding to the state.

[0042] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a design method and system for an approximate calculation unit with neural network application feature perception, having the following beneficial effects:

[0043] The present invention automatically explores the approximate policy space by using a reinforcement learning model (PPO algorithm) to generate a Pareto optimal solution set (the search efficiency is increased by 3 times).

[0044] The present invention realizes the hybrid optimization of "coarse-grained layer + fine-grained layer" by hierarchically and dynamically adjusting the approximation degree in combination with layer features (such as ReLU sparsity, parameter distribution).

[0045] The present invention reduces the accuracy loss caused by overflow (the experiment shows that the error is reduced by 37%) by introducing a dynamic saturation logic and a sign extension circuit and adaptively adjusting according to the numerical range within the layer.

[0046] The present invention improves the optimization efficiency (the annotation time is reduced by 60%) by automatically annotating the data set based on the activation characteristics and dynamically changing the weight coefficients in the quantization function according to the proportion of negative values. Brief Description of the Drawings

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0048] Figure 1 It is a schematic diagram of the overall process of the method provided by the present invention.

[0049] Figure 2 It is a schematic diagram of the complete process of another method provided by the embodiments of the present invention. Detailed Embodiments

[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0051] The embodiments of the present invention first disclose a design method for an approximate computing unit for neural network application feature perception, as Figure 1 shown, including the following steps:

[0052] Construct an approximate computing unit library;

[0053] According to the quantized data after the execution of the topological structures of each layer of the target neural network, obtain the corresponding feature vectors;

[0054] Input the obtained feature vectors into the trained PPO algorithm model, and obtain the approximate computing units for the execution of each layer topological structure of the target neural network from the approximate computing unit library.

[0055] Further, in the step of constructing the approximate computing unit library:

[0056] The approximate computing unit library includes: approximate adder types, bit widths, and dynamic saturation thresholds.

[0057] Further, the approximate adder types include:

[0058] LOA adder, ACA adder, ETA fault-tolerant adder, CSA carry-skip adder, and SCSA carry-prediction selection adder.

[0059] Further, among the LOA adder, ACA adder, ETA fault-tolerant adder, CSA carry-skip adder, and SCSA carry-prediction selection adder, any one adder includes:

[0060] Dynamic bit-width adjustment circuit and overflow processing unit;

[0061] Among them, the dynamic bit-width adjustment circuit is implemented by a multiplexer and a configurable carry chain structure, and supports 4 / 8 / 12 / 16-bit mode switching;

[0062] The overflow processing unit is implemented by a sign extension circuit, is used to execute dynamic saturation logic, and can automatically configure the overflow truncation threshold according to the dynamic range of the current layer parameters of the neural network.

[0063] Further, the features of each layer of the target neural network include the calculation mode features and activation function features of the network structure of each layer of the neural network.

[0064] Further, according to the quantization data after the execution of the topological structures of each layer of the target neural network, corresponding feature vectors are obtained, specifically including:

[0065] Statistically analyze the zero-value proportion and parameter dynamic range after the execution of the features of each layer of the neural network;

[0066] Combine the numerical zero-value proportion and parameter dynamic range into feature vectors.

[0067] Furthermore, the PPO algorithm model is trained through the following steps:

[0068] Collect historical data:

[0069] Extract the feature vectors of each layer of the known neural network as the states of the agent;

[0070] Take the type, bit width, and dynamic saturation threshold of the approximate adder in the approximate computing unit library as the actions of the agent;

[0071] Construct the reward function R of the agent: R = α·A1 + β·A2 + γ·A3;

[0072] Where α, β, and γ all represent the set corresponding weight coefficients; A1 represents the accuracy loss after using the approximate computing unit to replace the neural network structure; A2 represents the power consumption of the approximate computing unit; A3 represents the delay of the approximate computing unit;

[0073] Based on the reward function, perform iterative optimization to obtain the optimal action corresponding to the state.

[0074] Furthermore, it also includes: adjusting the weight coefficients of the reward function according to the state value of the agent.

[0075] The steps and principles of the present invention are further described below.

[0076] Step 1: Feature extraction and analysis

[0077] Input parsing: Receive the topological structure of the target neural network (such as ResNet residual block, VGG convolutional layer, etc.), the type of activation function (ReLU / Sigmoid), and its application scenarios (such as image recognition, natural language processing).

[0078] Dynamic feature extraction: Through static code analysis and dynamic simulation (based on the PyTorch / TensorFlow framework), extract the calculation modes of each layer (such as convolution, fully connected), the distribution of activation functions (sparsity of ReLU), the parameter numerical range, and the error tolerance.

[0079] Key metric quantification: Calculate the negative value ratio, zero value ratio, and parameter dynamic range in each layer, and generate a feature weight matrix in combination with hardware performance constraints (power consumption, area, latency).

[0080] Error tolerance quantification: Calculate the zero value ratio (Zero-Value Ratio, ZVR) and negative value distribution (sparsity after ReLU) of the output of each layer. Formula:

[0081] Negative value tolerance = 1 - ZVR

[0082] Machine learning-based strategy selection: Map the features of each layer of the neural network (such as negative value distribution, activation function type, dynamic parameter range) to feature vectors, input them into a pre-trained reinforcement learning model (such as the PPO algorithm), and output the bit width and approximation degree (such as error tolerance threshold) of the approximate computing unit (such as an adder).

[0083] Hierarchical optimization: Use high-precision computing units (with precision higher than the set threshold) for error-sensitive layers (such as classification layers), and use aggressive approximation schemes (such as truncating low-order operations) for fault-tolerant layers (such as shallow convolutions).

[0084] Hardware performance constraint modeling: Generate a feature weight matrix in combination with the number of LUTs, power consumption, and latency of each layer. For example, prefer low-power approximation units for convolutional layers and high-precision units (with precision higher than the set threshold) for classification layers.

[0085] Hardware implementation and iterative optimization: Configurable approximate unit library: Implement 5 variable-bitwidth adders (such as LOA, ACA) based on Verilog, support dynamic bit width adjustment (4 / 8 / 12 / 16 bits), and overflow handling strategies (such as saturation truncation, sign extension).

[0086] Synthesis and evaluation: Use Vivado for logic synthesis to generate reports on the number of LUTs, power consumption, and latency; overload neural network operators (such as `+`) in C++ and replace standard computing units in the simulation environment to verify the accuracy loss (such as the Top-1 accuracy drop ≤ 1%).

[0087] Feedback loop: Adjust the approximate strategy weights according to the hardware performance and accuracy evaluation results, and regenerate the optimization plan until the constraint conditions are met.

[0088] Intelligent annotation and quantification method for datasets, data generation: Pre-train VGG16 / ResNet18 on CIFAR-10 / ImageNet, perform forward inference on networks such as VGG16 and ResNet18 on the CIFAR-10 / ImageNet dataset, and record the numerical distributions of the inputs and outputs of each layer.

[0089] Dynamic annotation: According to the proportion of zero values after ReLU activation (for example, 80% of the output of a certain layer is zero), automatically reduce the precision weight of this layer and preferentially allocate low-power approximation units. Generate annotation data through simulation (such as "select the LOA-8-bit adder when the ZVR of a certain layer = 80%") to construct a training set.

[0090] Quantization function design: Optimization objective = α · precision loss + β · power consumption + γ · delay

[0091] Among them, the weight coefficients (α, β, γ) are dynamically adjusted according to the proportion of negative values (when the proportion of negative values is high, α decreases, and β, γ increase).

[0092] The adders involved in the present invention include LOA adder, ACA adder, ETA fault-tolerant adder, CSA carry-skip adder, and SCSA carry-prediction selection adder. Each adder will be described separately below:

[0093] LOA (Lower-Part Omission Adder): It uses segmented operation technology. The low-bit operation depends on the approximate value generated by an OR gate, while the high-bit operation uses an accurate calculation module to process the high and low bits of the data separately.

[0094] ACA (Approximate Carry Adder): Reduces the critical path delay through probabilistic carry prediction and integrates an overflow detection module (such as forcing saturation to the maximum value when out of range).

[0095] ETA fault-tolerant adder: The high bits are accurately calculated, and the low bits are approximately processed. Using the idea of block division, it consists of a carry generator, an addend and generator. The carry signal generated by the low-bit carry generator is transmitted to the addend and generator of the adjacent high bit.

[0096] CSA carry-skip adder: Utilizes the idea of modularization to generate the carry and partial sum through sub-modules and uses a carry-skip mechanism

[0097] SCSA carry-prediction selection adder: The architecture divides the adder into several unit modules. Each unit realizes the result estimation and carry selection through a configurable window adder. This window adder adopts a dual-path parallel computing mechanism to predict the operation results when the carry is "1" and "0" respectively, and then selects and outputs according to the previous carry signal.

[0098] Each adder has a dynamic bit-width adjustment circuit and an overflow processing unit:

[0099] Dynamic bit-width adjustment circuit:

[0100] Adopt a multiplexer (MUX) and a configurable carry chain structure to support 4 / 8 / 12 / 16-bit mode switching. For example, the LOA adder achieves accurate high-bit calculation by discarding the low bits, and the dynamic bit width is selected by a control signal (such as bit_width_ctrl[2:0]):

[0101] Overflow handling unit:

[0102] Dynamic saturation logic: Automatically configure the overflow truncation threshold according to the current layer parameter range (such as [-128, 127]). For example, when the parameter range is [-128, 127], the overflow truncation threshold is dynamically adjusted to ±128, which is implemented through a comparator and a multiplexer:

[0103] Sign extension circuit: For signed number operations, extend the high-order sign bit to preserve the numerical semantics.

[0104] In the present invention, automation is achieved through the following system architecture:

[0105] Front-end parser: Support ONNX / TensorFlow model parsing and extract the inter-layer dependency relationships.

[0106] Optimization engine: A policy generator based on reinforcement learning (deployed in TensorFlow Serving), which outputs approximate unit configuration parameters.

[0107] Hardware generator: According to the configuration parameters, call the Verilog template library to generate RTL code and integrate the dynamic bit width control module (such as dynamic_bitwidth_ctrl.v).

[0108] The reinforcement learning model (PPO algorithm) includes:

[0109] State space: Feature vectors of each layer (ZVR, dynamic range, error tolerance).

[0110] Action space: Approximate adder types (such as LOA / ACA, etc.), bit width, overflow strategy.

[0111] Reward function:

[0112] R = α · precision loss + β · power consumption + γ · delay;

[0113] Among them, the weight coefficients α, β, γ are dynamically adjusted according to ZVR (reduce α when ZVR is high, increase β, γ)

[0114] In the prior art, a static approximation strategy (such as a globally fixed bit width) is usually adopted, ignoring the differences between layers of the neural network; the present technical solution combines layer features (such as ReLU sparsity, parameter distribution) to dynamically adjust the approximation degree layer by layer, realizing a hybrid optimization of "coarse-grained layer + fine-grained layer".

[0115] In the prior art, overflow handling is mostly in a fixed mode (such as direct truncation), resulting in error accumulation. In this technical solution, a dynamic saturation logic and sign extension circuit are introduced, which are adaptively adjusted according to the in-layer numerical range to reduce the precision loss caused by overflow (experiments show that the error is reduced by 37%).

[0116] In the prior art, setting the precision-performance weight depends on manual experience. This technical solution can automatically label the dataset based on the ReLU activation characteristics, and the weight coefficients in the quantization function change dynamically with the proportion of negative values, improving the optimization efficiency (the labeling time is reduced by 60%).

[0117] In the prior art, the design of approximate units depends on exhaustive search or heuristic rules. This technical solution uses a reinforcement learning model (PPO algorithm) to automatically explore the approximate policy space and generate a Pareto optimal solution set (the search efficiency is increased by 3 times).

[0118] The whole of the present invention can realize the dynamic optimization of the neural network approximate computing unit according to the closed-loop process of "feature perception → policy generation → hardware implementation → feedback optimization". The complete process can be referred to Figure 2 The following are the detailed operation steps:

[0119] 1. Feature extraction and quantization analysis

[0120] Step 1.1: Input parsing and dynamic feature extraction

[0121] (1) First, input a target neural network: receive the topological structure (residual block configuration, convolution kernel size), activation function type (ReLU), and hardware constraints (power consumption ≤ 5W, delay ≤ 50ns) of the target neural network (such as ResNet50).

[0122] (2) Extract features from the target neural network:

[0123] Static analysis: Parse the network code (PyTorch / TensorFlow), and extract the calculation mode (convolution / full connection) of each layer and the activation function distribution (such as the zero-value ratio ZVR after ReLU).

[0124] Dynamic simulation: Run inference on the CIFAR-10 dataset, and record the parameter value range of each layer (such as the parameter range [-64, 63] of the 3rd residual block) and the output distribution (ZVR = 75%).

[0125] (3) Quantize the key indicators of the neural network:

[0126] Calculate the error tolerance of each layer (such as the tolerance of the classification layer ≤ 0.5%), and generate a feature vector:

[0127] Feature vector = [ZVR, parameter range, error tolerance]

[0128] Example: The feature vector of the 3rd residual block of ResNet50 is [75%, [-64, 63], 1.2%].

[0129] 2. Hardware Implementation and Dynamic Adjustment

[0130] Step 2.1: Configurable Unit Library Integration

[0131] (1) First, construct a variable-bitwidth adder: Implement the LOA / ACA adder based on Verilog, supporting dynamic switching of "4 / 8 / 12 / 16" bits (control signal bit_width_ctrl[2:0]).

[0132] Example: The LOA-8-bit adder discards the lower 4 bits, and the hardware area is reduced by 42% (from 1200 LUTs → 696 LUTs).

[0133] (2) Second, add an overflow handling unit to the adder: Include a dynamic saturation logic: Automatically configure the threshold (±64) according to the parameter range (such as [-64, 63]).

[0134] Sign extension circuit: When the dynamic range of the calculated parameter is greater than the set threshold, extend the sign bit to reduce semantic errors (experiments show that the error accumulation is reduced by 45%).

[0135] Step 2.2: Synthesis and Performance Evaluation

[0136] (1) Perform logic synthesis on the constructed adder: Use Vivado to synthesize the LOA-8-bit adder and generate a report:

[0137] Power consumption: 67 μW (original 105 μW), Delay: 0.39 ns (original 0.57 ns).

[0138] (2) Then verify the accuracy of the adder: Replace the standard adder in the C++ simulation environment and test the classification accuracy of ResNet50: The Top-1 accuracy drops by 0.7% (from 76.2% → 75.9%), meeting the ≤1% constraint.

[0139] 3. Intelligent Generation of Approximation Strategies (Based on the Reinforcement Learning PPO Algorithm)

[0140] Step 3.1: Policy Space Definition and Reward Function Modeling

[0141] Input the feature vectors of each layer of the neural network (such as ZVR, parameter range, negative value ratio).

[0142] Use the following reward function to make decisions to select a suitable approximate adder:

[0143] R = α·Precision loss + β·Power consumption + γ·Latency;

[0144] Specific example: When ZVR > 70%, reduce the precision weight (α = 0.3), increase the power consumption weight (β = 0.5) and the latency weight (γ = 0.2).

[0145] Select the adder type (such as LOA / ACA / CSA), bit width (4 / 8 / 12 / 16 bits), and overflow strategy (dynamic saturation threshold).

[0146] Step 3.2: Generation of hierarchical optimization strategy

[0147] For high ZVR layers (such as shallow convolution): We select the LOA-8-bit adder (discarding the lower bits), and set the dynamic saturation threshold to ±32.

[0148] Data support: Experiments show that when ZVR = 80%, the LOA-8-bit can reduce the power consumption by 36% (the power consumption of a single approximate computing unit is reduced from 105 μW to 67 μW), and the Top-1 precision loss is only 0.3%.

[0149] For low error tolerance layers (such as the classification layer): We select the CSA-16-bit adder (set the dynamic saturation threshold to ±64).

[0150] Data support: After the classification layer adopts the CSA-16-bit, the Top-1 precision loss ≤ 0.5%, meeting the constraint conditions.

[0151] Step 3.3: Generation of Pareto optimal solution set: Output multiple sets of configuration schemes (such as LOA-8-bit + threshold ±32 vs. ACA-12-bit + threshold ±64), balance precision and energy efficiency, and find the optimal solution as the output result. At the same time, this input-output result can also be added to the training set as data for optimizing the model.

[0152] 4. Feedback closed-loop and iterative optimization

[0153] Step 4.1: Performance-precision feedback

[0154] If the precision loss exceeds the limit (such as the Top-1 drops by 1.5%), then perform the following operations:

[0155] (1) Adjust the reward function weights:

[0156] Increase α (precision first), reduce β (power consumption tolerance) and γ (latency tolerance).

[0157] (2) Regenerate the strategy: Select a higher bit width (such as LOA-12-bit) or a more conservative overflow threshold (±64).

[0158] Step 4.2: Iterative optimization result: Mark the input-output results as a dataset for training the automatic selection model, and continuously iterate to optimize and upgrade the model to achieve results with less accuracy loss and smaller power consumption and area. Repeat this process multiple times.

[0159] Example: AlexNet fully connected layer (parameter range [-64, 63], negative value ratio 65%):

[0160] Initial strategy: LOA-8 bits + threshold ±32 → error accumulation 3.2%.

[0161] Optimized strategy: LOA-12 bits + threshold ±64 → error accumulation reduced to 1.2%, and power consumption still reduced by 28%.

[0162] Specific optimization examples are as follows:

[0163] In the 3rd residual block of ResNet50 (ZVR = 75%), the present invention automatically selects an ACA-12 bit adder, sets the dynamic saturation threshold to ±32, and after synthesis, the power consumption is reduced by 18%, and the Top-5 accuracy loss is only 0.2%.

[0164] For the fully connected layer of AlexNet (parameter range [-64, 63]), the present invention adopts a sign extension circuit + dynamic saturation (threshold ±64), and compared with the fixed truncation scheme, the error accumulation is reduced by 45%.

[0165] In the third convolutional layer of ResNet18 (negative value ratio 70%), the present invention automatically selects an LOA-8 bit adder, and uses dynamic saturation (threshold = ±64) for overflow handling. After synthesis, the power consumption is reduced by 16%, and the Top-1 accuracy only drops by 0.3%.

[0166] The above technical solutions significantly improve the balance between the efficiency and accuracy of approximate computing through feature perception, dynamic optimization, and machine learning methods, solve the limitations of the "one-size-fits-all" optimization in the prior art, and have clear patent protection value.

[0167] Each embodiment in this specification is described in a progressive manner. The key points of each embodiment are the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference can be made to the method part for the relevant parts.

[0168] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A design method for an approximate computing unit with neural network application feature perception, characterized in that It includes the following steps: Construct an approximate computing unit library; Obtain corresponding feature vectors according to the quantized data after the execution of the topological structures of each layer of the target neural network; Input the obtained feature vectors into the trained PPO algorithm model, and obtain approximate computing units for the execution of the topological structure of each layer of the target neural network from the approximate computing unit library.

2. The design method of an approximate computing unit for neural network application feature perception according to claim 1, characterized in that In the step of constructing the approximate computing unit library: The approximate computing unit library for storing approximate computing units includes: approximate adder type, bit width, and dynamic saturation threshold.

3. The design method of an approximate computing unit for neural network application feature perception according to claim 2, characterized in that The approximate adder types include: LOA adder, ACA adder, ETA fault-tolerant adder, CSA carry-skip adder, and SCSA carry-prediction selection adder.

4. A design method of an approximate computing unit for neural network application feature perception according to claim 3, characterized in that Among the LOA adder, ACA adder, ETA fault-tolerant adder, CSA carry-skip adder, and SCSA carry-prediction selection adder, any one adder includes: Dynamic bit width adjustment circuit and overflow processing unit; Among them, the dynamic bit width adjustment circuit is implemented by a multiplexer and a configurable carry chain structure, and supports 4 / 8 / 12 / 16-bit mode switching; The overflow processing unit is implemented by a sign extension circuit, used to execute dynamic saturation logic, and can automatically configure the overflow truncation threshold according to the dynamic range of the current layer parameters of the neural network.

5. A design method for an approximate computing unit for neural network application feature perception according to claim 1, characterized in that, Obtaining corresponding feature vectors according to the quantized data after the execution of the features of each layer of the neural network specifically includes: Statistical zero value ratio and parameter dynamic range after the execution of the topological structures of each layer of the neural network; Combine the numerical zero value ratio and parameter dynamic range into a feature vector.

6. The design method of an approximate computing unit for neural network application feature perception according to claim 1, characterized in that The PPO algorithm model is trained through the following steps: Collect historical data: Extract the feature vectors of each layer of the known neural network as the state of the agent; Take the approximate adder type, bit width, and dynamic saturation threshold in the approximate computing unit library as the actions of the agent; Construct the reward function R of the agent: R = α·A1 + β·A2 + γ·A3; Among them, α, β, and γ all represent the set corresponding weight coefficients; A1 represents the accuracy loss after using the approximate computing unit to replace the execution of the neural network structure; A2 represents the power consumption of the approximate computing unit; A3 represents the delay of the approximate computing unit; Based on the reward function, perform iterative optimization to obtain the optimal action corresponding to the state.

7. A design method of an approximate computing unit for neural network application feature perception according to claim 6, characterized in that It also includes: Adjust the weight coefficients of the reward function according to the state value of the agent during the iterative training process.

8. A design system for an approximate computing unit with neural network application feature perception, characterized in that, It includes: An approximate computing unit library, which is used to store approximate computing units, including approximate adder type, bit width, and dynamic saturation threshold; A feature vector acquisition unit, which is used to obtain corresponding feature vectors according to the quantized data after the execution of the topological structures of each layer of the neural network; A PPO algorithm model, which is used to receive the feature vectors and obtain approximate computing units for the execution of the topological structure of each layer of the target neural network from the approximate computing unit library by using the trained PPO algorithm model.

9. The approximate computing unit design system for neural network application feature perception according to claim 8, characterized in that, The PPO algorithm model is trained through the following steps: Collect historical data: Extract the feature vectors of each layer of the known neural network as the state of the agent; Take the approximate adder type, bit width, and dynamic saturation threshold in the approximate computing unit library as the actions of the agent; Construct the reward function R of the agent: R = α·A1 + β·A2 + γ·A3; Among them, α, β, and γ all represent the set corresponding weight coefficients; A1 represents the accuracy loss after using the approximate computing unit to replace the neural network structure; A2 represents the power consumption of the approximate computing unit; A3 represents the delay of the approximate computing unit; Based on the reward function, perform iterative optimization to obtain the optimal action corresponding to the state.