Multistage compression collaborative optimization neural network deployment method and device based on memristor and storage medium
Through the combined optimization methods of dynamic gradient sensitive pruning, multi-stage quantization perception training and hardware-aware knowledge distillation, the problems caused by insufficient storage density and non-ideal characteristics in the memristor memory and computing architecture are solved, and efficient and reliable neural network deployment is achieved, which improves hardware utilization and computing energy efficiency.
Patent Information
- Application Number
- CN202510458084.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-22
AI Technical Summary
In the integrated memristor memory architecture, the quantization error caused by insufficient storage density, non-ideal characteristics, and the lack of algorithm and hardware coordination have made it difficult to take into account both accuracy and energy efficiency during deployment.
A combined optimization method of dynamic gradient sensitive pruning, multi-stage quantization perception training and hardware-aware knowledge distillation is adopted to generate a sparse weight structure through a dual-drive scoring mechanism, combining nonlinear conductance modeling and noise injection, network weighting and conductance mapping parameters are optimized, and the robustness of the model in a noisy environment is improved.
It significantly improves the efficiency and reliability of memristor hardware deployment, realizes high storage utilization, nonlinear quantization error suppression and robustness in noise environments, reduces power consumption, and supports reliable inference under low-bit quantization.
Smart Images

Figure CN120354904A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence hardware acceleration, and particularly relates to a method, device, and storage medium for multi-level compression collaborative optimization neural network deployment based on memristors. Background Art
[0002] With the rapid development of fields such as big data, artificial intelligence, and the Internet of Things, traditional von Neumann architecture chips have gradually been unable to meet the current high-concurrency and low-latency intelligent computing requirements due to the limitations of the "power consumption wall" and "memory wall". In the traditional memory-computation separation architecture, frequent data transfer consumes a large amount of energy, severely restricting the computing power energy efficiency. To solve this fundamental bottleneck, research on memory-computation integrated architecture chips has begun, and among them, neural network chips based on memristors have stood out due to the characteristics of the deep integration of non-volatile storage and analog multiplication-addition computing. Compared with the traditional architecture, memristor chips exhibit three revolutionary advantages: an order-of-magnitude improvement in energy efficiency, a breakthrough increase in storage density, and large-scale parallel computing capabilities. However, currently, neural network chips based on memristors are still in the research stage, and there are still some problems before full commercial application:
[0003] (1) The storage density is limited while the model scale is huge. Although multi-array stacking can expand the capacity, it leads to a significant increase in interconnection latency and power consumption.
[0004] (2) Memristors have non-ideal characteristics such as conductance drift, read-write noise, and non-linear mapping. Traditional uniform quantization schemes (such as linear weight-conductance mapping) will amplify hardware errors. For example, when the weights are mapped to discrete conductance states, the non-linear conductance response leads to the accumulation of quantization errors. In actual measurements, the quantization error of 4 bits can reach 8% - 12%, severely reducing the model inference accuracy.
[0005] (3) Lack of algorithm-hardware collaboration: Existing methods mostly adopt isolated optimizations of techniques such as pruning, quantization, and distillation, lacking systematic collaboration. For example, directly quantizing after pruning may exacerbate the mapping error due to the distortion of the weight distribution, while distillation training that ignores hardware noise is difficult to ensure the robustness of actual deployment. This fragmented optimization makes it difficult to achieve both the compression rate at the algorithm level and the energy efficiency improvement at the hardware level for the model. Summary of the Invention
[0006] The technical problem to be solved by the present invention is: in the field of artificial intelligence hardware acceleration, how to overcome the problems of insufficient storage density, quantization errors caused by non-ideal characteristics, and lack of algorithm-hardware collaboration in the memristor memory-computation integrated architecture, and through the joint optimization method of dynamic pruning, non-linear quantization, and hardware-aware distillation, achieve high energy efficiency, high precision, and strong robustness deployment of the neural network model on memristor hardware.
[0007] To achieve the above object, the present invention adopts the following technical means:
[0008] The present invention provides a lightweight neural network deployment method based on the collaborative optimization of a memristor in-memory computing architecture and multi-level compression, including the following steps:
[0009] (1) Dynamic gradient-sensitive pruning: Based on a dual-drive scoring mechanism (L1 norm scoring and gradient sensitivity scoring) and hardware mapping efficiency feedback, perform progressive structured pruning on the pre-trained neural network model to generate a hardware-friendly sparse weight structure;
[0010] (2) Multi-level quantization-aware training: Simulate the memristor conductance characteristics through a non-linear conductance modeling function, perform mixed-precision quantization encoding on the pruned model, and synchronously optimize the network weights and conductance mapping parameters to reduce quantization errors;
[0011] (3) Hardware-aware knowledge distillation: Inject memristor non-ideal characteristic noise into the vectorized model, and improve the robustness of the model in a noisy environment through anti-interference distillation training to obtain an optimized lightweight neural network model;
[0012] (4) Map the optimized lightweight neural network model to a memristor in-memory computing architecture, perform weight encoding, array mapping, and in-memory computing inference, and output a deployable lightweight neural network model.
[0013] In the above solution, the specific steps of the dynamic gradient-sensitive pruning include:
[0014] 1.1) Load the pre-trained floating-point neural network model and extract the weight matrices of each convolutional layer;
[0015] 1.2) Input the training data to perform forward propagation and calculate the cross-entropy loss;
[0016] 1.3) For each element in the weight matrix, calculate its L1 norm score and gradient sensitivity score. The L1 norm score is the normalized value of the absolute value of the weight relative to the L1 norm of the weight matrix of the layer where it is located, and the gradient sensitivity score is the normalized value of the absolute value of the weight gradient relative to the maximum gradient of the layer where it is located;
[0017] 1.4) Generate a comprehensive score by weighted fusion of the L1 norm score and the gradient sensitivity score, and sort the weights according to the comprehensive score;
[0018] 1.5) Adopt a progressive pruning scheduling strategy, gradually increase the pruning rate until the target compression rate is reached, and restore the model accuracy through fine-tuning after each round of pruning;
[0019] 1.6) Based on the hardware mapping efficiency feedback of the memristor array, perform adaptive adjustment on the pruned sparse weight structure, and preferentially perform convolutional kernel-level or channel-level structured pruning to ensure that the hardware utilization rate is higher than the preset threshold;
[0020] 1.7) Generate a binary mask according to the dynamic scoring threshold, where a mask value of 1 indicates retaining the weight and 0 indicates pruning, and complete model sparsification through the combined action of the mask and gradient update.
[0021] In the above solution, the multi-level quantization-aware training includes the following steps:
[0022] 2.1) Define a differentiable non-linear conductance modeling function G(w) = a·tanh(b·w) + c, where a is the conductance range scaling factor, b is the slope factor for adjusting the non-linear saturation rate, and c is the conductance reference offset. Determine the initial parameter values by fitting the measured IV curve of the memristor using non-linear least squares.
[0023] 2.2) Adopt a mixed-precision quantization strategy. Perform 4-bit quantization on the convolutional layer weights to retain 16 conductance states, and perform 2-bit quantization on the fully connected layer weights to compress to 4 conductance states. Achieve forward propagation quantization through the differentiable quantization function Q(w):
[0024]
[0025] w: Neural network weight parameter;
[0026] q min / q max : The minimum / maximum value of the quantization dynamic range, used to limit the truncation boundary of the weight value;
[0027] Δ: Quantization step size;
[0028] clip(w, q min , q max ): Truncation function, which limits the weight w within the interval (q min q max );
[0029] round(·): Rounding operation, which maps continuous values to the nearest quantization level;
[0030] 2.3) Apply the gradient approximation rule. Keep the gradient directly backpropagated within the quantization interval and set the gradient to zero outside the interval to eliminate noise interference. Synchronously optimize the network weights and the conductance mapping parameters {a, b, c};
[0031] 2.4) Map the quantized weights to the simulated memristor array every 10 training cycles, and calculate the conductance mapping error ε map , if ε map > 5%, re-initialize {a, b, c} and roll back the model to the previous checkpoint.
[0032] In the above solution, the hardware-aware knowledge distillation includes the following steps:
[0033] 3.1) Model the non-ideal characteristics of memristors, including the conductance drift model and the read / write noise model, where:
[0034] The conductance drift model is:
[0035]
[0036] where \(G_0\) is the initial conductance value of the quantization weight mapping, \(\sigma\) d is the drift coefficient based on the 1T1R array, \(t\) is the time step, and \(N(0, 1)\) is the standard Gaussian distribution;
[0037] The read / write noise model is:
[0038] \(G\) read \( = G+\sigma\) r \(\cdot N(0, 1)\)
[0039] where \(\sigma\) r is the noise intensity, and \(G\) represents the original conductance value of the memristor array.
[0040] 3.2) Dynamic noise injection and anti-interference training:
[0041] In each batch of training, randomly sample the time \(t\sim Uniform(0, T\) max ) and the noise intensity \(\sigma'\) r \(\sim N(\sigma\) r , 0.02), generate the composite perturbation conductance value \(G\) perturbed , and use the perturbed conductance value to perform forward propagation to obtain the output \(y\) student of the student model:
[0042] \(G\) perturbed \( = G(t)+\sigma\) r \(\cdot N(0, 1)\)
[0043] where \(\sigma'\) r is the dynamic noise intensity;
[0044] 3.3) Adaptive noise enhancement control:
[0045] Every 10 training epochs, increase the noise intensity \(\sigma\) r by 10%, and monitor the accuracy on the validation set. If the accuracy drops by more than 2%, roll back to the previous stage parameters and terminate the enhancement;
[0046] 3.4) Design a dual-loss joint optimization strategy. Through the KL divergence loss \(L\) KL force the output distribution of the student model in the noise environment to approximate that of the teacher model, and at the same time retain the cross-entropy loss \(L\) task to maintain the basic classification ability. The total loss function is:
[0047]
[0048] where λ1, λ2, and λ3 are adjustable weighting coefficients, is the L2 regularization term;
[0049] 3.5) Freeze the parameters of the teacher model, and only update the weights and quantization parameters of the student model through backpropagation of gradients until the model converges in the noisy environment.
[0050] The present invention provides a multi-level compression collaborative optimization neural network deployment device based on memristors, including the following modules:
[0051] Dynamic gradient-sensitive pruning module: Based on a dual-drive scoring mechanism (L1 norm scoring and gradient sensitivity scoring) and hardware mapping efficiency feedback, perform progressive structured pruning on the pre-trained neural network model to generate a hardware-friendly sparse weight structure;
[0052] Multi-level quantization-aware training module: Simulate the conductance characteristics of memristors through a non-linear conductance modeling function, perform mixed-precision quantization encoding on the pruned model, and synchronously optimize the network weights and conductance mapping parameters to reduce quantization errors;
[0053] Hardware-aware knowledge distillation module: Inject memristor non-ideal characteristic noise into the quantized model, and improve the robustness of the model in the noisy environment through anti-interference distillation training to obtain an optimized lightweight neural network model;
[0054] Map the optimized lightweight neural network model to a memristor in-memory computing architecture, perform weight encoding, array mapping, and in-memory computing inference, and output a deployable lightweight neural network model.
[0055] In the above solution, the specific steps of the dynamic gradient-sensitive pruning module include:
[0056] 1.1) Load the pre-trained floating-point neural network model and extract the weight matrices of each convolutional layer;
[0057] 1.2) Input the training data to perform forward propagation and calculate the cross-entropy loss;
[0058] 1.3) For each element in the weight matrix, calculate its L1 norm score and gradient sensitivity score. The L1 norm score is the normalized value of the absolute value of the weight relative to the L1 norm of the weight matrix of the layer where it is located, and the gradient sensitivity score is the normalized value of the absolute value of the weight gradient relative to the maximum gradient of the layer where it is located;
[0059] 1.4) Generate a comprehensive score by weighted fusion of the L1 norm score and the gradient sensitivity score, and sort the weights according to the comprehensive score;
[0060] 1.5) Adopt a progressive pruning scheduling strategy to gradually increase the pruning rate until the target compression rate is reached. After each round of pruning, fine-tune to restore the model accuracy.
[0061] 1.6) Based on the hardware mapping efficiency feedback of the memristor array, adaptively adjust the pruned sparse weight structure, and preferentially perform convolutional kernel-level or channel-level structured pruning to ensure that the hardware utilization rate is higher than the preset threshold.
[0062] 1.7) Generate a binary mask according to the dynamic scoring threshold. A mask value of 1 indicates retaining the weight, and 0 indicates pruning. Complete model sparsification through the combined action of the mask and gradient update.
[0063] In the above scheme, the multi-level quantization-aware training module includes the following steps:
[0064] 2.1) Define a differentiable non-linear conductance modeling function G(w) = a·tanh(b·w) + c, where a is the conductance range scaling factor, b is the slope factor for adjusting the non-linear saturation rate, and c is the conductance reference offset. Determine the initial parameter values by fitting the measured IV curve of the memristor using non-linear least squares method.
[0065] 2.2) Adopt a mixed-precision quantization strategy. Perform 4-bit quantization on the convolutional layer weights to retain 16 conductance states, and perform 2-bit quantization on the fully connected layer weights to compress to 4 conductance states. Implement forward propagation quantization through the differentiable quantization function Q(w):
[0066]
[0067] w: Neural network weight parameter;
[0068] q min / q max : The minimum / maximum value of the quantization dynamic range, used to limit the truncation boundary of the weight value;
[0069] Δ: Quantization step size;
[0070] clip(w, q min , q max ): Truncation function, which limits the weight w within the interval (q min , q max );
[0071] round(·): Rounding operation, which maps continuous values to the nearest quantization level;
[0072] 2.3) Apply the gradient approximation rule. Keep the gradient directly backpropagated within the quantization interval, and set the gradient to zero outside the interval to eliminate noise interference. Synchronously optimize the network weights and conductance mapping parameters {a, b, c}.
[0073] 2.4) Map the quantized weights to the simulated memristor array every 10 training cycles, and calculate the conductance mapping error ε map , if ε map > 5%, re-initialize {a, b, c} and roll back the model to the previous checkpoint.
[0074] In the above solution, the hardware-aware knowledge distillation module includes the following steps:
[0075] 3.1) Model the non-ideal characteristics of memristors, including the conductance drift model and the read / write noise model, where:
[0076] The conductance drift model is:
[0077]
[0078] where G0 is the initial conductance value mapped by the quantized weights, σ d is the drift coefficient based on the 1T1R array, t is the time step, and N(0, 1) is the standard Gaussian distribution;
[0079] The read / write noise model is:
[0080] G read = G + σ r ·N(0, 1)
[0081] where σ r is the noise intensity, and G represents the original conductance value of the memristor array.
[0082] 3.2) Dynamic noise injection and anti-interference training:
[0083] Randomly sample the time t ~ Uniform(0, T max ) and the noise intensity σ′ r ~ N(σ r , 0.02) in each batch of training to generate the composite perturbation conductance value G perturbed , and use the perturbation conductance value to perform forward propagation to obtain the output y student of the student model:
[0084] G perturbed = G(t) + σ r ·N(0, 1)
[0085] where σ′ r is the dynamic noise intensity;
[0086] 3.3) Adaptive noise enhancement control:
[0087] Every 10 training cycles, increase the mean noise intensity by 10%, and monitor the accuracy on the validation set. If the accuracy drops by more than 2%, roll back to the previous stage parameters and terminate the enhancement;
[0088] 3.4) Design a dual-loss joint optimization strategy. Through the KL divergence loss L KL Force the output distribution of the student model in the noisy environment to approximate that of the teacher model, while retaining the cross-entropy loss L task Maintain the basic classification ability. The total loss function is:
[0089]
[0090] where λ1, λ2, and λ3 are adjustable weighting coefficients, is the L2 regularization term;
[0091] 3.5) Freeze the parameters of the teacher model, and only update the weights and quantization parameters of the student model through gradient backpropagation until the model converges in the noisy environment.
[0092] The present invention also provides a storage medium. When a program in the storage medium is executed by a processor, the described method for deploying a multi-level compression collaborative optimization neural network based on memristors is implemented.
[0093] The lightweight neural network joint optimization deployment method for memristor hardware proposed by the present invention significantly improves the deployment efficiency and reliability of memristor hardware while ensuring the model accuracy through the collaborative optimization of dynamic pruning, multi-level quantization-aware training, and hardware-aware distillation techniques. The specific beneficial effects are as follows:
[0094] I. Hardware-friendly model compression and high storage utilization
[0095] Based on the dynamic pruning strategy of the dual-drive scoring mechanism (L1 norm + gradient sensitivity), combined with the feedback of hardware mapping efficiency, it can adaptively generate structured sparse weights, effectively reducing the number of model parameters. By optimizing the sparse pattern through the convolutional kernel-level and channel-level pruning strategies, the fragmented weight distribution of the memristor array is reduced, significantly improving the utilization rate of hardware storage resources.
[0096] The progressive pruning scheduling and mask fine-tuning mechanism avoid a sudden drop in model accuracy, maintain the task performance while compressing the model scale, and solve the problem that it is difficult to balance the hardware utilization rate and model accuracy in traditional pruning methods.
[0097] II. Nonlinear quantization error suppression and improvement of hardware adaptability
[0098] The conductance characteristics of memristors are simulated through a differentiable conductance mapping function (G(w) = a·tanh(b·w) + c), combined with mixed-precision quantization (4 bits for convolutional layers / 2 bits for fully connected layers), effectively reducing the non-linear error of weight mapping. During the quantization-aware training process, the network parameters and conductance mapping coefficients are optimized synchronously to achieve a deep adaptation between the algorithm and the hardware conductance characteristics.
[0099] The gradient approximation rule for the quantization interval (retaining the gradient within the interval and truncating outside the interval) suppresses noise interference, combined with a periodic hardware simulation verification mechanism (conductance mapping error threshold control), ensuring the compatibility between the quantization model and the memristor array and avoiding accuracy loss caused by non-ideal characteristics.
[0100] III. Enhancement of Robustness and Reliability in a Noisy Environment
[0101] Hardware-aware knowledge distillation injects conductance drift noise G(t) and read / write noise G read to simulate the non-ideal characteristics in the actual operating environment of memristors. The student model undergoes anti-interference training under dynamic noise perturbations and approximates the output distribution of the teacher model through the KL divergence loss, significantly improving the generalization ability of the model in a noisy environment.
[0102] The combined action of the adaptive noise enhancement mechanism (gradually increasing the noise intensity in stages) and the validation set monitoring strategy (accuracy rollback mechanism) balances the anti-interference ability and task performance, avoids overfitting to the noisy scenario, and ensures the stability of the model during long-term operation.
[0103] IV. Algorithm-Hardware Co-Optimization and Energy Efficiency Enhancement
[0104] Through a three-stage closed-loop optimization of pruning, quantization, and distillation, a deep coupling between algorithm compression and hardware constraints is achieved, reducing the sub-optimal problems caused by traditional isolated optimization. After compression, the model can be directly mapped to the memristor array for in-memory computing, avoiding frequent data movement, reducing power consumption, and improving computing energy efficiency.
[0105] The joint optimization scheme significantly reduces the model's dependence on high-precision ADC / DAC, supports reliable inference under low-bit quantization, provides a feasible solution for low-power and high-energy-efficient intelligent computing at the edge, and promotes the practical application process of the in-memory computing architecture.
[0106] In summary, through the collaborative design of algorithm and hardware characteristics, the present invention forms multiple technical advantages in terms of model lightweight, hardware adaptability, and noise robustness, solves the core problems such as insufficient storage density, non-ideal characteristic interference, and algorithm-hardware mismatch in memristor deployment, and provides efficient and reliable technical support for the low-power neural network deployment of edge intelligent devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0107] Figure 1This is the overall flowchart of the lightweight neural network collaborative optimization and deployment method for memristor hardware in the embodiments of the present invention;
[0108] Figure 2 This is a schematic diagram of the memristor crossbar array of the neural network accelerator in the embodiments of the present invention. Detailed implementation manners
[0109] The following will give a detailed description of the embodiments of the present invention. Although the present invention will be described and explained in conjunction with some specific implementation manners, it should be noted that the present invention is not limited to these implementation manners only. On the contrary, any modifications or equivalent replacements made to the present invention should be covered within the scope of the claims of the present invention.
[0110] In addition, in order to better illustrate the present invention, numerous specific details are given in the following detailed implementation manners. Those skilled in the art will understand that the present invention can also be implemented without these specific details.
[0111] In response to the problems raised in the background art, the present invention proposes a dynamic pruning - non - linear quantization - hardware - aware distillation collaborative optimization method, aiming to achieve deep coupling between the algorithm and the hardware through three - stage cascade optimization: First, based on the dual - drive scoring mechanism (L1 norm + gradient sensitivity) and the feedback of hardware mapping efficiency, dynamic pruning is performed to achieve preliminary compression of the model parameter quantity; Subsequently, 4 - bit / 2 - bit mixed - precision quantization is achieved through a differentiable conductance mapping function to reduce quantization errors; Finally, hardware noise is injected and KL - divergence loss is used for anti - interference training to enable the model to maintain accuracy in a noisy environment. Through three - stage closed - loop optimization, this method achieves the global optimum of storage density, computing precision, and anti - interference ability, breaks the limitations of isolated optimization in traditional solutions, and has three major advantages: high energy efficiency, high hardware utilization rate, and strong robustness, providing an efficient neural network deployment solution for edge intelligent devices and promoting the practical process of the memory - in - computing architecture.
[0112] The following further illustrates the technical solutions of the present invention with reference to the accompanying drawings and embodiments. This embodiment proposes a lightweight neural network collaborative optimization and deployment method based on memristors, Figure 1 This is the overall flowchart of the lightweight neural network collaborative optimization and deployment method for memristor hardware in this embodiment. In the lightweight neural network collaborative optimization and deployment method based on memristors proposed in this embodiment, the following steps are included:
[0113] 1) First, adopt the dynamic gradient-sensitive pruning algorithm, which aims to generate a sparse weight structure friendly to memristor hardware through a dual-driven scoring mechanism (L1 norm + gradient sensitivity) and hardware mapping efficiency feedback, solving the problems of low hardware utilization and large model accuracy loss in traditional pruning methods. The specific steps are as follows: First, perform input preprocessing, load the pre-trained floating-point model, and extract the weight matrices of each convolutional layer where C in is the number of input channels, C out is the number of output channels, and K is the convolutional kernel size.
[0114] 1-1) Input a batch of training data, perform forward propagation, and calculate the cross-entropy loss L:
[0115]
[0116] where N is the batch size, C is the number of classes, y i,c is the true label of the c-th class of the i-th sample, and p i,c is the model prediction probability.
[0117] 1-2) For each element l in the weight matrix W calculate its absolute value The larger the absolute value of the weight, the more stable its contribution to the output activation value, and the less destructive it is to the model after pruning. Then calculate the L1 norm score of the weight:
[0118]
[0119] where ||W l ||1 is the L1 norm of the weight matrix (the sum of the absolute values of all elements), and S L1 ∈[0, 1]. The larger the value, the higher the static importance of the weight.
[0120] 1-3) Perform backpropagation, calculate the gradient of the loss L with respect to the weight . The larger the absolute value of the gradient, the higher the sensitivity of the weight to the task loss, and the greater the impact on the model accuracy after removal. Then perform normalization processing and calculate the gradient sensitivity score:
[0121]
[0122] where the denominator is the maximum value of the gradient of the current layer, ensuring that S grad ∈[0, 1], represents the weight at the p-th spatial position (unrolled row by row) of the convolutional kernel from the m-th input channel to the n-th output channel in the l-th layer, and max m,n,p () means traversing all the weights of the l-th layer to find the maximum value of the absolute value of the gradient.
[0123] 1 - 4) Comprehensive scoring and ranking, weighted fusion of L1 - norm scoring and gradient sensitivity scoring, ranking from high to low according to the comprehensive score, and retaining the top (1 - p)×100% of the weights.
[0124] 1 - 5) Adopt progressive pruning scheduling, set the initial pruning rate p0 = 20%, remove the lowest 20% of the weights in the first round, and after each round of pruning, the pruning rate increases by Δp = 10% until the total pruning rate p total = 80%, avoiding model collapse caused by one - time pruning, allowing the network to gradually adapt to the sparse structure. After multiple rounds of pruning, the number of neural network parameters can be significantly reduced.
[0125] 1 - 6) Calculate the actual utilization rate η of the weights after pruning in the memristor array, and adopt a structured pruning strategy, including convolutional - kernel - level pruning: calculate the mean comprehensive score of each convolutional kernel and remove the kernel with the lowest score; channel - level pruning: remove the input or output channels to reduce the feature map dimension.
[0126]
[0127] where the denominator is the total number of array units (for example, a 256×512 array has a total of 131,072 units). If η < 85%, it means that the sparse pattern is severely fragmented and structured pruning needs to be triggered.
[0128] 1 - 7) Generate a binary mask M according to the scoring threshold l , and the threshold is adaptively adjusted according to the current pruning rate p to ensure uniform compression of each layer.
[0129]
[0130] 1 - 8) Adopt a fine - tuning strategy where the learning rate decays to 10% of the original value, only update the unpruned weights to avoid zero - weight interference with gradient propagation. After each round of fine - tuning, evaluate the model accuracy on the validation set. If the accuracy drops by more than 2%, roll back to the previous pruning state and terminate pruning. The weights after fine - tuning are as follows:
[0131]
[0132] where ∈ is the learning rate (such as the initial learning rate of 0.1 and the decayed learning rate of 0.01), and ⊙ is element - wise multiplication.
[0133] 2) Multi-Level Quantization-Aware Training (MQAT) aims to encode floating-point weights into memristor conductance states, and minimize the hardware quantization error through non-linear conductance modeling and mixed-precision quantization strategies. The specific steps are as follows: Define a differentiable function to simulate the conductance-weight relationship G(w) = a·tanh(b·w) + c, where a is the scaling factor controlling the conductance range, b is the slope factor adjusting the non-linear saturation rate, and c is the conductance reference offset. Determine the initial values by fitting the measured IV curve using non-linear least squares method.
[0134] 2-1) Synchronously optimize the network weights and conductance mapping parameters, reducing the quantization error to less than 3.5%.
[0135] 2-2) The convolutional layer is responsible for feature extraction and needs to retain a relatively high precision (4 bits corresponding to 16 levels of conductance). The fully connected layer has a high parameter redundancy and can be compressed to 2 bits (4 levels of conductance). The differentiable quantization forward propagation is achieved through the following quantization function.
[0136]
[0137] w: Neural network weight parameter;
[0138] q min / q max : The minimum / maximum value of the quantization dynamic range, used to limit the truncation boundary of the weight value;
[0139] Δ: Quantization step size;
[0140] clip(w, q min , q max ): Truncation function, which limits the weight w within the interval (q min , q max );
[0141] round(·): Rounding operation, which maps continuous values to the nearest quantization level;
[0142] where clip(w, q min , q max ) truncates the weight to the quantization range, round rounds to the nearest integer, and Δ is selected according to whether it is a convolutional layer or a fully connected layer. If the weight exceeds the dynamic range, it is forced to be truncated to q min or q max .
[0143] 2-3) Gradient approximation rule. Within the quantization interval, assuming the quantization error is negligible, the gradient is directly backpropagated; outside the interval, the gradient is set to zero to avoid noise interference. Jointly optimize the forward propagation and backward propagation processes. The weight first passes through the conductance mapping function G(w), and then through the quantizer Q(w) to generate discrete values to calculate the task loss L task and the mapping loss Lmap ; Backpropagate the gradients through STE to update the network weights W and the conductance parameters {a, b, c}.
[0144]
[0145] 2 - 4) Every 10 training epochs, map the quantized weights to the simulated memristor array, calculate the conductance mapping error. If the error ∈ map > 5%, it is determined that there is hardware incompatibility, re - initialize {a, b, c} and roll back the model to the previous checkpoint.
[0146]
[0147] 3) Hardware - aware knowledge distillation aims to improve the anti - interference ability of the student model (compressed model) by injecting memristor non - ideal characteristic noise, and solve the problem of accuracy degradation caused by quantization and hardware errors. The specific steps are as follows: Model the conductance drift. The conductance value decays with a square - root dependence over time. After long - term operation, the conductance value offset causes the accumulation of calculation errors and an increase in the confidence fluctuation of the model output.
[0148]
[0149] Where G0 is the initial conductance value (mapped from the quantization result Q(w)), σ d is the drift coefficient based on the 1T1R array, and t is the time step, simulating the continuous working time during the inference process.
[0150] 3 - 1) Model the read - write noise. Gaussian noise injection exposes the student model to a perturbation environment close to real hardware, enhancing the generalization ability.
[0151] G read = G + σ r ·N(0, 1)
[0152] Where σ r is the noise intensity, and N(0, 1) is the standard Gaussian distribution.
[0153] 3 - 2) In each training batch, randomly sample the time t ~ Uniform(0, T max ) and the noise intensity σ r ~ N(σ r , 0.02) to generate the anti - interference conductance value, and perform forward propagation using the perturbed conductance value to obtain the output y' of the student model student .
[0154] G perturbed = G'(t)+σ' r ·N(0, 1)
[0155] Where σ'r is the dynamic noise intensity, with an initial mean of σ r = 0.1 and a variance of 0.02.
[0156] 3 - 3) Adaptive noise enhancement. Every 10 training cycles, increase the mean noise intensity σ r by 10%. Monitor the accuracy on the validation set. If it drops by more than 2%, roll back to the previous stage parameters and terminate the enhancement.
[0157] 3 - 4) Anti - interference loss design. Ensure the basic classification ability of the student model through the task loss (cross - entropy) L task and force the student model to approximate the teacher model's output distribution under noise through the KL - divergence loss L KL . Finally, obtain a total loss function, specifically:
[0158]
[0159] where N is the batch size, C is the number of classes, y i,c is the true label of the c - th class of the i - th sample, is the predicted probability of the c - th class of the i - th sample by the student model.
[0160]
[0161] where is the output probability of the teacher model, is the output probability of the student model under noise perturbation.
[0162]
[0163] where λ1, λ2, λ3 are determined by experiments, is the L2 regularization term to prevent overfitting, and W represents the weight matrix.
[0164] 3.5) Adopt an optimization strategy of freezing the teacher model and updating the student model. The parameters of the teacher model are fixed, and only provide supervision signals through forward propagation. The gradient is only backpropagated to the student model to update its weights and quantization parameters.
[0165] The present invention provides a memristor - based multi - level compression collaborative optimization neural network deployment device, including the following modules:
[0166] Dynamic gradient - sensitive pruning module: Based on a dual - drive scoring mechanism (L1 - norm scoring and gradient sensitivity scoring) and hardware mapping efficiency feedback, perform progressive structured pruning on the pre - trained neural network model to generate a hardware - friendly sparse weight structure;
[0167] Multi-level quantization-aware training module: Simulate the conductance characteristics of memristors through a non-linear conductance modeling function, perform mixed-precision quantization encoding on the pruned model, and synchronously optimize the network weights and conductance mapping parameters to reduce quantization errors;
[0168] Hardware-aware knowledge distillation module: Inject memristor non-ideal characteristic noise into the vectorized model, and improve the robustness of the model in a noisy environment through anti-interference distillation training to obtain an optimized lightweight neural network model;
[0169] Map the optimized lightweight neural network model to a memristor in-memory computing architecture, perform weight encoding, array mapping, and in-memory computing inference, and output a deployable lightweight neural network model.
[0170] In the above solution, the specific steps of the dynamic gradient-sensitive pruning module include:
[0171] 1.1) Load a pre-trained floating-point neural network model and extract the weight matrices of each convolutional layer;
[0172] 1.2) Input training data to perform forward propagation and calculate the cross-entropy loss;
[0173] 1.3) For each element in the weight matrix, calculate its L1 norm score and gradient sensitivity score. The L1 norm score is the normalized value of the absolute value of the weight relative to the L1 norm of the weight matrix of the current layer, and the gradient sensitivity score is the normalized value of the absolute value of the weight gradient relative to the maximum gradient of the current layer;
[0174] 1.4) Generate a comprehensive score by weighted fusion of the L1 norm score and the gradient sensitivity score, and sort the weights according to the comprehensive score;
[0175] 1.5) Adopt a progressive pruning scheduling strategy to gradually increase the pruning rate until the target compression rate is reached. After each round of pruning, fine-tune to restore the model accuracy;
[0176] 1.6) Based on the feedback of the hardware mapping efficiency of the memristor array, adaptively adjust the pruned sparse weight structure, and preferentially perform convolutional kernel-level or channel-level structured pruning to ensure that the hardware utilization rate is higher than the preset threshold;
[0177] 1.7) Generate a binary mask according to the dynamic scoring threshold. A mask value of 1 indicates that the weight is retained, 0 indicates pruning, and the sparsification of the model is completed through the joint action of the mask and gradient update.
[0178] In the above solution, the multi-level quantization-aware training module includes the following steps:
[0179] 2.1) Define the differentiable non-linear conductance modeling function \(G(w)=a\cdot\tanh(b\cdot w)+c\), where \(a\) is the conductance range scaling factor, \(b\) is the slope factor for adjusting the non-linear saturation rate, and \(c\) is the conductance reference offset. Determine the initial parameter values by fitting the measured IV curve of the memristor using non-linear least squares method;
[0180] 2.2) Adopt a mixed-precision quantization strategy. Perform 4-bit quantization on the convolutional layer weights to retain 16 conductance states, and perform 2-bit quantization on the fully-connected layer weights to compress to 4 conductance states. Implement forward propagation quantization through the differentiable quantization function \(Q(w)\):
[0181]
[0182] w: Neural network weight parameter;
[0183] q min / q max : The minimum / maximum value of the quantization dynamic range, used to limit the truncation boundary of the weight value;
[0184] Δ: Quantization step size;
[0185] clip(w, q min , q max ): Truncation function, which limits the weight \(w\) within the interval \((q min , q max );
[0186] round(·): Rounding operation, which maps continuous values to the nearest quantization level;
[0187] 2.3) Apply the gradient approximation rule. Keep the gradient directly backpropagated within the quantization interval, and set the gradient to zero outside the interval to eliminate noise interference. Synchronously optimize the network weights and the conductance mapping parameters \(\{a, b, c\}\);
[0188] 2.4) Map the quantized weights to the simulated memristor array every 10 training epochs, and calculate the conductance mapping error \(\varepsilon map . If \(\varepsilon map > 5\%\), re-initialize \(\{a, b, c\}\) and roll back the model to the previous checkpoint.
[0189] In the above scheme, the hardware-aware knowledge distillation module includes the following steps:
[0190] 3.1) Model the non-ideal characteristics of the memristor, including the conductance drift model and the read / write noise model, where:
[0191] The conductance drift model is:
[0192]
[0193] where G0 is the initial conductance value of the quantization weight mapping, and σ d is the drift coefficient based on the 1T1R array, t is the time step, and N(0, 1) is the standard Gaussian distribution;
[0194] The read / write noise model is:
[0195] G read = G + σ r ·N(0, 1)
[0196] where σ r is the noise intensity, and G represents the original conductance value of the memristor array.
[0197] 3.2) Dynamic noise injection and anti-interference training:
[0198] In each batch of training, randomly sample the time t ~ Uniform(0, T max ) and the noise intensity σ' r ~ N(σ r , 0.02), generate the composite perturbation conductance value G perturbed , and use the perturbed conductance value to perform forward propagation to obtain the output y' student of the student model:
[0199] G perturbed = G'(t) + σ' r ·N(0, 1)
[0200] where σ' r is the dynamic noise intensity;
[0201] 3.3) Adaptive noise enhancement control:
[0202] Every 10 training epochs, increase the mean value of the noise intensity by 10%, and monitor the accuracy on the validation set. If the accuracy drops by more than 2%, roll back to the previous stage parameters and terminate the enhancement;
[0203] 3.4) Design a dual-loss joint optimization strategy. Through the KL divergence loss L KL force the output distribution of the student model in the noise environment to approximate that of the teacher model, and at the same time retain the cross-entropy loss L task to maintain the basic classification ability. The total loss function is:
[0204]
[0205] where λ1, λ2, λ3 are adjustable weighting coefficients, is the L2 regularization term;
[0206] 3.5) Freeze the parameters of the teacher model, and only update the weights and quantization parameters of the student model through gradient backpropagation until the model converges in the noise environment.
[0207] The present invention also provides a storage medium. When a program in the storage medium is executed by a processor, the described method for deploying a multi-level compression collaborative optimization neural network based on memristors is implemented.
[0208] This embodiment proposes a lightweight neural network deployment method based on a memristor in-memory computing architecture and multi-level compression collaborative optimization, as Figure 2 shown, which is a schematic diagram of the memristor cross array of the neural network accelerator in this embodiment; the computing unit of the neural network accelerator is a memristor cross array composed of memristors, and the method is used to accelerate the calculation of the neural network based on memristors; this method innovatively constructs a dual-drive pruning strategy based on the L1 norm and gradient sensitivity, combines the differentiable non-linear quantization coding of the memristor conductance characteristics, and the anti-interference distillation training with non-ideal characteristic noise injection, significantly improving the hardware utilization rate and energy efficiency ratio while ensuring the model accuracy, providing a feasible idea for key problems such as insufficient storage density and non-ideal characteristic interference in memristor deployment, and providing an efficient neural network hardware deployment solution for edge computing scenarios.
Claims
1. A method for deploying a multi-level compression collaborative optimization neural network based on memristors, characterized in that, It includes the following steps: (1) Dynamic gradient-sensitive pruning: Based on a dual-drive scoring mechanism (L1 norm scoring and gradient sensitivity scoring) and hardware mapping efficiency feedback, perform progressive structured pruning on the pre-trained neural network model to generate a hardware-friendly sparse weight structure; (2) Multi-level quantization-aware training: Simulate the memristor conductance characteristics through a non-linear conductance modeling function, perform mixed-precision quantization encoding on the pruned model, and synchronously optimize the network weights and conductance mapping parameters to reduce quantization errors; (3) Hardware-aware knowledge distillation: Inject memristor non-ideal characteristic noise into the vectorized model, and improve the robustness of the model in a noisy environment through anti-interference distillation training to obtain an optimized lightweight neural network model; (4) Map the optimized lightweight neural network model to a memristor in-memory computing architecture, perform weight encoding, array mapping, and in-memory computing inference, and output a deployable lightweight neural network model.
2. The method according to claim 1, characterized in that The specific steps of the dynamic gradient-sensitive pruning include: 1.1) Load the pre-trained floating-point neural network model and extract the weight matrices of each convolutional layer; 1.2) Input the training data to perform forward propagation and calculate the cross-entropy loss; 1.3) For each element in the weight matrix, calculate its L1 norm score and gradient sensitivity score. The L1 norm score is the normalized value of the absolute weight relative to the L1 norm of the weight matrix of the current layer, and the gradient sensitivity score is the normalized value of the absolute weight gradient relative to the maximum gradient of the current layer; 1.4) Generate a comprehensive score by weighted fusion of the L1 norm score and the gradient sensitivity score, and sort the weights according to the comprehensive score; 1.5) Adopt a progressive pruning scheduling strategy to gradually increase the pruning rate until the target compression rate is reached. After each round of pruning, fine-tune to restore the model accuracy; 1.6) Based on the hardware mapping efficiency feedback of the memristor array, adaptively adjust the pruned sparse weight structure, and preferentially perform convolutional kernel-level or channel-level structured pruning to ensure that the hardware utilization rate is higher than the preset threshold; 1.7) Generate a binary mask according to the dynamic scoring threshold. A mask value of 1 indicates that the weight is retained, and 0 indicates pruning, and complete model sparsification through the joint action of the mask and gradient update.
3. The method according to claim 1, wherein The multi-level quantization-aware training includes the following steps: 2.1) Define a differentiable non-linear conductance modeling function \(G(w)=a\cdot\tanh(b\cdot w)+c\), where \(a\) is the conductance range scaling factor, \(b\) is the slope factor for adjusting the non-linear saturation rate, and \(c\) is the conductance reference offset. Determine the initial parameter values by non-linear least squares fitting of the measured IV curve of the memristor; 2.2) Adopt a mixed-precision quantization strategy, perform 4-bit quantization on the convolutional layer weights to retain 16 conductance states, perform 2-bit quantization on the fully connected layer weights to compress to 4 conductance states, and implement forward propagation quantization through a differentiable quantization function \(Q(w)\): w: Neural network weight parameter; q min / q max : The minimum / maximum value of the quantization dynamic range, which is used to limit the truncation boundary of the weight value; Δ: Quantization step size; clip(w, q min , q max ): Truncation function that restricts the weight w within the interval (q min , q max ); round(·): Rounding operation, mapping continuous values to the nearest quantization level; 2.3) Apply the gradient approximation rule, keep the direct backpropagation of the gradient within the quantization interval, and set the gradient to zero outside the interval to eliminate noise interference. Synchronously optimize the network weights and the conductance mapping parameters {a, b, c}; 2.4) Map the quantized weights to the simulated memristor array every 10 training cycles, and calculate the conductance mapping error ε map’ If ε map > 5%, re-initialize {a, b, c} and roll back the model to the previous checkpoint.
4. The method according to claim 1, wherein The hardware-aware knowledge distillation includes the following steps: 3.1) Model the non-ideal characteristics of memristors, including the conductance drift model and the read / write noise model, where: The conductance drift model is: where G0 is the initial conductance value of the quantization weight mapping, σ d is the drift coefficient based on the 1T1R array, t is the time step, and N(0, 1) is the standard Gaussian distribution; The read / write noise model is: G read = G + σ r ·N(0, 1) where σ r is the noise intensity, and G represents the original conductance value of the memristor array. 3.2) Dynamic noise injection and anti-interference training: At each batch of training, randomly sample time \(t\sim Uniform(0, T)\) max ) and noise intensity \(\sigma'\) r \(\sim N(\sigma\) r , 0.02), generate the composite perturbation conductance value \(G\) perturbed , and use the perturbation conductance value to perform forward propagation to obtain the output \(y'\) of the student model student : G perturbed = G′(t) + σ′ r ·N(0, 1) where σ′ r is the dynamic noise intensity; 3.3) Adaptive noise enhancement control: Every 10 training cycles, increase the average noise intensity by 10%, and monitor the accuracy on the validation set. If the accuracy drops by more than 2%, roll back to the previous stage parameters and terminate the enhancement; 3.4) Design a dual-loss joint optimization strategy. Through the KL divergence loss L KL force the output distribution of the student model in a noisy environment to approximate that of the teacher model, while retaining the cross-entropy loss L task to maintain the basic classification ability. The total loss function is: where λ1, λ2, and λ3 are adjustable weighting coefficients, is the L2 regularization term; 3.5) Freeze the parameters of the teacher model, and only update the weights and quantization parameters of the student model through gradient backpropagation until the model converges in the noise environment.
5. A multi-level compression collaborative optimization neural network deployment device based on a memristor, characterized in that It includes the following modules: Dynamic gradient-sensitive pruning module: Based on the dual-drive scoring mechanism (L1 norm scoring and gradient sensitivity scoring) and the feedback of hardware mapping efficiency, perform progressive structured pruning on the pre-trained neural network model to generate a hardware-friendly sparse weight structure; Multi-level quantization-aware training module: Simulate the conductance characteristics of memristors through a non-linear conductance modeling function, perform mixed-precision quantization coding on the pruned model, and synchronously optimize the network weights and conductance mapping parameters to reduce the quantization error; Hardware-aware knowledge distillation module: Inject the non-ideal characteristics noise of memristors into the quantized model, and improve the robustness of the model in the noise environment through anti-interference distillation training to obtain an optimized lightweight neural network model; Map the optimized lightweight neural network model to the memristor in-memory computing architecture, perform weight encoding, array mapping, and in-memory computing inference, and output a deployable lightweight neural network model.
6. The device according to claim 5, characterized in that The specific steps of the dynamic gradient-sensitive pruning module include: 1.1) Load the pre-trained floating-point neural network model and extract the weight matrices of each convolutional layer; 1.2) Input the training data to perform forward propagation and calculate the cross-entropy loss; 1.3) For each element in the weight matrix, calculate its L1 norm score and gradient sensitivity score. The L1 norm score is the normalized value of the absolute value of the weight relative to the L1 norm of the weight matrix of the current layer, and the gradient sensitivity score is the normalized value of the absolute value of the weight gradient relative to the maximum gradient of the current layer; 1.4) Generate a comprehensive score by weighted fusion of the L1 norm score and the gradient sensitivity score, and sort the weights according to the comprehensive score; 1.5) Adopt a progressive pruning scheduling strategy to gradually increase the pruning rate until the target compression rate is reached. After each round of pruning, restore the model accuracy through fine-tuning; 1.6) Based on the feedback of the hardware mapping efficiency of the memristor array, adaptively adjust the pruned sparse weight structure, and preferentially perform convolutional kernel-level or channel-level structured pruning to ensure that the hardware utilization rate is higher than the preset threshold; 1.7) Generate a binary mask according to the dynamic scoring threshold. A mask value of 1 indicates that the weight is retained, and 0 indicates pruning. Complete the model sparsification through the joint action of the mask and gradient update.
7. The device according to claim 5, characterized in that, The multi-level quantization-aware training module includes the following steps: 2.1) Define the differentiable non-linear conductance modeling function \(G(w)=a\cdot\tanh(b\cdot w)+c\), where \(a\) is the conductance range scaling factor, \(b\) is the slope factor for adjusting the non-linear saturation rate, and \(c\) is the conductance reference offset. Determine the initial parameter values by fitting the measured IV curve of the memristor using the non-linear least squares method; 2.2) Adopt a mixed-precision quantization strategy. Perform 4-bit quantization on the weights of the convolutional layer to retain 16 conductance states, and perform 2-bit quantization on the weights of the fully connected layer to compress to 4 conductance states. Implement forward propagation quantization through the differentiable quantization function \(Q(w)\): \(w\): Neural network weight parameter; q min / q max : The minimum / maximum value of the quantization dynamic range, used to limit the truncation boundary of the weight value; \(\Delta\): Quantization step size; clip(w, q min q max ): Truncation function that restricts the weight w within the interval (q min , q max ); round(·): Rounding operation that maps continuous values to the nearest quantization level; 2.3) Apply the gradient approximation rule. Keep the gradient directly backpropagated within the quantization interval and set the gradient to zero outside the interval to eliminate noise interference. Synchronously optimize the network weights and the conductance mapping parameters \(\{a, b, c\}\); 2.4) Map the quantized weights to the simulated memristor array every 10 training cycles, and calculate the conductance mapping error ε map , if ε map > 5%, re-initialize {a, b, c} and roll back the model to the previous checkpoint.
8. The device according to claim 5, characterized in that, The hardware-aware knowledge distillation module includes the following steps: 3.1) Model the non-ideal characteristics of the memristor, including the conductance drift model and the read / write noise model, where: The conductance drift model is: where G0 is the initial conductance value of the quantization weight mapping, σ d is the drift coefficient based on the 1TlR array, t is the time step, and N(0, 1) is the standard Gaussian distribution; The read / write noise model is: G read = G + σ r ·N(0, 1) where σ r is the noise intensity, and G represents the original conductance value of the memristor array. 3.2) Dynamic noise injection and anti-interference training: In each batch of training, the random sampling time t~Uniform(0, T max ) and noise intensity σ′ r ~N(σ r , 0.02), generating the composite perturbation conductance value G perturbed , perform forward propagation using the perturbed conductance value to obtain the student model output y′ student : G perturbed = G′(t) + σ′ r ·N(0, 1) where σ′ r is the dynamic noise intensity; 3.3) Adaptive noise enhancement control: Every 10 training cycles, increase the mean noise intensity by 10%, and monitor the accuracy on the validation set. If the accuracy drops by more than 2%, roll back to the previous stage parameters and terminate the enhancement; 3.4) Design a dual-loss joint optimization strategy. Through the KL divergence loss L KL force the output distribution of the student model in a noisy environment to approximate that of the teacher model, while retaining the cross-entropy loss L task to maintain the basic classification ability. The total loss function is as follows: where λ1, λ2, and λ3 are adjustable weighting coefficients, is the L2 regularization term; 3.5) Freeze the parameters of the teacher model, and only update the weights of the student model and the quantization parameters through gradient backpropagation until the model converges in the noise environment.
9. A storage medium, characterized in that, When the processor executes the program in the storage medium, the method described in any one of claims 1-4 is implemented.
Citation Information
Cited By
Model deployment method based on pruning compression in edge device
CN120633749A
A model deployment method based on pruning and compression in edge devices
CN120633749B
MEMS sensor high-bandwidth disturbance control method based on truncation distribution
CN120705824A
Manufacturing production line defect detection method and system based on big data intelligent algorithm
CN120991953A
Large model-oriented split privacy protection training method
CN121561979A