Heterogeneous Neural Network Acceleration Device Based on the Collaboration of DSP and Memristor

Through the heterogeneous neural network acceleration device that cooperates with DSP and memristor, the problem of insufficient energy efficiency ratio and drift resistance of HNN network acceleration circuit is solved, and efficient acceleration performance improvement and signal-to-noise ratio improvement are achieved.

CN120046677BActive Publication Date: 2025-07-22HUNAN GREAT WALL GALAXY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510509568.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-22
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

Traditional HNN network acceleration circuits have problems with insufficient energy efficiency ratio and drift resistance, resulting in low acceleration performance.

Method used

Using a heterogeneous neural network acceleration device based on the collaboration of DSP and memristors, through the hybrid computing architecture design, combined with dynamic accuracy switching, real-time closed-loop feedback mechanism of Liyapunov stability theory, adaptive cross-domain quantization module and hardware instruction set optimization, the in-situ update of memristor weights and efficient nonlinear calculation of DSP are realized.

Benefits of technology

The energy efficiency ratio and drift resistance of the HNN network acceleration circuit have been improved, the acceleration performance has been improved, the signal-to-noise ratio has been improved by more than 20%, the calculation energy efficiency ratio has been improved by 5-8 times, and the training convergence speed has been increased by more than 30%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046677B_ABST
    Figure CN120046677B_ABST
Patent Text Reader

Abstract

The present invention relates to a heterogeneous neural network acceleration device based on the cooperation of DSP and memristor. Through the design of a hybrid computing architecture of DSP and memristor cross-array, such as dynamic precision switching, a real-time closed-loop feedback mechanism based on Lyapunov stability theory, an adaptive cross-domain quantization module, and hardware instruction set optimization, etc., the hybrid architecture has good balance between performance and versatility. Moreover, the stability is significantly improved through conductance monitoring, adaptive PWM pulse generation, and quantization table update. The adaptive cross-domain quantization module is combined with noise reduction to achieve a higher signal-to-noise ratio at a lower cost. The hardware level is optimized through a customized instruction set, which is more suitable for HNN computing tasks, improves the energy efficiency ratio and anti-drift ability of the HNN network acceleration circuit, and enhances the acceleration performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of neural network hardware acceleration and hybrid computing, and relates to a heterogeneous neural network acceleration device based on the cooperation of DSP and memristor. Background Art

[0002] As the fourth basic circuit element, the memristor has core characteristics (the resistance value changes with the flowing charge and remains unchanged when powered off), making it an ideal device for simulating biological synapses. It is widely used in fields such as artificial neural networks (ANNs), secure communication, and in-memory computing (IMC). The in-memory computing technology solves the "memory wall" problem of the von Neumann architecture by directly performing calculations in the memory cells, and the energy consumption can be reduced to 12% of the traditional architecture.

[0003] As a recurrent neural network, the Hopfield neural network (HNN) is good at dealing with combinatorial optimization problems (such as the traveling salesman problem) and associative memory tasks. Its core lies in the dynamic convergence characteristics of the energy function. However, traditional CMOS (complementary metal oxide semiconductor) devices face technical bottlenecks of high power consumption and low scalability when used for network implementation. Memristors are widely used to construct the synaptic weights of HNNs due to their non-volatility and analog computing capabilities. In this field, there are already memristor-HNN hybrid architectures based on FPGAs, 1T1R memristor crossbar array acceleration schemes, analog in-memory computing (IMC) schemes, and dynamic calibration and hybrid quantization technologies. However, the above traditional technologies have insufficient energy efficiency ratio and anti-drift ability in practical applications, resulting in the technical problem of low acceleration performance in the HNN network acceleration circuit. Summary of the Invention

[0004] Aiming at the problems existing in the above traditional technologies, the present invention proposes a heterogeneous neural network acceleration device based on the cooperation of DSP and memristor, which can improve the energy efficiency ratio and anti-drift ability of the HNN network acceleration circuit and enhance the acceleration performance.

[0005] To achieve the above object, the embodiments of the present invention adopt the following technical solutions:

[0006] Provide a heterogeneous neural network acceleration device based on the cooperation of DSP and memristor, including a DSP coprocessor and a memristor crossbar array. The memristor crossbar array interacts with the DSP coprocessor through a PWM interface and maps the conductance value of the memristor to the memory space of the DSP coprocessor;

[0007] The memristive crossbar array integrates a 12-bit high-precision signal conversion module, a conductance monitoring circuit, and a PWM driving circuit. The high-precision signal conversion module is used to achieve seamless analog-to-digital signal conversion. The conductance monitoring circuit is used to monitor the change in the memristive conductance value and transmit it to the DSP coprocessor. The PWM driving circuit is used to receive the PWM pulses generated by the DSP coprocessor after calculating the error gradient based on the change in the memristive conductance value, adjust the conductance of the memristive crossbar array, and calibrate and update the conductance-digital activation mapping table;

[0008] The DSP coprocessor is used to adopt high-precision quantization in the signal-sensitive area and low-precision quantization in the signal-insensitive area through a non-uniform piecewise quantization strategy. The DSP coprocessor is deployed with an extended and customized SIMD instruction set and is used for hardware-level acceleration of heterogeneous neural networks; The SIMD instruction set includes an extended instruction set and a newly added instruction set. The extended instruction set is used to support SIMD parallel computing to update the neuron states of heterogeneous neural networks, and the newly added instruction set is used to dynamically calibrate the memristive weights of heterogeneous neural networks based on the Lyapunov stability theory.

[0009] One of the above technical solutions has the following advantages and beneficial effects:

[0010] The above heterogeneous neural network acceleration device based on the cooperation of DSP and memristor, through the design of a hybrid computing architecture of DSP and memristive crossbar array, such as dynamic precision switching, real-time closed-loop feedback mechanism based on the Lyapunov stability theory, adaptive cross-domain quantization module, and hardware instruction set optimization, etc., makes the hybrid architecture have good balance between performance and generality, and significantly improves stability through conductance monitoring, adaptive PWM pulse generation, and quantization table update, realizes higher signal-to-noise ratio at lower cost through joint noise reduction by the adaptive cross-domain quantization module, and realizes hardware-level optimization through the customized instruction set, is more suitable for HNN computing tasks, improves the energy efficiency ratio and anti-drift ability of the HNN network acceleration circuit, and improves the acceleration performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0012] Figure 1 It is a schematic diagram of the architecture of a heterogeneous neural network acceleration device based on the cooperation of DSP and memristor in an embodiment;

[0013] Figure 2 It is a schematic diagram of the cross-domain quantization coding architecture in an embodiment;

[0014] Figure 3 Schematic diagram of the dynamic weight calibration process in one embodiment;

[0015] Figure 4 Schematic diagram of SIMD parallel computing in one embodiment. Detailed implementation manners

[0016] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the description of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0017] It should be noted that referring to "embodiment" herein means that a specific feature, structure or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. The phrase is shown at various positions in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art can understand that the embodiments described herein can be combined with other embodiments. The term "and / or" used in the description and claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0018] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings in the embodiments of the present invention.

[0019] In the memristor-HNN hybrid architecture based on FPGA, the FPGA is used as a digital controller, the memristor array is used to store the synaptic weights, and the analog-to-digital converter ADC / digital-to-analog converter DAC is used to realize the digital-analog interaction. In the image encryption task, the delay is 8.5 ms / frame and the energy efficiency ratio is 8.3 GOPS / W. However, limited by the FPGA logic unit, the maximum supported network is 256×256 and there is a lack of a dynamic calibration mechanism. The conductance value of the memristor is easily affected by temperature and voltage fluctuations, resulting in a weight drift error >5% and obvious quantization noise of the analog signal. In addition, the memristor-HNN hybrid architecture implemented by FPGA has a higher delay and a slower training convergence speed.

[0020] In the 1T1R memristor crossbar array acceleration scheme, each memristor cell is connected in series with a transistor (1T1R) structure to suppress sneak path problems and support high-density integration. It is used for convolutional neural network acceleration, with a 10-fold increase in computing power. However, the transistors increase the chip area and power consumption, making it difficult to achieve ultra-large-scale expansion. In the analog in-memory computing (IMC) scheme, analog multiply-accumulation operations are directly performed in the memristor array to avoid data movement. For example, a commercially available Flash-based chip has an energy efficiency ratio as high as 25 TOPS / W, but it relies on high-precision ADC / DAC, is costly, and is vulnerable to noise interference.

[0021] Dynamic calibration and hybrid quantization techniques use a memristor weight calibration algorithm based on Lyapunov theory to suppress drift through feedback regulation, reducing the weight drift error to 2%. However, its calibration period is long (greater than 10 ms), with poor real-time performance. Some systems rely on pre-training or complex compensation steps and do not combine non-uniform quantization, resulting in insufficient accuracy in the low-current region. Among them, the resistance change of traditional memristors (TiO2) depends on ion migration, is vulnerable to environmental interference, and the above traditional techniques mostly use uniform quantization without optimizing for the non-linear characteristics of memristors, leading to insufficient accuracy in the low-current region.

[0022] When using a memory-computation separation architecture in the above traditional techniques, it is necessary to frequently read and write external storage, resulting in a low energy efficiency ratio. The serial computing mode lacks parallel acceleration units (such as SIMD (Single Instruction, Multiple Data)), and cannot efficiently process large-scale matrix operations. In addition, there is a lack of a closed-loop feedback mechanism, and it does not combine the Lyapunov stability theory to achieve real-time conductance adjustment. Moreover, there is insufficient hardware support, lacking dedicated instructions to accelerate gradient calculation and pulse generation.

[0023] Therefore, the present invention will provide a heterogeneous neural network acceleration device based on the cooperation of DSP and memristors, which solves the memory-computation separation problem through a "digital-analog hybrid computing architecture" and "dynamic adaptive instruction optimization", realizes in-situ update of memristor weights and efficient non-linear computing of DSP, improves the energy efficiency ratio and anti-drift ability of the HNN network acceleration circuit, and thus improves the acceleration performance.

[0024] In one embodiment, as Figure 1As shown in the figure, a heterogeneous neural network acceleration device based on the cooperation of DSP and memristor is provided, including a DSP coprocessor (DSP) and a memristive crossbar array. The memristive crossbar array interacts with the DSP coprocessor through a PWM interface and maps the conductance value of the memristor to the memory space of the DSP coprocessor. The memristive crossbar array integrates a 12-bit high-precision signal conversion module (ADC / DAC), a conductance monitoring circuit, and a PWM driving circuit. The high-precision signal conversion module is used to achieve seamless analog-digital signal conversion. The conductance monitoring circuit is used to monitor the change of the memristor conductance value and transmit it to the DSP coprocessor. The PWM driving circuit is used to receive the PWM pulse generated by the DSP coprocessor after calculating the error gradient according to the change of the memristor conductance value, adjust the conductance of the memristive crossbar array, and calibrate and update the conductance-digital activation mapping table. The DSP coprocessor is used to adopt high-precision quantization in the signal-sensitive area and low-precision quantization in the signal-insensitive area through a non-uniform segmented quantization strategy. The DSP coprocessor is deployed with an extended and customized SIMD instruction set (represented by a SIMD unit) and is used for hardware-level acceleration of heterogeneous neural networks; the SIMD instruction set includes an extended instruction set and a new instruction set. The extended instruction set is used to support SIMD parallel computing to update the neuron states of heterogeneous neural networks, and the new instruction set is used to dynamically calibrate the memristive weights of heterogeneous neural networks (i.e., the calibration algorithm) based on the Lyapunov stability theory.

[0025] It can be understood that in the first part of the design, such as Figure 1 In the architecture of the heterogeneous neural network acceleration device based on the cooperation of DSP and memristor shown in the figure, the memristive crossbar array serves as a synaptic weight storage and analog multiply-add unit, represents the weight magnitude through the conductance value (the reciprocal of the resistance), and supports in-situ update (without external storage) to reduce power consumption; it integrates a 12-bit high-precision signal conversion module ADC / DAC, which is used to achieve seamless analog-digital signal conversion, enabling the subsequent circuit to directly process analog signals. The memristive crossbar array directly interacts with the DSP through a PWM interface and supports mapping the dynamic conductance value to the memory space of the DSP.

[0026] The DSP coprocessor integrates an "adaptive cross-domain quantization module" and adopts a non-uniform segmented quantization strategy. It uses high-precision quantization for the signal-sensitive area (such as small current) to reduce noise sensitivity, while using low-precision quantization for the signal-insensitive area to effectively save resources. It has extended and customized an instruction set (such as vector Lyapunov function calculation instructions, memristor conductance gradient acceleration instructions) to achieve hardware-level acceleration of the HNN dynamic equation. It deploys a "dynamic weight calibration engine", which is a real-time feedback system based on the Lyapunov stability theory and is used to suppress the conductance drift of the memristor.

[0027] The implementation principle of the "dynamic weight calibration engine" is as follows: First, the conductance monitoring circuit monitors the change of the memristor conductance value. Then, the DSP calculates the error gradient and generates PWM pulses to adjust the conductance (for example, long pulses are used to increase the conductance, and short pulses are used to decrease the conductance). Finally, the "conductance-digital activation mapping table" is updated to maintain network stability. The conductance monitoring circuit can adopt the existing monitoring circuits in the field of memristors, which can automatically measure the output conductance through the current-conductance conversion relationship of the memristor.

[0028] The above heterogeneous neural network acceleration device based on the cooperation of DSP and memristor adopts a hybrid computing architecture design of DSP and memristive crossbar array, such as dynamic precision switching, real-time closed-loop feedback mechanism based on Lyapunov stability theory, adaptive cross-domain quantization module, and hardware instruction set optimization, etc., making the hybrid architecture have good balance between performance and versatility. And through conductance monitoring, adaptive PWM pulse generation, and quantization table update, the stability is significantly improved. Through the joint noise reduction of the adaptive cross-domain quantization module, a higher signal-to-noise ratio is achieved at a lower cost. Through the customized instruction set, the hardware level is optimized, which is more suitable for HNN computing tasks, improves the energy efficiency ratio and anti-drift ability of the HNN network acceleration circuit, and improves the acceleration performance.

[0029] Furthermore, regarding the cross-domain quantization coding algorithm of DSP and memristor, its corresponding cross-domain quantization coding architecture is as Figure 2 shown, including non-uniform segmented quantization (adaptive quantization), adaptive reference correction (dynamic weight calibration), noise suppression coding, and algorithm implementation process, etc. Specifically, it can be divided into the digital domain (DSP) and the analog domain (memristor). The improvements in the digital domain (DSP) include parallel computing, adaptive quantization, dynamic weight calibration, and instruction set extension. The improvements in the analog domain (memristor) include PWM drive, weight mapping, and analog multiply-accumulate. The process of dynamic weight calibration is as Figure 3 shown.

[0030] The quantization table structure of non-uniform segmented quantization is shown in Table 1:

[0031] Table 1

[0032]

[0033] The design basis of non-uniform segmented quantization is: according to the volt-ampere characteristic curve of the memristor, more quantization bits are allocated in the low current region (i.e., the sensitive region) to suppress noise; sparse quantization is performed in the high current region (i.e., the non-sensitive region) to save resources. The formula expression of non-uniform segmented quantization is as follows:

[0034] ;

[0035] Among them, represents the quantization result, roundThe () function is used to round a numerical value to a specified number of digits. V min represents the minimum voltage amplitude of the memristor. V max represents the maximum voltage amplitude of the memristor. Through piecewise linear mapping, the analog signal x is mapped to a 12-bit digital code (0 - 4095).

[0036] The reference sampling for adaptive reference correction is that the DSP periodically reads the output current when the memristor has zero input I offset (drift noise). The offset compensation for adaptive reference correction is to update the V min and V max in the formula for non-uniform piecewise quantization in real time through the following formula:

[0037] V min = V min + α I offset;

[0038] V max = V max + β I offset;

[0039] where α and β are calibration coefficients that can be calibrated through experiments.

[0040] The quantization table update for adaptive reference correction is to recalculate the current segmentation interval according to the new corrected reference and write it into the lookup table (LUT) of the DSP.

[0041] Noise suppression coding includes differential coding and Kalman filtering. Among them, differential coding is to continuously perform a difference operation on adjacent sampled values in time Q t and Q t-1 :

[0042] Δ Q = Q t - Q t-1;

[0043] Only the sign bit and magnitude bit of the differential signal Δ Q are transmitted (for example, 4 bits represent the change amount) to reduce the transmission bandwidth. Kalman filtering is to filter the differential signal at the DSP side to suppress high-frequency noise:

[0044] ;

[0045] Among them, K is the dynamically adjusted Kalman gain.

[0046] In some embodiments, the conductance of the memristive crossbar array is adjusted between the DSP coprocessor and the memristive crossbar array through a cross-domain quantization coding algorithm. Among them, the process steps of the cross-domain quantization coding algorithm may include:

[0047] Calibrate the initial conductance-current curve of the memristors in the memristive crossbar array to generate a default quantization table;

[0048] Configure the sampling rate of the high-precision signal conversion module and the interrupt response of the DSP coprocessor;

[0049] Perform ADC sampling through the high-precision signal conversion module, look up the table to determine the current segment interval to which the current signal obtained by the ADC sampling belongs, and then apply the corresponding quantization formula for quantization;

[0050] Perform differential coding on the quantized sampling values and then perform Kalman filtering, and write the obtained digital code into the memory mapping area of the DSP coprocessor; among them, adaptive reference calibration is performed every 1 ms;

[0051] Convert the PWM pulse generated by the DSP coprocessor into an output voltage through non-uniform DAC conversion by the high-precision signal conversion module to adjust the conductance of the memristive crossbar array.

[0052] Specifically, the algorithm implementation process of the cross-domain quantization coding algorithm is as follows:

[0053] Initialization stage; including (1) calibrating the initial conductance-current curve of the memristors to generate a default quantization table (for predefined current segment intervals), and (2) configuring the sampling rate (1 MSPS) of the high-precision signal conversion module ADC / DAC and the DSP interrupt response.

[0054] Enter the real-time quantization stage; including While (system running) and reverse control (DAC output). Among them, in the While (system running) stage, ADC sampling is performed to obtain the current signal I(t), look up the default quantization table to determine the current segment interval to which the current signal I(t) belongs, then apply the corresponding quantization formula, perform differential coding and Kalman filtering, and write the digital code obtained after differential coding and Kalman filtering into the memory mapping area of the DSP. Adaptive reference calibration is performed every 1 ms (update V min and V max ).

[0055] The reverse control (DAC output) converts the gradient pulses (PWM pulses) generated by the DSP into an output voltage through a non-uniform DAC. V out :

[0056] ;

[0057] The second part of the design is a SIMD-based parallel microarchitecture for updating neuron states. The design includes a SIMD data path, data loading and storage, a control unit, and instruction set extensions. By customizing SIMD + mixed precision, high-performance computing is achieved.

[0058] In one embodiment, the SIMD data path of the DSP co-processor includes a 128-bit SIMD register and a dedicated memristor weight cache. Each register of the 128-bit SIMD register stores 4 32-bit floating-point numbers or 8 16-bit fixed-point numbers. The dedicated memristor weight cache is a 128B buffer for storing the quantized weights read from the memristor crossbar array. The SIMD data path is deployed with a parallel multiply-accumulate unit using Booth encoding and Wallace tree structure to support 4-way parallel analog multiply-accumulate operations. The activation function accelerator of the parallel multiply-accumulate unit uses piecewise linear approximation of tanh. A schematic diagram of SIMD parallel computing is shown as Figure 4 shown.

[0059] Specifically, the register bank of the SIMD data path includes a 128-bit SIMD register (V0 - V15) and a dedicated memristor weight cache. Each register of the 128-bit SIMD register stores 4 32-bit floating-point numbers (FP32) or 8 16-bit fixed-point numbers (FP16). The dedicated memristor weight cache is a 128B buffer for storing the quantized weights (12-bit precision) read from the memristor crossbar array.

[0060] The parallel multiply-accumulate unit (PMAC) supports 4-way parallel analog multiply-accumulate operations and completes the following multiply-accumulate output in a single cycle:

[0061] ;

[0062] where, represents the quantized weight of the memristor in the i th row and j th column, X j represents the j th input signal, B i represents the reference threshold for adjusting the output. Hardware optimization uses Booth (Booth) encoding and Wallace (Wallace) tree structure to reduce the multiplier delay.

[0063] The activation function accelerator uses piecewise linear approximation of tanh: the non-linear function is split into 4 segments of linear approximation, implemented through a look-up table (LUT), and supports 4-way parallel computing.

[0064] ;

[0065] Among them, y represents the output of the activation function, x represents the neuron input signal (since the activation function is part of the nervous system, the signal x and the neuron state value are the same expression and can be an analog signal). 4 activation values are calculated in a single cycle through SIMD instructions.

[0066] In one embodiment, the DSP coprocessor uses a double-buffer mechanism or a broadcast mode for neuron state caching.

[0067] Specifically, the data loading and storage part includes: first, the memristive weight is loaded. The quantized weight (12 bits) of the memristive crossbar array is batch-loaded into the dedicated memristive weight cache through DMA and automatically expanded to the FP16 / FP32 format; then data alignment is performed, and the quantized weights are aligned according to the 4-way parallel requirement (for example, the weight matrix is stored separately in 4×4); then the neuron state caching is completed, and a double-buffer mechanism (that is, two 128B buffer areas alternate for reading and writing to hide the data transfer delay) or a broadcast mode (that is, a single input signal is broadcast to 4 parallel computing units (applicable to the fully connected layer)) can be used.

[0068] The control unit part includes an instruction scheduler and dynamic precision switching. The instruction scheduler is used to parse custom SIMD instructions (such as VMM4 vector matrix multiplication), dynamically allocate computing resources, and support instruction-level parallelism (ILP), such as simultaneously executing multiply-accumulate and activation functions. Dynamic precision switching is to switch the FP16 / FP32 format through the status register (CR0) to enable the hardware to automatically adjust the data path bit width.

[0069] In one embodiment, the instruction set extension is for supporting SIMD parallel computing, and the extended instruction set is shown in Table 2 below:

[0070] Table 2

[0071]

[0072] Among them, Vd represents the target vector register, Vw represents the weight vector register, and Vx represents the neuron iThe status vector register, Vs represents the gradient value vector register, Rs represents the source base address register for calculating the base address of the memory address, Rd represents the destination base address register for storing the destination address base of the memory to be written, and offset represents the address offset, which is added to the base address to generate the effective address. Based on the foregoing design, an example of the HNN status update is as follows, parallelly calculating the status updates of 4 neurons:

[0073] Call the instruction VLDMEMV1, [R0]; Load the weight block W[0:3];

[0074] Call the instruction VLDMEMV2, [R1]; Load the input vector X[0:3];

[0075] Call the instruction VMM4V3, V1, V2; (4-way multiply-accumulate);

[0076] Call the instruction VTANHV4, V3; V4 = tanh(V3);

[0077] Call the instruction VSTNEURON[R2], V4; Store the neuron status Y[0:3].

[0078] In some embodiments, when performing neuron status update between the DSP coprocessor and the memristive crossbar array, the following processing may be included:

[0079] Adopt crossbar dynamic routing for weights and input data;

[0080] In the FP16 format, a 32-bit accumulator is used inside the multiply-accumulate unit for accumulation. When the accumulated result exceeds the FP16 range, an interrupt is automatically triggered to switch to the FP32 format;

[0081] The 5-stage pipeline adopted is instruction fetch, decode, memristor load, execute, and write-back. Among them, in the memristor load stage, the next set of weights is prefetched.

[0082] Specifically, the corresponding hardware optimizations include the weight-input alignment network, dynamic precision fusion, and memristor-DSP collaborative pipeline. The weight-input alignment network uses a crossbar to dynamically route weights and input data, supports non-contiguous memory access, and its application scenario can be when dealing with sparse connections, skipping zero-weight calculations, reducing power consumption by 30%.

[0083] Dynamic precision fusion includes mixed-precision calculation, that is, in the FP16 format, a 32-bit accumulator is used inside the multiply-accumulate unit to avoid precision loss; and overflow protection, that is, when the accumulated result exceeds the FP16 range, an interrupt is automatically triggered to switch to the FP32 format. The memristor-DSP collaborative pipeline adopts a 5-stage pipeline: 1. Instruction Fetch (IF) → 2. Decode (ID) → 3. Memristor Load (MEM) → 4. Execute (EX) → 5. Write Back (WB); the optimization of the MEM stage can hide the memristor read latency by prefetching the next set of weights.

[0084] In one embodiment, for the third part of the design, the new instruction set in the DSP instruction set extension for HNN is designed as shown in Table 3 below:

[0085] Table 3

[0086]

[0087] The detailed description of the key instructions is as follows:

[0088] VGRAD: Parallel gradient calculation instruction, whose function is to calculate the weight gradient based on the Lyapunov stability theory , the formula is:

[0089] ;

[0090] Where, is the learning rate, x i , x j is the neuron state.

[0091] The hardware implementation of the instruction can be as follows: The inputs are Vw (weight), V x (neuron i state) and V y (neuron j state); the calculation unit uses a 4-way parallel floating-point multiply-accumulate unit, which supports the fast approximation of the non-linear term ; the output is the gradient value stored in the Vd register, in the FP32 format. An example code is: VGRAD V0, V1, V2, V3; .

[0092] VPULSEGEN: Gradient pulse generation instruction, whose function is to convert the gradient value into the memristor PWM control parameters (pulse width, amplitude); the mapping rule is pulse width = , polarity = , where, K p is the proportionality coefficient (configured by the register Rt). The hardware implementation of the instruction can be as follows: The inputs are Vs (gradient value) and Rt ( Kp and pulse parameters); the output is to store 4 PWM control words (32 bits, including width and polarity) in Vd; the pipeline optimization directly connects the pulse parameter prediction calculation module to the DAC interface to reduce the control delay.

[0093] VLYAP: Lyapunov energy function instruction, whose function is to calculate the HNN energy function E , for stability judgment:

[0094] ;

[0095] Its hardware implementation can be as follows: the input is the neuron state V state (4 neuron states x0 - x3), the calculation process is to first load a 4×4 weight block from the dedicated memristor weight cache, and then calculate in parallel , then sum and take the negative, and store the obtained result in Vd. The example code is as: VLYAP V0, V1; .

[0096] VMEMSYNC: Memristor weight synchronization instruction, whose function is to synchronize the weight update in the DSP memory to the memristor crossbar array, supporting batch transmission. Its hardware implementation is as follows: first, align the data address by 4×4 weight blocks (128 bits), and then send the conductance value update command through the SPI interface to automatically trigger the PWM drive circuit of the memristor crossbar array.

[0097] In one embodiment, the instruction set hardware architecture adopted by the new instruction set includes a dedicated computing unit and a weight broadcast network. The dedicated computing unit includes a gradient calculation unit and a pulse generation unit. The gradient calculation unit is a 4-way parallel floating-point multiplier that supports the Taylor expansion approximation of non-linear terms, used to perform dynamic precision switching. The pulse generation unit is used to map the gradient value to the PWM pulse width and select the voltage direction according to the gradient sign bit. The weight broadcast network is used to broadcast a single weight value to 4 dedicated computing units simultaneously.

[0098] Specifically, the hardware architecture of the new instruction set in the DSP instruction set extension for HNN includes a dedicated computing unit, data path optimization, and pipeline design, and application examples are given. Among them, the dedicated computing unit includes a gradient calculation unit (GCU, which is a 4-way parallel floating-point multiplier that supports the Taylor expansion approximation (3rd order) of non-linear terms ) and a pulse generation unit (PCU). The pulse generation unit consists of a time-to-digital converter (TDC) and a polarity controller. The time-to-digital converter is used to map the gradient value to the PWM pulse width (resolution 1ns), and the polarity controller is used to select the voltage direction according to the gradient sign bit (such as +Vdd / -Vdd).

[0099] The data path optimization adopts a weight broadcast network and dynamic precision switching. Among them, the weight broadcast network allows a single weight value to be broadcast to 4 dedicated computing units simultaneously, which is applicable to fully connected HNNs; dynamic precision switching is used to reduce the memory width in the FP16 format by adopting a compressed weight format (12-bit mantissa) inside the gradient calculation unit GCU.

[0100] In some embodiments, the pipeline design corresponding to the new instruction set adopts a 6-stage pipeline: instruction fetch, decoding, weight loading, gradient calculation, pulse generation, and write-back. Its conflict resolution includes data hazards and control hazards. Data hazards are avoided through register renaming and out-of-order execution, while control hazards are optimized by using a branch predictor for loop unrolling (the number of HNN iterations is fixed).

[0101] The application example can be the following HNN weight update process, which updates 4 memristor weights in parallel:

[0102] Call the instruction VLDMEMV1, [R0]; load the current weight W [0:3];

[0103] Call the instruction VLDMEMV2, [R1]; load the neuron state X[0:3];

[0104] Call the instruction VGRADV3, V1, V2, V2; calculate the gradient ;

[0105] Call the instruction VPULSEGENV4, V3, R2; generate PWM pulse parameters (R2 = K p )

[0106] Call the instruction VMEMSYNC[R3], 4; synchronize the weight to the memristor crossbar array (R3 is the memristor address).

[0107] The specific implementation method can be as follows: In the hardware configuration, the memristor crossbar array adopts a 1T1R structure, with a scale of 256×256, TiO2 memristor devices, and an integrated PWM drive circuit; the DSP chip selects an existing DSP and expands a custom instruction set (such as the VLYAP instruction to accelerate the calculation of the Lyapunov function); the cross-domain interface adopts a 12-bit 1MSPS ADC / DAC module, supporting a dynamic range 5V.

[0108] The software implementation process includes: (1) In the initialization stage, the DSP reads the memristor conductance value through the SPI interface, constructs the weight matrix and loads it into the L2 cache. (2) In the inference stage, the input signal undergoes analog multiplication and addition through the memristor crossbar array, and the output current is converted into a digital signal by the ADC. The DSP calls the extended instruction set to execute the activation function (piecewise linear approximation of tanh) and state iteration, and the result is transmitted to the output cache through DMA, with the whole process processed in a pipelined manner. (3) In the training stage, the DSP generates gradient pulses based on the error backpropagation, adjusts the memristor conductance through the DAC, and the dynamic weight calibration engine monitors the conductance value in real time, updates the parameters of the conductance-digital activation mapping table and feeds back to the training algorithm.

[0109] The above heterogeneous neural network acceleration device based on the cooperation of DSP and memristor, compared with traditional technologies, through the integration of memory and computing + SIMD, the weights are stored in the memristor crossbar array and directly participate in the analog multiplication and addition operation, eliminating the "memory wall" problem of frequent off-chip memory access by traditional DSPs and reducing the data transfer power consumption; through the 4-way parallel multiplication and addition unit (PMAC) and custom instructions (such as VMM4), 4 neuron state updates are completed in a single cycle, and the computing density is increased by 4 times, ultimately increasing the computing energy efficiency ratio by 5 - 8 times. Through the dynamic weight calibration engine based on Lyapunov stability theory, the memristor conductance value is monitored in real time and fed back to the DSP to generate an adaptive PWM pulse (instruction VPULSEGEN) to compensate for the drift. At the same time, non-uniform quantization + reference correction is adopted, 4-bit high-precision quantization (0.05 μA / LSB) is used in the low-current sensitive area, combined with periodic zero-input calibration to suppress non-linear distortion, and finally closed-loop calibration + dynamic quantization is achieved, reducing the error to less than 1%.

[0110] In addition, through hybrid precision dynamic switching, FP16 fixed-point operation is used in the inference stage (the delay is reduced by 40%), and the FP32 floating-point mode is switched to in the training stage to ensure the gradient calculation accuracy. Combined with the hardware acceleration of gradient pulses, 4-way gradient calculations are completed in a single cycle using the custom instruction VGRAD, which is 32 times faster than the software implementation, reducing the iteration cycle, achieving dynamic precision + hardware acceleration, and increasing the training convergence speed by more than 30%. By performing first-order difference compression on the ADC sampling values, combined with the Kalman filter (instruction VKFILTER) to suppress high-frequency noise, and skipping zero-value connections in the weight loading stage to reduce the invalid calculations introduced by noise, non-uniform quantization + noise suppression coding is achieved, and the signal-to-noise ratio is increased by more than 20%. Through instruction-level pipeline optimization, such as a 6-stage pipeline (including memristor prefetch) to hide 90% of the memory latency, the computing throughput reaches 4 neurons / cycle, and the PWM pulse is directly connected for driving. For example, the pulse parameters generated by the VPULSEGEN instruction are directly output through the DAC, bypassing the software protocol stack, and the control delay is greatly reduced. Finally, hardware acceleration + direct connection control is achieved, and the delay is reduced by more than 60%.

[0111] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0112] The above embodiments only express several implementation manners of the present invention, and the description is relatively specific and detailed. However, it should not be construed as a limitation on the protection scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, which all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.

Claims

1. A heterogeneous neural network acceleration device based on the cooperation of DSP and memristor, characterized in that It includes a DSP coprocessor and a memristive crossbar array. The memristive crossbar array interacts with the DSP coprocessor through a PWM interface and maps the conductance values of the memristors to the memory space of the DSP coprocessor; The memristive crossbar array integrates a 12-bit high-precision signal conversion module, a conductance monitoring circuit, and a PWM driving circuit. The high-precision signal conversion module is used to achieve seamless analog-digital signal conversion. The conductance monitoring circuit is used to monitor the change of the memristor conductance value and transmit it to the DSP coprocessor. The PWM driving circuit is used to receive the PWM pulses generated by the DSP coprocessor after calculating the error gradient according to the change of the memristor conductance value, adjust the conductance of the memristive crossbar array, and calibrate and update the conductance-digital activation mapping table; The DSP coprocessor is used to adopt high-precision quantization in the signal-sensitive area and low-precision quantization in the signal-insensitive area through a non-uniform piecewise quantization strategy. The DSP coprocessor is deployed with an extended and customized SIMD instruction set and is used for hardware-level acceleration of heterogeneous neural networks; The SIMD instruction set includes an extended instruction set and a new instruction set. The extended instruction set is used to support SIMD parallel computing to update the neuron states of heterogeneous neural networks, and the new instruction set is used to dynamically calibrate the memristive weights of heterogeneous neural networks based on the Lyapunov stability theory; The conductance of the memristive crossbar array is adjusted through a cross-domain quantization coding algorithm between the DSP coprocessor and the memristive crossbar array; Among them, the process steps of the cross-domain quantization coding algorithm include: Calibrate the initial conductance-current curve of the memristors in the memristive crossbar array to generate a default quantization table; Configure the sampling rate of the high-precision signal conversion module and the interrupt response of the DSP coprocessor; Execute ADC sampling through the high-precision signal conversion module, look up the table to determine the current segment interval to which the current signal obtained by the ADC sampling belongs, and then apply the corresponding quantization formula for quantization; Perform differential coding on the quantized sampling values and then perform Kalman filtering, and write the obtained digital code into the memory mapping area of the DSP coprocessor; Among them, adaptive reference calibration is performed every 1ms; Convert the PWM pulses generated by the DSP coprocessor into an output voltage through non-uniform DAC conversion by the high-precision signal conversion module to adjust the conductance of the memristive crossbar array.

2. The heterogeneous neural network acceleration device based on the cooperation of DSP and memristor according to claim 1, characterized in that The SIMD data path of the DSP coprocessor includes a 128-bit SIMD register and a dedicated memristive weight cache. Each register of the 128-bit SIMD register stores 4 32-bit floating-point numbers or 8 16-bit fixed-point numbers. The dedicated memristive weight cache is a 128B buffer area used to store the quantized weights read from the memristive crossbar array; The SIMD data path is deployed with a parallel multiply-accumulate unit using Booth coding and Wallace tree structure to support 4-way parallel analog multiply-accumulate operations. The activation function accelerator of the parallel multiply-accumulate unit adopts piecewise linear approximation tanh.

3. The heterogeneous neural network acceleration device based on the cooperation of DSP and memristor according to claim 2, characterized in that The DSP coprocessor adopts a double-buffer mechanism or a broadcast mode for neuron state caching.

4. The heterogeneous neural network acceleration device based on the cooperation of DSP and memristor according to claim 2, characterized in that, The extended instruction set includes instruction VLDMEM, instruction VMM4, instruction VTANH4, and instruction VSTNEURON; The instruction VLDMEM is used to load 4 weights from the dedicated memristor weight cache into the SIMD register; The instruction VMM4 is used to perform 4-way parallel multiply-accumulate; The instruction VTANH4 is used to run the 4-way parallel tanh activation function; The instruction VSTNEURON is used to store 4 neuron states into memory.

5. The heterogeneous neural network acceleration device based on the cooperation of DSP and memristor according to claim 4, wherein When the DSP coprocessor updates the neuron states with the memristive crossbar array, it includes: Using a crossbar switch for dynamic routing of weights and input data; In the FP16 format, a 32-bit accumulator is used for accumulation inside the multiply-accumulate unit. When the accumulated result exceeds the FP16 range, an interrupt is automatically triggered to switch to the FP32 format; Among them, the cooperation between the DSP coprocessor and the memristive crossbar array adopts a 5-stage pipeline: instruction fetch, decode, memristor load, execute, and write-back; in the memristor load stage, the next set of weights is prefetched.

6. The heterogeneous neural network acceleration device based on the cooperation of DSP and memristor according to claim 2, wherein The new instruction set includes the instructions VGRAD, VPULSEGEN, VLYAP, and VMEMSYNC; The instruction VGRAD is used to calculate the gradients of 4 weights in parallel; The instruction VPULSEGEN is used to generate a PWM pulse signal corresponding to the gradient; The instruction VLYAP is used to calculate the Lyapunov energy function; The instruction VMEMSYNC is used to synchronize the memristor weight cache of the DSP coprocessor and the memristive crossbar array.

7. The heterogeneous neural network acceleration device based on the cooperation of DSP and memristor according to claim 6, wherein The instruction set hardware architecture adopted by the new instruction set includes a dedicated computing unit and a weight broadcast network. The dedicated computing unit includes a gradient calculation unit and a pulse generation unit; The gradient calculation unit is a 4-way parallel floating-point multiplier that supports the Taylor expansion approximation of the non-linear term and is used to perform dynamic precision switching. The pulse generation unit is used to map the gradient value to the PWM pulse width and select the voltage direction according to the gradient sign bit. The weight broadcast network is used to broadcast a single weight value to 4 of the dedicated computing units simultaneously.

8. The heterogeneous neural network acceleration device based on the cooperation of DSP and memristor according to claim 6, characterized in that, The pipeline corresponding to the new instruction set is a 6-stage pipeline: instruction fetch, decode, weight load, gradient calculation, pulse generation, and write-back; among them, conflict resolution includes data hazards and control hazards.

9. The heterogeneous neural network acceleration device based on the cooperation of DSP and memristor according to claim 1, characterized in that, The non-uniform segmented quantization strategy includes: When the output circuit of the memristive crossbar array is in the low current region, the quantization bit number is allocated as 4 bits and the quantization interval is set to 0.05 μA; When the output circuit of the memristive crossbar array is in the medium current region, the quantization bit number is allocated as 6 bits and the quantization interval is set to 0.2 μA; When the output circuit of the memristive crossbar array is in the high current region, the quantization bit number is allocated as 2 bits and the quantization interval is set to 100 μA.

Citation Information

Patent Citations

  • Neural network face recognition system based on memristor

    CN110443168A

  • Digital architecture supporting analog co-processor

    CN111542826A