Heterogeneous neural network acceleration device based on cooperation of DSP and memristor
By adopting a heterogeneous architecture that coordinates DSP and memristors in the HNN network acceleration circuit, the problems of high power consumption and low scalability in traditional technologies are solved, and more efficient acceleration performance and drift resistance are achieved.
Patent Information
- Application Number
- CN202510509568.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-22
AI Technical Summary
Traditional CMOS devices face high power consumption and low scalability technical bottlenecks when used in the implementation of Hopfield neural networks, resulting in insufficient energy efficiency ratio and drift resistance of the HNN network acceleration circuit.
Using a heterogeneous neural network acceleration device based on the collaboration of DSP and memristors, a hybrid computing architecture of DSP coprocessor and memristor cross-array is combined with dynamic accuracy switching, real-time closed-loop feedback mechanism of Lyapunov stability theory, adaptive cross-domain quantization module and hardware instruction set optimization.
It significantly improves the energy efficiency ratio and drift resistance of the HNN network acceleration circuit, improves the acceleration performance, and achieves higher signal-to-noise ratio and faster training convergence speed.
Smart Images

Figure CN120046677A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of neural network hardware acceleration and hybrid computing, and relates to a heterogeneous neural network acceleration device based on the cooperation of DSP and memristor. Background Art
[0002] As the fourth basic circuit element, the memristor has core characteristics (the resistance value changes with the flowing charge and remains unchanged when powered off), making it an ideal device for simulating biological synapses. It is widely used in fields such as artificial neural networks (ANNs), secure communication, and in-memory computing (IMC). The in-memory computing technology solves the "memory wall" problem of the von Neumann architecture by directly performing calculations within the memory cells, and the energy consumption can be reduced to 12% of the traditional architecture.
[0003] As a recurrent neural network, the Hopfield neural network (HNN) is good at processing combinatorial optimization problems (such as the traveling salesman problem) and associative memory tasks. Its core lies in the dynamic convergence characteristics of the energy function. However, traditional CMOS (complementary metal oxide semiconductor) devices face technical bottlenecks of high power consumption and low scalability when used for network implementation. Memristors are widely used to construct the synaptic weights of HNNs due to their non-volatility and analog computing capabilities. In this field, there are already memristor-HNN hybrid architectures based on FPGAs, 1T1R memristor crossbar array acceleration schemes, analog in-memory computing (IMC) schemes, and dynamic calibration and hybrid quantization technologies. However, the above traditional technologies have insufficient energy efficiency ratio and anti-drift ability in practical applications, resulting in the technical problem of low acceleration performance in the HNN network acceleration circuit. Summary of the Invention
[0004] Aiming at the problems existing in the above traditional technologies, the present invention proposes a heterogeneous neural network acceleration device based on the cooperation of DSP and memristor, which can improve the energy efficiency ratio and anti-drift ability of the HNN network acceleration circuit and enhance the acceleration performance.
[0005] To achieve the above object, the embodiments of the present invention adopt the following technical solutions: Provide a heterogeneous neural network acceleration device based on the cooperation of DSP and memristor, including a DSP coprocessor and a memristor crossbar array. The memristor crossbar array interacts with the DSP coprocessor through a PWM interface and maps the conductance value of the memristor to the memory space of the DSP coprocessor; The memristive crossbar array integrates a 12-bit high-precision signal conversion module, a conductance monitoring circuit, and a PWM driving circuit. The high-precision signal conversion module is used to achieve seamless analog-digital signal conversion. The conductance monitoring circuit is used to monitor the change of the memristive conductance value and transmit it to the DSP co-processor. The PWM driving circuit is used to receive the PWM pulses generated by the DSP co-processor after calculating the error gradient according to the change of the memristive conductance value, adjust the conductance of the memristive crossbar array, and calibrate and update the conductance-digital activation mapping table; The DSP co-processor is used to adopt high-precision quantization in the signal-sensitive area and low-precision quantization in the signal-insensitive area through a non-uniform piecewise quantization strategy. The DSP co-processor is deployed with an extended and customized SIMD instruction set and is used for hardware-level acceleration of heterogeneous neural networks; The SIMD instruction set includes an extended instruction set and a newly added instruction set. The extended instruction set is used to support SIMD parallel computing to update the neuron states of heterogeneous neural networks, and the newly added instruction set is used to dynamically calibrate the memristive weights of heterogeneous neural networks based on Lyapunov stability theory.
[0006] One technical solution in the above technical solutions has the following advantages and beneficial effects: The above heterogeneous neural network acceleration device based on the cooperation of DSP and memristor, through the design of a hybrid computing architecture of DSP and memristive crossbar array, such as dynamic precision switching, real-time closed-loop feedback mechanism based on Lyapunov stability theory, adaptive cross-domain quantization module, and hardware instruction set optimization, etc., makes the hybrid architecture have good balance between performance and versatility, and significantly improves stability through conductance monitoring, adaptive PWM pulse generation, and quantization table update. The adaptive cross-domain quantization module combines noise reduction to achieve a higher signal-to-noise ratio at a lower cost. The hardware level is optimized through a customized instruction set, which is more suitable for HNN computing tasks, improves the energy efficiency ratio and anti-drift ability of the HNN network acceleration circuit, and improves the acceleration performance. Description of the Drawings
[0007] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the traditional technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or the traditional technologies. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0008] Figure 1 It is a schematic diagram of the architecture of a heterogeneous neural network acceleration device based on the cooperation of DSP and memristor in an embodiment; Figure 2 It is a schematic diagram of the cross-domain quantization coding architecture in an embodiment; Figure 3 It is a schematic diagram of the process of dynamic weight calibration in an embodiment; Figure 4 Schematic diagram of SIMD parallel computing in an embodiment. Detailed implementation manners
[0009] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, but not to limit the present invention. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art belonging to the technical field of the present invention. The terms used in the description of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention.
[0010] It should be noted that referring to "embodiment" herein means that a specific feature, structure or characteristic described in connection with the embodiment can be included in at least one embodiment of the present invention. The phrase is shown at various positions in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art can understand that the embodiments described herein can be combined with other embodiments. The term "and / or" used in the description and claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0011] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings in the embodiments of the present invention.
[0012] In the memristor-HNN hybrid architecture based on FPGA, the FPGA is used as a digital controller, the memristor array is used for analog synaptic weight storage, and the analog-digital interaction is realized through the analog-to-digital converter ADC / digital-to-analog converter DAC. In the image encryption task, the delay is 8.5 ms / frame and the energy efficiency ratio is 8.3 GOPS / W. However, limited by the FPGA logic unit, it supports a maximum of 256×256 network and lacks a dynamic calibration mechanism. The conductance value of the memristor is easily affected by temperature and voltage fluctuations, resulting in a weight drift error >5% and obvious quantization noise of the analog signal. In addition, the memristor-HNN hybrid architecture implemented by FPGA has a high delay and a slow training convergence speed.
[0013] In the 1T1R memristor crossbar array acceleration scheme, each memristor unit is connected in series with a transistor (1T1R) structure to suppress the sneak path problem and support high-density integration. It is used for convolutional neural network acceleration, with a 10-fold increase in computing power. However, the transistors increase the chip area and power consumption, making it difficult to achieve ultra-large-scale expansion. In the analog in-memory computing (IMC) scheme, analog multiply-accumulate operations are directly performed in the memristor array to avoid data movement. For example, a commercially available Flash-based chip has an energy efficiency ratio as high as 25 TOPS / W, but it relies on high-precision ADC / DAC, has a high cost, and is vulnerable to noise interference.
[0014] Dynamic calibration and hybrid quantization techniques use a memristor weight calibration algorithm based on Lyapunov theory to suppress drift through feedback regulation, reducing the weight drift error to 2%. However, its calibration period is long (greater than 10 ms), resulting in poor real-time performance. Some systems rely on pre-training or complex compensation steps and do not combine non-uniform quantization, resulting in insufficient accuracy in the low-current region. Among them, the resistance change of traditional memristors (TiO 2 ) depends on ion migration, is vulnerable to environmental interference, and the above traditional technologies mostly use uniform quantization and do not optimize for the non-linear characteristics of memristors, resulting in insufficient accuracy in the low-current region.
[0015] When using a memory-computation separation architecture in the above traditional technologies, off-chip memory needs to be frequently read and written, resulting in a low energy efficiency ratio. The serial computing mode lacks parallel acceleration units (such as SIMD (Single Instruction, Multiple Data)), and cannot efficiently process large-scale matrix operations. In addition, there is no closed-loop feedback mechanism, and real-time conductance adjustment is not achieved by combining Lyapunov stability theory, and there is insufficient hardware support, lacking dedicated instructions to accelerate gradient calculation and pulse generation.
[0016] Therefore, the present invention will provide a heterogeneous neural network acceleration device based on the cooperation of DSP and memristors, which solves the memory-computation separation problem through a "digital-analog hybrid computing architecture" and "dynamic adaptive instruction optimization", realizes in-situ update of memristor weights and efficient non-linear computing of DSP, improves the energy efficiency ratio and anti-drift ability of the HNN network acceleration circuit, and thus improves the acceleration performance.
[0017] In one embodiment, as Figure 1As shown in the figure, a heterogeneous neural network acceleration device based on the cooperation of DSP and memristor is provided, including a DSP coprocessor (DSP) and a memristive crossbar array. The memristive crossbar array interacts with the DSP coprocessor through a PWM interface and maps the conductance value of the memristor to the memory space of the DSP coprocessor. The memristive crossbar array integrates a 12-bit high-precision signal conversion module (ADC / DAC), a conductance monitoring circuit, and a PWM driving circuit. The high-precision signal conversion module is used to achieve seamless analog-digital signal conversion. The conductance monitoring circuit is used to monitor the change of the memristor conductance value and transmit it to the DSP coprocessor. The PWM driving circuit is used to receive the PWM pulses generated by the DSP coprocessor after calculating the error gradient according to the change of the memristor conductance value, adjust the conductance of the memristive crossbar array, and calibrate and update the conductance-digital activation mapping table. The DSP coprocessor is used to adopt high-precision quantization in the signal-sensitive area and low-precision quantization in the signal-insensitive area through a non-uniform segmented quantization strategy. The DSP coprocessor is deployed with an extended and customized SIMD instruction set (represented by SIMD units) and is used for hardware-level acceleration of heterogeneous neural networks; the SIMD instruction set includes an extended instruction set and a new instruction set. The extended instruction set is used to support SIMD parallel computing to update the neuron states of heterogeneous neural networks, and the new instruction set is used to dynamically calibrate the memristive weights of heterogeneous neural networks (i.e., the calibration algorithm) based on the Lyapunov stability theory.
[0018] It can be understood that in the first part of the design, such as Figure 1 In the architecture of the heterogeneous neural network acceleration device based on the cooperation of DSP and memristor as shown in the figure, the memristive crossbar array serves as a synaptic weight storage and analog multiply-accumulate unit, represents the weight magnitude through the conductance value (the reciprocal of the resistance), and supports in-situ update (without external storage) to reduce power consumption; it integrates a 12-bit high-precision signal conversion module ADC / DAC for seamless analog-digital signal conversion, enabling the subsequent circuit to directly process analog signals. The memristive crossbar array directly interacts with the DSP through a PWM interface and supports dynamic conductance value mapping to the memory space of the DSP.
[0019] The DSP coprocessor integrates an "adaptive cross-domain quantization module" and adopts a non-uniform segmented quantization strategy. It uses high-precision quantization for signal-sensitive areas (such as small currents) to reduce noise sensitivity, and uses low-precision quantization for signal-insensitive areas to effectively save resources. It has extended and customized instruction sets (such as vector Lyapunov function calculation instructions, memristor conductance gradient acceleration instructions) to achieve hardware-level acceleration of the HNN dynamic equation. It deploys a "dynamic weight calibration engine", which is a real-time feedback system based on the Lyapunov stability theory and is used to suppress the conductance drift of the memristor.
[0020] The implementation principle of the "Dynamic Weight Calibration Engine" is as follows: First, the conductance monitoring circuit monitors the change in the memristor conductance value. Then, the DSP calculates the error gradient and generates PWM pulses to adjust the conductance (for example, long pulses are used to increase the conductance, and short pulses are used to decrease the conductance). Finally, the "Conductance-Digital Activation Mapping Table" is updated to maintain network stability. The conductance monitoring circuit can adopt the existing monitoring circuits in the field of memristors, which can automatically measure the output conductance through the current-conductance conversion relationship of the memristor.
[0021] The above heterogeneous neural network acceleration device based on the cooperation of DSP and memristor, through the design of a hybrid computing architecture that combines DSP and memristive crossbar arrays, such as dynamic precision switching, real-time closed-loop feedback mechanism based on Lyapunov stability theory, adaptive cross-domain quantization module, and hardware instruction set optimization, etc., enables the hybrid architecture to have good balance between performance and generality. And it significantly improves stability through conductance monitoring, adaptive PWM pulse generation, and quantization table update, realizes higher signal-to-noise ratio at lower cost through the joint noise reduction of the adaptive cross-domain quantization module, and achieves hardware-level optimization through the customized instruction set, making it more suitable for HNN computing tasks, improving the energy efficiency ratio and anti-drift ability of the HNN network acceleration circuit, and enhancing the acceleration performance.
[0022] Furthermore, regarding the cross-domain quantization coding algorithm for DSP and memristor, its corresponding cross-domain quantization coding architecture is as Figure 2 shown, including non-uniform segmented quantization (adaptive quantization), adaptive reference correction (dynamic weight calibration), noise suppression coding, and algorithm implementation process, etc. Specifically, it can be divided into the digital domain (DSP) and the analog domain (memristor). The improvements in the digital domain (DSP) include parallel computing, adaptive quantization, dynamic weight calibration, and instruction set extension. The improvements in the analog domain (memristor) include PWM drive, weight mapping, and analog multiply-accumulate. The process of dynamic weight calibration is as Figure 3 shown.
[0023] The quantization table structure of non-uniform segmented quantization is shown in Table 1: Table 1
[0024] The design basis of non-uniform segmented quantization is: according to the volt-ampere characteristic curve of the memristor, more quantization bits are allocated in the low-current region (i.e., the sensitive region) to suppress noise; sparse quantization is performed in the high-current region (i.e., the non-sensitive region) to save resources. The formula expression of non-uniform segmented quantization is as follows: ; where represents the quantization result, round the () function is used to round the value to the specified number of bits, V minRepresents the minimum voltage amplitude of the memristor, V max and represents the maximum voltage amplitude of the memristor. Through piecewise linear mapping, the analog signal x is mapped to a 12-bit digital code (0 - 4095).
[0025] The reference sampling for adaptive reference calibration is that the DSP periodically reads the output current when the input of the memristor is zero I offset (drift noise). The offset compensation for adaptive reference calibration is to update the V min and V max in the formula of non-uniform piecewise quantization in real time through the following formula: V min = V min + α I offset; V max = V max + β I offset; where α and β are calibration coefficients and can be calibrated through experiments.
[0026] The quantization table update for adaptive reference calibration is to recalculate the current segmentation interval according to the new calibrated reference and write it into the look-up table (LUT) of the DSP.
[0027] The noise suppression coding includes differential coding and Kalman filtering. Among them, differential coding is to continuously perform differential operations on the adjacent sampling values Q t and Q t-1 in time: Δ Q = Q t - Q t-1; Only the sign bit and magnitude bit of the differential signal Δ Q are transmitted (for example, 4 bits represent the change amount) to reduce the transmission bandwidth. Kalman filtering is to filter the differential signal at the DSP end to suppress high-frequency noise: ; where K is the dynamically adjusted Kalman gain.
[0028] In some embodiments, the conductance of the memristive crossbar array is adjusted by a cross-domain quantization coding algorithm between the DSP coprocessor and the memristive crossbar array. The process steps of the cross-domain quantization coding algorithm may include: Calibrate the initial conductance-current curve of the memristors in the memristive crossbar array to generate a default quantization table; Configure the sampling rate of the high-precision signal conversion module and the interrupt response of the DSP coprocessor; Perform ADC sampling through the high-precision signal conversion module, look up the table to determine the current segment interval to which the current signal obtained by the ADC sampling belongs, and then apply the corresponding quantization formula for quantization; Perform differential coding on the quantized sampling values and then perform Kalman filtering, and write the obtained digital code into the memory mapping area of the DSP coprocessor; among them, adaptive reference calibration is performed every 1 ms; Convert the PWM pulse generated by the DSP coprocessor into an output voltage through non-uniform DAC conversion by the high-precision signal conversion module to adjust the conductance of the memristive crossbar array.
[0029] Specifically, the algorithm implementation process of the cross-domain quantization coding algorithm is as follows: Initialization stage; including (1) calibrating the initial conductance-current curve of the memristors to generate a default quantization table (for predefined current segment intervals), and (2) configuring the sampling rate (1 MSPS) of the high-precision signal conversion module ADC / DAC and the DSP interrupt response.
[0030] Enter the real-time quantization stage; including While (system running) and reverse control (DAC output). Among them, in the While (system running) stage, ADC sampling is performed to obtain the current signal I(t), look up the default quantization table to determine the current segment interval to which the current signal I(t) belongs, then apply the corresponding quantization formula, perform differential coding and Kalman filtering, and write the digital code obtained after differential coding and Kalman filtering into the memory mapping area of the DSP. Adaptive reference calibration is performed every 1 ms (update V min and V max ).
[0031] Reverse control (DAC output) is to convert the gradient pulse (PWM pulse) generated by the DSP into an output voltage through non-uniform DAC V out : ; The second part of the design, the SIMD-based parallel update microarchitecture of neuron states, is designed to include a SIMD data path, data loading and storage, a control unit, and an instruction set extension, and realizes high-performance computing through customizing SIMD + mixed precision.
[0032] In one embodiment, the SIMD data path of the DSP co-processor includes a 128-bit SIMD register and a dedicated memristive weight cache. Each register of the 128-bit SIMD register stores 4 32-bit floating-point numbers or 8 16-bit fixed-point numbers. The dedicated memristive weight cache is a 128B buffer for storing the quantized weights read from the memristive crossbar array. The SIMD data path is deployed with a parallel multiply-accumulate unit using Booth encoding and Wallace tree structure to support 4-way parallel analog multiply-accumulate operations. The activation function accelerator of the parallel multiply-accumulate unit uses piecewise linear approximation of tanh. The schematic diagram of SIMD parallel computing is as Figure 4 shown.
[0033] Specifically, the register bank of the SIMD data path includes a 128-bit SIMD register (V0-V15) and a dedicated memristive weight cache. Each register of the 128-bit SIMD register stores 4 32-bit floating-point numbers (FP32) or 8 16-bit fixed-point numbers (FP16). The dedicated memristive weight cache is a 128B buffer for storing the quantized weights (12-bit precision) read from the memristive crossbar array.
[0034] The parallel multiply-accumulate unit (PMAC) supports 4-way parallel analog multiply-accumulate operations and completes the following multiply-accumulate output in a single cycle: ; where, represents the quantized weight of the memristor in the i th row and j th column, X j represents the j th input signal, B i represents the reference threshold for adjusting the output. Hardware optimization uses Booth encoding and Wallace tree structure to reduce the multiplier delay.
[0035] The activation function accelerator uses piecewise linear approximation of tanh: splitting the non-linear function into 4 segments of linear approximation, implemented by looking up a table (LUT), and supporting 4-way parallel computing.
[0036] ; where y represents the activation function output, x represents the neuron input signal (since the activation function is part of the nervous system, the signal x here is the same expression as the neuron state value and can be an analog signal). The single-cycle completion of 4 activation value calculations is achieved through SIMD instructions.
[0037] In one embodiment, the DSP coprocessor uses a double-buffer mechanism or a broadcast mode to cache neuron states.
[0038] Specifically, the data loading and storage part includes: First, load the memristive weights. Batch load the quantized weights (12 bits) of the memristive crossbar array into the dedicated memristive weight cache through DMA and automatically expand them to the FP16 / FP32 format. Then, perform data alignment. Align the quantized weights according to the 4-way parallel requirement (for example, store the weight matrix separately in 4×4). Next, complete the neuron state caching, which can use a double-buffer mechanism (that is, two 128B buffer areas read and write alternately to hide the data transfer latency) or a broadcast mode (that is, broadcast a single input signal to 4 parallel computing units (applicable to the fully connected layer)).
[0039] The control unit part includes an instruction scheduler and dynamic precision switching. The instruction scheduler is used to parse custom SIMD instructions (such as VMM4 vector matrix multiplication), dynamically allocate computing resources, and support instruction-level parallelism (ILP), such as simultaneously executing multiply-add and activation functions. Dynamic precision switching is to switch the FP16 / FP32 format through the status register (CR0) to enable the hardware to automatically adjust the data path bit width.
[0040] In one embodiment, the instruction set extension is to support SIMD parallel computing. The extended instruction set is shown in Table 2 below: Table 2
[0041] Among them, Vd represents the destination vector register, Vw represents the weight vector register, Vx represents the neuron i state vector register, Vs represents the gradient value vector register, Rs represents the source base register, which is used to calculate the base address of the memory address, Rd represents the destination base register, which is used to store the target address base address to be written to the memory, and offset represents the address offset, which is added to the base address to generate a valid address. Based on the foregoing design, an example of HNN state update is as follows, parallel computing the state updates of 4 neurons: Call the instruction VLDMEMV1, [R0]; load the weight block W[0:3]; Call the instruction VLDMEMV2, [R1]; load the input vector X[0:3]; Call the instruction VMM4V3, V1, V2; (4-way multiply-add); Call the instruction VTANHV4, V3; V4 = tanh(V3); Call the instruction VSTNEURON[R2], V4; store the neuron state Y[0:3].
[0042] In some embodiments, when updating the neuron state between the DSP coprocessor and the memristive crossbar array, the following processing may be included: Use a crossbar switch to dynamically route weights and input data; In the FP16 format, a 32-bit accumulator is used inside the multiply-accumulate unit for accumulation. When the accumulated result exceeds the FP16 range, an interrupt is automatically triggered to switch to the FP32 format; The 5-stage pipeline adopted is instruction fetch, decode, memristor load, execute, and write-back. Among them, in the memristor load stage, the next set of weights is prefetched.
[0043] Specifically, the corresponding hardware optimizations include a weight-input alignment network, dynamic precision fusion, and a memristor-DSP cooperative pipeline. The weight-input alignment network uses a crossbar switch to dynamically route weights and input data, supports non-contiguous memory access, and its application scenario can be when processing sparse connections, skipping zero-weight calculations, reducing power consumption by 30%.
[0044] Dynamic precision fusion includes mixed-precision calculation, that is, in the FP16 format, a 32-bit accumulator is used inside the multiply-accumulate unit to avoid precision loss; and overflow protection, that is, when the accumulated result exceeds the FP16 range, an interrupt is automatically triggered to switch to the FP32 format. The memristor-DSP cooperative pipeline adopts a 5-stage pipeline: 1. Instruction Fetch (IF) → 2. Decode (ID) → 3. Memristor Load (MEM) → 4. Execute (EX) → 5. Write-back (WB); The optimization of the MEM stage can hide the memristor read latency by prefetching the next set of weights.
[0045] In one embodiment, for the third part of the design, the newly added instruction set in the DSP instruction set extension for HNN is designed as shown in Table 3 below: Table 3
[0046] The detailed description of the key instructions is as follows: VGRAD: Parallel gradient calculation instruction, whose function is to calculate the weight gradient based on the Lyapunov stability theory , and the formula is: ; Where is the learning rate, x i , x j is the neuron state.
[0047] The hardware implementation of the instruction can be as follows: The inputs are Vw (weights), V x (neuron i state) and V y(Neuron j state); The computing unit uses a 4-way parallel floating-point multiplier-accumulator, supporting fast approximation of non-linear terms ; The output is the gradient value stored in the Vd register, in the FP32 format. An example code is: VGRADV0, V1, V2, V3; .
[0048] VPULSEGEN: Gradient pulse generation instruction, whose function is to convert the gradient value into the PWM control parameters (pulse width, amplitude) of the memristor; The mapping rule is that the pulse width = , polarity = , where K p is the scaling factor (configured by the register Rt). The hardware implementation of the instruction can be as follows: The inputs are Vs (gradient value) and Rt ( K p and pulse parameters); The output is Vd storing 4 PWM control words (32 bits, including width and polarity); The pipeline optimization uses a pulse parameter prediction calculation module directly connected to the DAC interface to reduce control latency.
[0049] VLYAP: Lyapunov energy function instruction, whose function is to calculate the HNN energy function E , for stability judgment: ; Its hardware implementation can be as follows: The input is the neuron state V state (4 neuron states x0 - x3), and the calculation process is to first load a 4×4 weight block from the dedicated memristive weight cache, then calculate in parallel, then sum and take the negative, and store the obtained result in Vd. An example code is: VLYAP V0, V1; .
[0050] VMEMSYNC: Memristive weight synchronization instruction, whose function is to synchronize the weight update in the DSP memory to the memristive crossbar array, supporting batch transfer. Its hardware implementation is as follows: First, align the data address in 4×4 weight blocks (128 bits), and then send the conductance value update command through the SPI interface to automatically trigger the PWM drive circuit of the memristive crossbar array.
[0051] In one embodiment, the instruction set hardware architecture adopted by the new instruction set includes a dedicated computing unit and a weight broadcast network. The dedicated computing unit includes a gradient computing unit and a pulse generation unit. The gradient computing unit is a 4-way parallel floating-point multiplier supporting Taylor expansion approximation of non-linear terms, used to perform dynamic precision switching. The pulse generation unit is used to map the gradient value to the PWM pulse width and select the voltage direction according to the gradient sign bit. The weight broadcast network is used to broadcast a single weight value to 4 dedicated computing units simultaneously.
[0052] Specifically, the hardware architecture of the newly added instruction set in the DSP instruction set extension for HNN includes a dedicated computing unit, data path optimization, and pipeline design, and an application example is given. Among them, the dedicated computing unit includes a gradient calculation unit (GCU, a 4-way parallel floating-point multiplier that supports the Taylor expansion approximation (3rd order) of non-linear terms and a pulse generation unit (PCU). The pulse generation unit consists of a time-to-digital converter (TDC) and a polarity controller. The time-to-digital converter is used to map the gradient value to the PWM pulse width (resolution 1 ns), and the polarity controller is used to select the voltage direction according to the gradient sign bit (such as +Vdd / -Vdd).
[0053] Data path optimization adopts a weight broadcast network and dynamic precision switching. Among them, the weight broadcast network allows a single weight value to be broadcast to 4 dedicated computing units simultaneously, which is suitable for fully connected HNN; dynamic precision switching is used to reduce the memory width in the FP16 format by using a compressed weight format (12-bit mantissa) inside the gradient calculation unit GCU.
[0054] In some embodiments, the pipeline design corresponding to the newly added instruction set adopts a 6-stage pipeline: instruction fetch, decode, weight loading, gradient calculation, pulse generation, and write-back. Its conflict resolution includes data hazards and control hazards. Data hazards are avoided through register renaming and out-of-order execution, and control hazards are optimized by using a branch predictor for loop unrolling (the number of HNN iterations is fixed).
[0055] The application example can be the following HNN weight update process, which updates 4 memristor weights in parallel: Call instruction VLDMEMV1, [R0]; Load the current weight W [0:3]; Call instruction VLDMEMV2, [R1]; Load the neuron state X[0:3]; Call instruction VGRADV3, V1, V2, V2; Calculate the gradient ; Call instruction VPULSEGENV4, V3, R2; Generate PWM pulse parameters (R2 = K p ); Call instruction VMEMSYNC[R3], 4; Synchronize the weight to the memristor crossbar array (R3 is the memristor address).
[0056] The specific implementation method can be as follows: In the hardware configuration, the memristor crossbar array adopts a 1T1R structure with a scale of 256×256, TiO 2Memristive device, integrated PWM drive circuit; the DSP chip selects an existing DSP and expands the custom instruction set (such as the VLYAP instruction to accelerate the calculation of the Lyapunov function); the cross-domain interface uses a 12-bit 1MSPS ADC / DAC module, supporting the dynamic range 5V.
[0057] The software implementation process includes: (1) Initialization stage, the DSP reads the memristive conductance value through the SPI interface, constructs the weight matrix and loads it into the L2 cache. (2) Inference stage, the input signal passes through the analog multiplication and addition of the memristive crossbar array, and the output current is converted into a digital signal by the ADC. The DSP calls the extended instruction set to execute the activation function (piecewise linear approximation of tanh) and state iteration, and the result is transmitted to the output cache through DMA, with the whole process pipelined. (3) Training stage, the DSP generates gradient pulses based on the error backpropagation, adjusts the memristive conductance through the DAC, and the dynamic weight calibration engine monitors the conductance value in real time, updates the parameters of the conductance-digital activation mapping table and feeds back to the training algorithm.
[0058] The above heterogeneous neural network acceleration device based on the cooperation of DSP and memristor, compared with the traditional technology, through the integration of memory and computing + SIMD, the weights are stored in the memristive crossbar array and directly participate in the analog multiplication and addition operation, eliminating the "memory wall" problem of frequent off-chip memory access by the traditional DSP and reducing the data transfer power consumption; through the 4-way parallel multiplication and addition unit (PMAC) and custom instructions (such as VMM4), 4 neuron state updates are completed in a single cycle, and the computing density is increased by 4 times, ultimately making the computing energy efficiency ratio increase by 5-8 times. Through the dynamic weight calibration engine based on the Lyapunov stability theory, the memristive conductance value is monitored in real time and fed back to the DSP to generate an adaptive PWM pulse (instruction VPULSEGEN) to compensate for the drift. At the same time, non-uniform quantization + reference correction is adopted, 4-bit high-precision quantization (0.05μA / LSB) is used in the low-current sensitive area, combined with periodic zero-input calibration to suppress nonlinear distortion, and finally closed-loop calibration + dynamic quantization is realized, reducing the error to less than 1%.
[0059] In addition, through dynamic switching of mixed precision, FP16 fixed-point arithmetic is used in the inference stage (with a 40% reduction in latency), and the training stage switches to the FP32 floating-point mode to ensure the accuracy of gradient calculation. Combined with hardware acceleration of gradient pulses, the custom instruction VGRAD can complete 4-way gradient calculation in a single cycle, which is 32 times faster than software implementation, reducing the iteration cycle. By achieving dynamic precision + hardware acceleration, the training convergence speed is increased by more than 30%. Through first-order difference compression of ADC sampling values, combined with a Kalman filter (instruction VKFILTER) to suppress high-frequency noise, and skipping zero-value connections during the weight loading stage to reduce invalid calculations introduced by noise, non-uniform quantization + noise suppression coding is achieved, with the signal-to-noise ratio increased by more than 20%. Through instruction-level pipeline optimization, such as a 6-stage pipeline (including memristor prefetch) to hide 90% of the memory latency, the computing throughput reaches 4 neurons per cycle, and PWM pulse direct connection drive is used. For example, the pulse parameters generated by the VPULSEGEN instruction are directly output through the DAC, bypassing the software protocol stack, and the control latency is greatly reduced. Finally, hardware acceleration + direct connection control is achieved, with the latency reduced by more than 60%.
[0060] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0061] The above embodiments only represent several implementation manners of the present invention, and the description is relatively specific and detailed, but it should not be understood as a limitation on the protection scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, which all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.
Claims
1. A heterogeneous neural network acceleration device based on DSP and memristor collaboration, characterized in that: It includes a DSP coprocessor and a memristor cross array, wherein the memristor cross array interacts with the DSP coprocessor through a PWM interface and maps the conductance value of the memristor to the memory space of the DSP coprocessor; The memristor crossbar array integrates a 12-bit high-precision signal conversion module, a conductance monitoring circuit and a PWM drive circuit. The high-precision signal conversion module is used to realize seamless conversion of analog-to-digital signals. The conductance monitoring circuit is used to monitor the change of the memristor conductance value and transmit it to the DSP coprocessor. The PWM drive circuit is used to receive the PWM pulse generated by the DSP coprocessor after calculating the error gradient according to the change of the memristor conductance value, adjust the conductance of the memristor crossbar array, and calibrate and update the conductance-digital activation mapping table. The DSP coprocessor is used to adopt high-precision quantization in signal-sensitive areas and low-precision quantization in signal-insensitive areas through a non-uniform segmented quantization strategy. The DSP coprocessor is deployed with an extended customized SIMD instruction set and is used for hardware-level acceleration of heterogeneous neural networks; the SIMD instruction set includes an extended instruction set and a newly added instruction set. The extended instruction set is used to support SIMD parallel computing to update the neuron state of the heterogeneous neural network, and the newly added instruction set is used to dynamically calibrate the memristor weights of the heterogeneous neural network based on Lyapunov stability theory.
2. The heterogeneous neural network acceleration device based on DSP and memristor collaboration according to claim 1 is characterized in that: The DSP coprocessor and the memristor crossbar array adjust the conductance of the memristor crossbar array through a cross-domain quantization coding algorithm; wherein the process steps of the cross-domain quantization coding algorithm include: Calibrate the initial conductance-current curve of the memristor of the memristor crossbar array to generate a default quantization table; Configure the sampling rate of the high-precision signal conversion module and the interrupt response of the DSP coprocessor; ADC sampling is performed by the high-precision signal conversion module, and a table is looked up to determine the current segment interval to which the current signal obtained by ADC sampling belongs, and then a corresponding quantization formula is applied for quantization; Performing differential encoding on the quantized sampled values and then performing Kalman filtering, and writing the obtained digital code into the memory mapping area of the DSP coprocessor; wherein the adaptive reference correction is performed once every 1 ms; The PWM pulse generated by the DSP coprocessor is converted into an output voltage by non-uniform DAC via the high-precision signal conversion module, thereby adjusting the conductance of the memristor cross array.
3. The heterogeneous neural network acceleration device based on DSP and memristor collaboration according to claim 1 or 2, characterized in that: The SIMD data path of the DSP coprocessor includes a 128-bit SIMD register and a dedicated memristor weight cache, each of the 128-bit SIMD registers stores 4 32-bit floating-point numbers or 8 16-bit fixed-point numbers, and the dedicated memristor weight cache is a 128B cache area for storing quantized weights read from the memristor crossbar array; The SIMD data path is deployed with a parallel multiplication-addition unit using Bozeman coding and Wallace tree structure to support 4-way parallel simulation multiplication-addition operations. The activation function accelerator of the parallel multiplication-addition unit uses piecewise linear approximation tanh.
4. The heterogeneous neural network acceleration device based on DSP and memristor collaboration according to claim 3 is characterized in that: The DSP coprocessor adopts a double buffer mechanism or a broadcast mode to cache neuron states.
5. The heterogeneous neural network acceleration device based on DSP and memristor collaboration according to claim 3 is characterized in that: The extended instruction set includes instruction VLDMEM, instruction VMM4, instruction VTANH4 and instruction VSTNEURON; The instruction VLDMEM is used to load 4 weights from a dedicated memristor weight cache into a SIMD register; The instruction VMM4 is used to perform 4-way parallel multiplication and addition; The instruction VTANH4 is used to run a 4-way parallel tanh activation function; The instruction VSTNEURON is used to store the states of 4 neurons into the memory.
6. The heterogeneous neural network acceleration device based on DSP and memristor collaboration according to claim 5 is characterized in that: When the DSP coprocessor and the memristor crossbar array perform neuron state updating, it includes: Use crossbar switches to dynamically route weights and input data; In FP16 format, the multiplication and addition unit uses a 32-bit accumulator for accumulation. When the accumulated result exceeds the FP16 range, an interrupt is automatically triggered to switch to FP32 format. The DSP coprocessor and the memristor crossbar array cooperate in a 5-stage pipeline: instruction fetch, decoding, memristor loading, execution and write-back; the next set of weights are pre-fetched in the memristor loading stage.
7. The heterogeneous neural network acceleration device based on DSP and memristor collaboration according to claim 3 is characterized in that: The newly added instruction set includes instruction VGRAD, instruction VPULSEGEN, instruction VLYAP and instruction VMEMSYNC; The instruction VGRAD is used to calculate the gradients of 4 weights in parallel; The instruction VPULSEGEN is used to generate a PWM pulse signal corresponding to the gradient; The instruction VLYAP is used to calculate the Lyapunov energy function; The instruction VMEMSYNC is used to synchronize the DSP coprocessor with the memristor weight cache of the memristor crossbar array.
8. The heterogeneous neural network acceleration device based on DSP and memristor collaboration according to claim 7 is characterized in that: The instruction set hardware architecture adopted by the newly added instruction set includes a dedicated computing unit and a weight broadcast network, and the dedicated computing unit includes a gradient computing unit and a pulse generating unit; The gradient calculation unit is a 4-way parallel floating-point multiplier that supports Taylor expansion approximation of nonlinear terms, and is used to perform dynamic precision switching. The pulse generation unit is used to map the gradient value to a PWM pulse width and select the voltage direction according to the gradient sign bit. The weight broadcast network is used to broadcast a single weight value to the four dedicated computing units at the same time.
9. The heterogeneous neural network acceleration device based on DSP and memristor collaboration according to claim 7 is characterized in that: The pipeline corresponding to the newly added instruction set is a 6-stage pipeline: instruction fetch, decoding, weight loading, gradient calculation, pulse generation and write back; wherein conflict resolution includes data hazard and control hazard.
10. The heterogeneous neural network acceleration device based on DSP and memristor collaboration according to claim 1 is characterized in that: The non-uniform segmented quantization strategy includes: When the output circuit of the memristor crossbar array is in a low current region, the number of quantization bits is allocated to 4 bits and the quantization interval is set to 0.05 μA; When the output circuit of the memristor crossbar array is in the medium current region, the number of quantization bits is allocated to 6 bits and the quantization interval is set to 0.2 μA; When the output circuit of the memristor crossbar array is in a high current region, the number of quantization bits is allocated to 2 bits and the quantization interval is set to 100 μA.
Citation Information
Patent Citations
Automatic vectorizing method for heterogeneous SIMD expansion components
CN103279327A
Neural network online learning system based on a memristor
CN109800870A
Neural network face recognition system based on memristor
CN110443168A
Training of artificial neural networks
CN111279366A
Digital architecture supporting analog co-processor
CN111542826A
Cited By
Parallel polynomial gradient calculation system and method based on CMOS (complementary metal oxide semiconductor) neural synaptic transistor
CN121981183A
Parallel polynomial gradient computation system and method based on cmos synapse transistor
CN121981183B