Convolutional Neural Network Hardware Acceleration System Based on Hall Bar Array

By designing a convolutional neural network hardware acceleration system based on Hall bar array, using FPGA and multiple DAC modules to reduce data transfer, the performance bottleneck caused by data transfer during convolutional neural network computing is solved, and efficient calculation and low power consumption are achieved.

CN114997383BActive Publication Date: 2025-06-13HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210609637.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-31
Publication Date
2025-06-13
Estimated Expiration
2042-05-31

AI Technical Summary

Technical Problem

Convolutional neural networks require a large amount of data to be moved during the computing process, resulting in the transfer power consumption exceeding the computing power consumption, causing performance bottlenecks, and it is difficult to achieve intelligent edge computing with high computing power consumption ratio.

Method used

A hardware acceleration system for convolutional neural network based on Hall bar array is designed, and the data is input and output controlled by FPGA, and the data is converted into analog current signals through multiple DAC modules, and it is directly input into the Hall bar array for convolutional operations to reduce data transfer.

Benefits of technology

It realizes efficient convolutional operations, reduces data transfer, improves calculation speed, reduces power consumption, and opens up a new path to intelligent edge computing with high computing power consumption ratio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114997383B_ABST
    Figure CN114997383B_ABST
Patent Text Reader

Abstract

The present invention discloses a convolutional neural network hardware acceleration system based on a Hall bar array. The system includes an FPGA, a multiplexed DAC module, a convolutional operation module, an activation pooling module, and an ADC circuit. The FPGA schedules and controls the input and output data of the system and each operation link, and utilizes the characteristic of the Hall bar device for storage and computing integration to complete the convolutional operation in the form of a hardware circuit in the Hall bar array. Then, the convolutional result is used to implement the activation and maximum pooling operations through a hardware integrated circuit with a controllable discharge path, and then the ADC circuit samples and sends it to the FPGA. This system is a storage and computing integrated convolutional neural network system that implements convolutional, activation, and pooling operations in hardware in an analog circuit manner, with the advantages of high speed and high degree of parallelization, which can reduce the computing burden of the processor in the application system and provide high-cost-effective convolutional neural network computing capabilities for edge devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of convolutional neural network hardware design, and relates to an acceleration system with an FPGA control and a spin Hall bar device as the core, specifically to a convolutional neural network hardware acceleration system based on a Hall bar array. Background Art

[0002] With the rapid development of the Internet and the continuous improvement of computer hardware computing power, the application of convolutional neural networks has been continuously expanded, playing an important role in fields such as economy and military. With the development of convolutional neural networks, the large number of weight parameters included therein also impose increasing requirements on computing power, storage space, and power consumption. If traditional memories are used, a large amount of data transfer is necessarily involved in the operation process, resulting in the transfer power consumption exceeding the computing power consumption.

[0003] If a system can reduce the data transfer during the operation while improving the convolutional operation speed, it can not only solve the current bottleneck of convolutional operations, but also open up a new path for realizing intelligent edge computing with a high computing power consumption ratio. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention proposes a convolutional neural network hardware acceleration system based on a Hall bar array, which makes full use of the input and output of data by the FPGA and the control of each operation link, and designs a hardware acceleration system integrating storage and calculation based on a Hall bar device with storage and calculation in one, and realizes convolutional, activation, and pooling operations in the form of an analog circuit to obtain a convolutional neural network system with multiple parallel high computing speeds.

[0005] The convolutional neural network hardware acceleration system based on a Hall bar array includes an FPGA, a multi-channel DAC module, a convolutional operation module, an activation pooling module, and an ADC circuit.

[0006] The FPGA serves as the input and output interface of data and schedules and controls other modules in the system. The FPGA maps the convolutional kernels obtained by training to the anomalous Hall resistance values of the corresponding Hall bars in the Hall bar array, and sends the image data to the multi-channel DAC module.

[0007] The multi-channel DAC module receives the image data according to the scheduling instructions of the FPGA and converts the sampled data into an output analog current signal.

[0008] Preferably, the amplitude of the analog current signal output by the multi-channel DAC module does not exceed the critical current value that causes the Hall resistance value to change.

[0009] The convolution operation module includes a Hall bar array composed of multiple parallel Hall bars and an analog addition circuit; the parallel Hall bars respectively receive the current values output by multiple DAC modules as the multipliers for the convolution operation; the anomalous Hall resistance value of the Hall bar is used as the multiplicand; the Hall bar array is used to complete the dot product operation of the input current and the anomalous Hall resistance value, and the product is the anomalous Hall voltage at the output end of the Hall bar. The analog addition circuit performs an addition operation on the multiple anomalous Hall voltages output by the Hall bar array, and the output value is used as the result of the convolution operation of the convolution kernel.

[0010] Preferably, the convolution operation module further includes an amplifier module, which is used to amplify the anomalous Hall voltage at the output end of the Hall bar by the same multiple and then input it into the analog addition circuit.

[0011] Preferably, the convolution operation module further includes a consistency compensation circuit, which is used to perform consistency compensation on the anomalous Hall resistance and output bias of the Hall bar devices in the Hall bar array to improve the convolution accuracy.

[0012] Preferably, the consistency compensation circuit includes a first amplifier OP1, a second amplifier OP2, and a compensating resistor R with adjustable resistance. Two input terminals of the first amplifier OP1 are respectively connected to two output terminals of the Hall bar, the non-inverting input terminal of the second amplifier OP2 is connected to the output terminal of the first amplifier OP1, and the inverting input terminal is connected to one input terminal of the Hall bar. Two ends of the compensating resistor R are respectively connected to the inverting input terminal of the second amplifier OP2 and the ground wire. The other input terminal of the Hall bar is used to receive the input current. The consistency compensation circuit amplifies the anomalous Hall voltage at the output end of the Hall bar by A 1 times through the first amplifier OP1, and then performs a subtraction operation on the amplified result and the voltage drop across the two ends of the compensating resistor R through the second amplifier OP2. After the compensation of the consistency compensation circuit, the output V mul = A 2 ×(A 1 ×I in ×R H + V OS (I in ) - I in ×R), where A 1 , A 2 are the amplification factors of the first amplifier OP1 and the second amplifier OP2 respectively, I in is the input current, R H is the pre-stored anomalous Hall resistance value in the Hall bar, V OS (I in) represents the input offset error. By changing the resistance value of the compensation resistor R, the output offset of the Hall bar can be compensated, and by adjusting the gains of the two amplifiers, the anomalous Hall resistance of the Hall bar can be compensated. Finally, it is realized that under different input currents, the ratio between the output voltage and the input current in each multiplication operation branch is a constant.

[0013] The activation pooling module performs activation pooling on the result output by the analog addition circuit through a hardware integrated circuit with a controllable discharge path, and outputs the maximum voltage value of one-time pooling.

[0014] The hardware integrated circuit with a controllable discharge path is used to implement ReLu activation and maximum pooling, and includes a third amplifier OP3, a diode D1, a fourth amplifier OP4, a capacitor C, and a field effect transistor M1. Among them, the output terminal of the third amplifier OP3 is connected to the positive pole of the diode D1. The negative pole of the diode D1 is connected to the non-inverting input terminal of the fourth amplifier OP4. Both ends of the capacitor C are respectively connected to the non-inverting input terminal of the fourth amplifier OP4 and the source electrode of the field effect transistor M1. The source electrode of the field effect transistor M1 is grounded, the drain electrode is connected to the non-inverting input terminal of the fourth amplifier OP4, and the gate electrode is connected to the control signal Vctrl sent by the FPGA. The inverting input terminal of the fourth amplifier OP4 is connected to the output terminal and the inverting input terminal of the third amplifier OP3. The output of the convolution operation module is connected to the non-inverting input terminal of the third amplifier OP3, and the control signal V ctrl Controls the field effect transistor M1 to conduct, discharges the charge stored in the capacitor C, so that the previous pooling result does not affect the current calculation. During the pooling process, the control signal V ctrl Controls the field effect transistor M1 to cut off, the voltage value during the pooling process is maintained at both ends of the capacitor C, and the output terminal of the fourth amplifier OP4 outputs the maximum voltage value of one-time pooling as the result of activation pooling.

[0015] The ADC module is used to perform analog-to-digital conversion on the maximum voltage value of one-time pooling output by the activation pooling module and send it to the FPGA.

[0016] Compared with the prior art, the present application has the following beneficial effects:

[0017] 1. The multi-channel DAC module directly converts the data sent by the FPGA into the current signal input to the Hall bar array, reducing the time for converting the voltage signal into the current signal and then inputting it into the Hall bar array, and at the same time can also reduce the use of the hardware circuit.

[0018] 2. The activation pooling module is used to reduce the dimension of the results of multiple convolution operations, output the maximum value of one-time pooling, and output it to the ADC circuit for acquisition, greatly reducing the ADC acquisition frequency and improving the speed of the convolutional neural network system. Description of the Drawings

[0019] Figure 1 is the structural block diagram of the convolutional neural network hardware acceleration system based on the Hall bar array in the embodiment;

[0020] Figure 2 is the circuit schematic diagram of the convolutional operation module in the embodiment;

[0021] Figure 3 is the schematic diagram of the consistency compensation circuit;

[0022] Figure 4 is the schematic diagram of the hardware integrated circuit with a controllable discharge path;

[0023] Figure 5 is the flowchart of the convolutional neural network hardware acceleration based on the Hall bar array. Detailed implementation manners

[0024] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further explained below with reference to the accompanying drawings; it must be pointed out that the embodiments described below are only partial embodiments of the present invention, not all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0025] As Figure 1 shown, the convolutional neural network hardware acceleration system based on the Hall bar array

[0026] The convolutional neural network hardware acceleration system based on the Hall bar array includes an FPGA, a multi-channel DAC module, a convolutional operation module, an activation pooling module, and an ADC circuit.

[0027] The FPGA serves as the input / output interface of data and schedules and controls other modules in the system. The FPGA maps the convolutional kernels obtained from the previous training to the anomalous Hall resistance values of the corresponding Hall bars in the Hall bar array, and sends the image data to the multi-channel DAC module.

[0028] As Figure 2 shown, the multi-channel DAC module receives the image data according to the scheduling instructions of the FPGA and converts the sampled data into an output analog current signal. In order not to change the pre-stored anomalous Hall resistance values in the Hall bars, the amplitude of the analog current signal output by the multi-channel DAC module does not exceed the critical current value of the Hall bars.

[0029] The convolution operation module includes a Hall bar array composed of multiple parallel Hall bars, an amplifier module, an analog addition circuit, and a consistency compensation circuit. The parallel Hall bars respectively receive the current values output by the multiple DAC modules as the multipliers for the convolution operation, and the anomalous Hall resistance values of the Hall bars serve as the multiplicands. The Hall bar array is used to perform the dot product operation of the input current and the anomalous Hall resistance value, and the product is the anomalous Hall voltage at the output end of the Hall bar. Since the anomalous Hall voltage is relatively small, the anomalous Hall voltage output by the Hall bar is amplified by the same multiple through the amplifier module and then input into the analog addition circuit. The analog addition circuit performs an addition operation on the multiple anomalous Hall voltages output by the Hall bar array, and the output value serves as the result of the convolution operation of the convolution kernel.

[0030] As Figure 3 shown, the consistency compensation circuit includes a first amplifier OP1, a second amplifier OP2, and a compensating resistor R with adjustable resistance. The two input terminals of the first amplifier OP1 are respectively connected to the two output terminals of the Hall bar. The non-inverting input terminal of the second amplifier OP2 is connected to the output terminal of the first amplifier OP1, and the inverting input terminal is connected to one input terminal of the Hall bar. The two ends of the compensating resistor R are respectively connected to the inverting input terminal of the second amplifier OP2 and the ground wire. The other input terminal of the Hall bar is used to receive the input current. The consistency compensation circuit amplifies the anomalous Hall voltage at the output end of the Hall bar by A 1 times through the first amplifier OP1, and then performs a subtraction operation on the amplified result and the voltage drop across the compensating resistor R through the second amplifier OP2. After compensation by the consistency compensation circuit, the output V mul of the Hall bar array = A 2 × (A 1 × I in × R H + V OS (I in ) - I in × R), where A 1 , A 2 are the amplification factors of the first amplifier OP1 and the second amplifier OP2 respectively, I in is the input current, R H is the pre-stored anomalous Hall resistance value in the Hall bar, and V OS (I in ) represents the input offset error. By changing the resistance value of the compensating resistor R, the output offset of the Hall bar can be compensated, and by adjusting the gains of the two amplifiers, the anomalous Hall resistance of the Hall bar can be compensated. Finally, it is realized that under different input currents, the ratio between the output voltage and the input current in each multiplication operation branch is the same constant.

[0031] The activation pooling module performs activation pooling on the result output by the analog addition circuit through a hardware integrated circuit with a controllable discharge path, and outputs the maximum voltage value of one-time pooling.

[0032] As Figure 4 shown, the hardware integrated circuit with a controllable discharge path is used to implement ReLu activation and maximum pooling, and includes a third amplifier OP3, a diode D1, a fourth amplifier OP4, a capacitor C, and a field effect transistor M1. Among them, the non-inverting input terminal of the third amplifier OP3 is connected to the convolution calculation result output by the convolution operation module, and the output terminal is connected to the positive electrode of the diode D1. The negative electrode of the diode D1 is connected to the non-inverting input terminal of the fourth amplifier OP4. Both ends of the capacitor C are respectively connected to the non-inverting input terminal of the fourth amplifier OP4 and the source electrode of the field effect transistor M1. The source electrode of the field effect transistor M1 is grounded, the drain electrode is connected to the non-inverting input terminal of the fourth amplifier OP4, and the gate electrode is connected to the control signal V ctrl . The inverting input terminal of the fourth amplifier OP4 is connected to the output terminal and the inverting input terminal of the third amplifier OP3.

[0033] The ADC module is used to perform analog-to-digital conversion on the maximum voltage value of one-time pooling output by the activation pooling module and send it to the FPGA.

[0034] s1. As Figure 5 shown, the FPGA maps the convolution kernel to the anomalous Hall resistance value of the corresponding Hall bar in the Hall bar array by using a write pulse and write control;

[0035] s2. After completing the writing of the anomalous Hall resistance value of the Hall bar, the FPGA makes the field effect transistor M1 in the hardware integrated circuit with a controllable discharge path conduct through the control signal V ctrl to discharge the stored charge in the capacitor C and prepare for a new round of pooling calculation;

[0036] s3. After initializing the capacitor C, the FPGA sends the input data corresponding to the next group of convolutions to the multi-channel DAC module, and the multi-channel DAC module converts it into the corresponding Hall bar input current;

[0037] s4. After the Hall bar array in the convolution operation module receives the input current, it performs a dot product operation on the input current and the convolution kernel equivalent to the anomalous Hall resistance, and adds the multi-channel dot product results through an analog addition circuit to obtain the convolution operation result;

[0038] s5. The convolution operation result is input to the activation pooling module. The activation pooling module samples the maximum voltage value of the convolution operation result through feedback and the characteristics of the operational amplifier circuit and holds it in the capacitor C, and the activation and pooling circuit reaches a steady state;

[0039] s6. The FPGA determines whether the number of convolution times reaches the size requirement for pooling calculation. If a pooling is not completed, it returns to s3 to perform convolution calculation on the next set of data;

[0040] If the FPGA determines that a pooling is successfully completed, it acquires the output result of the activation pooling module through the ADC circuit and sends it to the FPGA for storage;

[0041] s7. The FPGA determines whether all the convolution and pooling calculations in this round are completed. If not, it returns to s2 to discharge the previous pooling result stored in the capacitor C to ensure that it will not affect the current pooling calculation; if completed, the convolutional neural network acceleration system completes one round of activation and pooling operations.

Claims

1. A convolutional neural network hardware acceleration system based on a Hall bar array, characterized in that: The convolutional neural network hardware acceleration system based on a Hall bar array includes an FPGA, a multi-channel DAC module, a convolution operation module, an activation pooling module, and an ADC circuit; The FPGA serves as the input / output interface for data and schedules and controls other modules in the system; The multi-channel DAC module receives the input image data according to the scheduling instructions of the FPGA and converts it into an output analog current signal; The convolution operation module includes a Hall bar array composed of multiple parallel Hall bars and an analog addition circuit; the parallel Hall bars respectively receive the current values output by the multi-channel DAC module as the multipliers for convolution operation; the anomalous Hall resistance value of the Hall bar is mapped by the FPGA from the pre-trained convolution kernel as the multiplicand; the anomalous Hall voltage at the output end of the Hall bar array is the result of the dot product operation of the input current and the anomalous Hall resistance value; the analog addition circuit performs an addition operation on the multiple anomalous Hall voltages output by the Hall bar array, and the output value is used as the result of the convolution operation of the convolution kernel; The activation pooling module performs activation pooling on the result output by the analog addition circuit through a hardware integrated circuit with a controllable discharge path, and outputs the maximum voltage value of one-time pooling; The hardware integrated circuit with a controllable discharge path is used to implement ReLu activation and maximum pooling, and includes a third amplifier OP3, a diode D1, a fourth amplifier OP4, a capacitor C, and a field effect transistor M1. Among them, the output terminal of the third amplifier OP3 is connected to the positive electrode of the diode D1; the negative electrode of the diode D1 is connected to the non-inverting input terminal of the fourth amplifier OP4; both ends of the capacitor C are respectively connected to the non-inverting input terminal of the fourth amplifier OP4 and the source electrode of the field effect transistor M1; the source electrode of the field effect transistor M1 is grounded, the drain electrode is connected to the non-inverting input terminal of the fourth amplifier OP4, and the gate electrode is connected to the control signal Vctrl sent by the FPGA; the inverting input terminal of the fourth amplifier OP4 is connected to the output terminal and the inverting input terminal of the third amplifier OP3; the output of the convolution operation module is connected to the non-inverting input terminal of the third amplifier OP3, and the control signal V ctrl controls the on or off of the field effect transistor M1, so that the voltage value across the capacitor C remains unchanged during the pooling process, and the output terminal of the fourth amplifier OP4 outputs the maximum voltage value of one pooling as the result of activation pooling. The ADC module is used to perform analog-to-digital conversion on the maximum voltage value of one-time pooling output by the activation pooling module and send it to the FPGA.

2. The convolutional neural network hardware acceleration system based on a Hall bar array as claimed in claim 1, characterized in that: The amplitude of the analog current signal output by the multi-channel DAC module does not exceed the critical current value that causes the Hall resistance value to change.

3. The convolutional neural network hardware acceleration system based on a Hall bar array as claimed in claim 1, characterized in that: The convolution operation module further includes an amplifier module for amplifying the anomalous Hall voltage at the output end of the Hall bar by the same multiple and then inputting it into the analog addition circuit.

4. The convolutional neural network hardware acceleration system based on a Hall bar array as claimed in claim 1, characterized in that: The convolution operation module further includes a consistency compensation circuit for compensating the consistency of the anomalous Hall resistance and output bias of the Hall bar devices in the Hall bar array to improve the convolution accuracy.

5. The convolutional neural network hardware acceleration system based on a Hall bar array as claimed in claim 4, characterized in that: The consistency compensation circuit includes a first amplifier OP1, a second amplifier OP2, and a compensating resistor R with adjustable resistance value. Two input terminals of the first amplifier OP1 are respectively connected to two output terminals of the Hall bar. The non-inverting input terminal of the second amplifier OP2 is connected to the output terminal of the first amplifier OP1, and the inverting input terminal is connected to an input terminal of the Hall bar. Two ends of the compensating resistor R are respectively connected to the inverting input terminal of the second amplifier OP2 and the ground wire. The other input terminal of the Hall bar is used to receive an input current. The consistency compensation circuit amplifies the anomalous Hall voltage at the output terminal of the Hall bar by the first amplifier OP1 by A 1 times, and then performs a subtraction operation on the amplified result and the voltage drop across the compensating resistor R by the second amplifier OP2. After compensation by the consistency compensation circuit, the output V mul of the Hall bar array is equal to A 2 ×(A 1 ×I in ×R H +V OS (l in )-I in ×R), where A 1 and A 2 are the amplification factors of the first amplifier OP1 and the second amplifier OP2 respectively, I in is the input current, R H is the pre-stored anomalous Hall resistance value in the Hall bar, and V OS (I in ) represents the input offset error. By changing the resistance value of the compensating resistor R to compensate for the output bias of the Hall bar and adjusting the gains of the two amplifiers to compensate for the anomalous Hall resistance of the Hall bar, finally, the ratio between the output voltage and the input current in each multiplication operation branch is a constant under different input currents.

6. The convolutional neural network hardware acceleration system based on a Hall bar array as claimed in any one of claims 1 to 5, characterized in that: The process of using this system for convolutional neural network acceleration includes the following steps: s1. The FPGA maps the convolution kernel to the anomalous Hall resistance values of the corresponding Hall bars in the Hall bar array by using a write pulse and write control; After the anomalous Hall resistance value of the Hall bar is written, the FPGA controls the signal V ctrl to turn on the field effect transistor M1 in the hardware integrated circuit with a controllable discharge path, complete the discharge of the stored charge in the capacitor C, and prepare for the new round of pooling calculation; s3. The FPGA sends the input data corresponding to the next group of convolutions to the multi-channel DAC module, and the multi-channel DAC module converts it into the corresponding Hall bar input current; After the Hall bar array in the convolution operation module receives the input current, it performs a dot product operation on the input current and the convolution kernel equivalent to the anomalous Hall resistance, and adds the multiplexed dot product results through an analog adder circuit to obtain the convolution operation result; s5. The convolution operation result is input to the activation pooling module. The activation pooling module samples the maximum voltage value of the convolution operation result through the characteristics of the feedback and operational amplifier circuit and holds it in the capacitor C, and the activation and pooling circuit reaches a steady state; s6. The FPGA determines whether the number of convolutions reaches the size requirement for pooling calculation. If a pooling is not completed, it returns to s3 to perform the convolution calculation of the next set of data; If the FPGA determines that a pooling is successfully completed, it acquires the output result of the activation pooling module through the ADC circuit and sends it to the FPGA for storage; s7. The FPGA determines whether all the convolution pooling calculations in this round are completed. If not, it returns to s2 to discharge the previous pooling result stored in the capacitor C to ensure that it will not affect the current pooling calculation; if completed, the convolutional neural network acceleration system completes one round of activation and pooling operations.

Citation Information

Patent Citations

  • Memory-based convolutional neural network system

    CN108805270A

  • Unsupervised learning synaptic unit circuit based on Hall strip

    CN112270409A