Object Detection Method for a Binary Neural Network Model and Its Hardware Acceleration Method

By building a binary neural network model with time information on the pulsed neural network, the problem of high-precision object detection on low-power devices is solved, and the detection accuracy is improved, reducing hardware storage resource consumption.

CN116434035BActive Publication Date: 2025-06-24PEKING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310343231.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-03
Publication Date
2025-06-24
Estimated Expiration
2043-04-03

AI Technical Summary

Technical Problem

The prior art is difficult to realize high-precision object detection on low-power devices, and the accuracy of binary neural networks in object detection is insufficient to replace pulsed neural networks.

Method used

A binary neural network model with time information is constructed based on the pulse neural network. By transforming the neurons of the pulse neural network, using step functions and using single-frame convolution results for threshold judgment, adding a time information fusion layer, and designing a hardware acceleration structure.

Benefits of technology

High-precision object detection on low-power devices is realized, and the target detection accuracy is improved compared with traditional binary neural networks and reduces hardware storage resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116434035B_ABST
    Figure CN116434035B_ABST
Patent Text Reader

Abstract

The present invention discloses an object detection method for a binary neural network model and its hardware acceleration method. A binary neural network detection model with time information is constructed based on a spiking neural network. The activation function of the spiking neural network is a step function, and the threshold judgment is performed using the convolution result of a single frame. Then, a time information fusion layer is added to the binary neural network. The model outputs the object detection box information. A hardware acceleration module is designed for the time information fusion layer of the constructed model, including a loading module, an accumulation module, a calculation module, and an output module. The accumulation process of neurons is completed through the loading module and the accumulation module, and the normalization and convolution processes are completed through the calculation module. At the same time, the convolution operation is performed first and then the normalization process. By using the method of the present invention, the object detection speed can be improved, and the storage resource consumption can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and relates to object detection and acceleration technologies, and particularly relates to an object detection method for a binary neural network model and a hardware acceleration method therefor. Background Art

[0002] In recent years, with the rapid development of autonomous driving and wearable devices, neural network models (CNNs) have been increasingly widely used in the field of object detection, among which convolutional neural network models are the most widely used. However, applying convolutional neural networks to object detection is difficult for low-power devices due to their huge computational amount.

[0003] In order to reduce the power consumption of neural network model computations, spiking neural network models are currently being used in the field of object detection technology. Due to the extremely low power consumption brought about by the sparse characteristics of spiking neural networks themselves, they are known as the third-generation neural network models. However, due to the characteristic that spiking neurons need to continuously maintain their own membrane voltages, it brings extremely high storage resource consumption in hardware implementation and extremely high bandwidth pressure to hardware acceleration, and the prior art is difficult to solve the problem of resource consumption caused by reducing the membrane voltage of spiking neural networks.

[0004] On the other hand, existing research has proposed a binary neural network (BNN network) for compressing CNN networks, that is, using single-bit weights or features to transfer parameters or information. Existing traditional binary neural networks consume less computing resources. However, compared with spiking neural networks, binary neural networks do not carry time information, so the accuracy of object detection is far less than that of spiking neural networks, and binary neural networks can only be used in relatively simple image classification scenarios and are difficult to be applied to the field of object detection technology.

[0005] In summary, on the one hand, spiking neural networks have the deficiency of extremely high storage resource waste or extremely high bandwidth pressure in hardware acceleration due to the characteristic that they always need to maintain membrane voltages. On the other hand, the existing traditional binary neural network BNN network model is relatively small and consumes less resources, but because it does not carry time information, it is difficult to be used for object detection. Summary of the Invention

[0006] To overcome the deficiencies of the above-mentioned existing technologies, the present invention provides a target detection method for a binary neural network model and its hardware acceleration method. Based on the spiking neural network, a binary neural network model with time information is proposed for target detection, enabling the constructed binary neural network target detection model to carry time information while having the low power consumption of the spiking neural network, and improving the target detection accuracy compared to most existing binary neural networks. Further, the present invention proposes a method for hardware acceleration on an FPGA (Field Programmable Gate Array), verifying the hardware implementation effect of the constructed binary neural network target detection model.

[0007] The technical solution provided by the present invention is as follows:

[0008] A target detection method for a binary neural network model, which constructs a binary neural network detection model with time information based on a spiking neural network. The activation function of the spiking neural network is a step function, and the threshold judgment is performed using the convolution result of a single frame; a time information fusion layer is added to the binary neural network; a hardware acceleration structure for the time information fusion layer is also designed; the method includes the following steps:

[0009] 1) Obtain multiple frames of image data for target detection. Images with targets such as people and vehicles taken by a camera can be used, and the images are binarized by frequency encoding to obtain multiple frames of binarized images, which are used as the images to be detected input into the model.

[0010] 2) Construct a binary neural network target detection model with binary time information based on a spiking neural network, including a feature extraction layer, a time information fusion layer, and a detection layer; specifically including:

[0011] 21) Modify the neurons of the spiking neural network;

[0012] Specifically, the neurons of the spiking neural network model (YOLO (You Only Look Once) neurons are used in specific implementation) are modified. The neurons of the spiking neural network (feature extraction layer) need to accumulate the neuron membrane voltages of the previous frames before judging whether to activate. After modification, only the recognition result of the current frame image is used for threshold comparison to judge whether to activate (compare the convolution calculation result of this frame with the reference activation threshold, and if it is greater than the activation threshold, emit (output) 1 (1 represents activating the neurons of this layer), and if it is less than the threshold, emit 0 (not activating the neurons of this layer)). At the same time, the time information of the previous frames is accumulated in the last layer of the model, and then normalized and convolved. Thus, the constructed network model not only retains part of the time information but also has the characteristic of single-bit signal transmission; the neurons of the spiking neural network use a step function; the modified neurons of the spiking neural network are represented as:

[0013]

[0014] Among them, y is the output image of the neurons in the spiking neural network, x is the input image to be detected; w is the model weight; x*w is the intermediate result of the convolution of the current layer of the spiking neural network, and TH is the activation threshold. After the transformation of the above neurons, there is no need to store the previous image information during the hardware implementation process, which greatly saves the use of on-chip storage resources.

[0015] 22) Add a time information fusion layer; the time information fusion layer includes a time information normalization module and a convolution module.

[0016] During specific implementation, place the time information fusion layer in front of the yolo layer of the YOLO neural network, that is, in front of the detection layer (i.e., the 16th and 24th layers in Table 1); the convolutional layers before the 16th and 24th layers are used for feature extraction of each frame of the image;

[0017] After feature extraction, perform feature fusion in the time information fusion layer; compared with normal convolution, the time information fusion layer has additional operations of accumulating and normalizing N frames of information, and the activation function of the time information fusion layer uses the RELU function; normalization means dividing the accumulated information (the number of 1s X in the previous time step (step) frames) by the time step step; the above normalization is normalization in the time dimension.

[0018] 3) Train the constructed binary neural network object detection model, and obtain the accurate information of the object detection frame through the detection layer of the object detection model;

[0019] During specific implementation, the training dataset uses the MS COCO image dataset. The detection frame is specifically represented by four coordinates x, y, w, and h of a rectangle. (x, y) is the center point coordinate of the rectangle, and (w, h) are the length and width of the rectangle.

[0020] During specific implementation, after the above feature fusion, the output channel depth is 85. Among them, 4 channels are the center point coordinates of the detection frame and the length and height of the detection frame, 1 channel is the confidence of the border, and the remaining 80 channels are the class confidences of the targets. Through the yolo layer, statistical processing is performed on the information of 85 channels to obtain the accurate information of the object detection frame, including the center point coordinates of the detection frame and the length and height of the detection frame.

[0021] During specific implementation, input the image to be detected into the trained binary neural network object detection model, and output the object detection frame information, thereby realizing the detection of the object.

[0022] Furthermore, the present invention also designs a hardware acceleration module and method for the time information fusion layer of the constructed model;

[0023] The time information fusion layer is mainly divided into a time information normalization module and a convolution module. The time normalization module includes the process of accumulating the activation of neurons in previous frames and the normalization process, and convolution is performed after normalization. When constructing the corresponding hardware acceleration module for the time information fusion layer, it includes a loading module, an accumulation module, a calculation module, and an output module. Among them, the accumulation process of neurons is completed through the loading module and the accumulation module, and the normalization and convolution processes are completed through the calculation module. At the same time, in order to make it easier for the hardware to implement, the normalization process is performed after the convolution operation.

[0024] 41) Loading module of the time information fusion layer: It includes states 1 to 4, so that the storage required for loading and calculation does not conflict with each other.

[0025] It includes a 4-line buffer. Taking the 0th, 1st, 2nd, and 3rd rows of the feature map as an example, three rows of the buffer store the 0th, 1st, and 2nd rows for (convolution) calculation, and one row is used for the loading of the 3rd row.

[0026] 42) The accumulation module of the time information fusion layer constructs a time window through BRAM (block RAM), and the results of frame operations are accumulated within the time window. Continuously accumulate the values accumulated within the window with the new input data of the next frame image, and at the same time discard the old data in the time window.

[0027] 43) The calculation module uses the method of performing convolution operation first and then normalization.

[0028] Specifically, the calculations of the two multiply-accumulate units in the calculation module are expressed as:

[0029]

[0030] Among them, the left side of the formula represents the operation process of the algorithm. Among them, X1 and X2 are the values accumulated by different multiply-accumulate units respectively, step is the time step, W1 and W2 are the weights of the convolution kernels respectively. After normalization, the accumulated value is multiplied by the weight at the corresponding position and then accumulated, which is the convolution operation. After optimization, the convolution operation is performed first and then the normalization according to the right side of the formula.

[0031] The computing module includes multiple multiply-accumulate units. For an image with three dimensions of length, width, and channels as input, the computing module first performs a convolution operation on the weights and features through N multiply-accumulate units. With each additional multiply-accumulate unit, the computing speed will increase. After the multiply-accumulation is completed, the results of the convolution for each channel (two-dimensional image) are accumulated and summed. Then, it is judged whether the convolution of the input channels is completed. If not, the above operations are repeated. If all the input channels are accumulated, the results are normalized, that is, divided. In an FPGA, it can be implemented by shifting. The normalized results are added with biases to complete the calculation of the time information fusion layer. The results of the convolution calculation are compared with 0 through a comparator. If it is greater than 0, the calculated results are output. If it is less than 0, 0 is output, which is the hardware implementation of the RELU activation function of the time information fusion layer.

[0032] Through the above steps, the object detection and hardware acceleration of the binary neural network model can be achieved.

[0033] Compared with the prior art, the beneficial effects of the present invention are:

[0034] The present invention provides an object detection method and a hardware acceleration method for a binary neural network model. A binary neural network model with time information is constructed based on a spiking neural network and used for object detection. Among them, the neurons of the spiking neural network adopt a step function, and the convolution results of the current frame are used for threshold judgment to convert the spiking neural network into a binary neural network; then a time information fusion layer is added to the binary neural network to construct a BNN with time information. In addition, a hardware acceleration structure for the time information fusion layer is designed to improve the computing speed of object detection. Using the BNN neural network detection method with binary time information provided by the present invention, after training, it can reach MAP (mean average precision, the average correct rate of all classes) of 0.2 on the COCO dataset (Common Objects in Context, a dataset that can be used for image recognition). At the same time, its corresponding FPGA hardware acceleration scheme is realized, the detection speed is improved, and compared with the original spiking neural network for hardware acceleration, the storage resource consumption is reduced by 22.5 MB when using the BNN network with time information for hardware acceleration. Description of the Drawings

[0035] Figure 1 It is a structural block diagram of the binary neural network detection model with time information provided by the present invention.

[0036] Figure 2 It is an operation flow block diagram of the time information fusion layer for constructing the model of the present invention.

[0037] Figure 3 It is a structural block diagram of the hardware implementation of the time information fusion layer.

[0038] Figure 4 Schematic diagram of four states of the loading module for hardware acceleration of the time information fusion layer. Detailed implementation manners

[0039] The present invention will be further described below in conjunction with the accompanying drawings through embodiments, but the scope of the present invention is not limited in any way.

[0040] The present invention provides a target detection method for a new binary neural network model and its hardware acceleration method. For a grayscale image containing a target, through the inference process of the neural network, a detection box of the target can be obtained, and the detection box is described by the coordinates of the center point of the box and the length and width of the box. By improving the spiking neural network, an algorithm for a BNN neural network detection model with binary time information is proposed. After training, it can reach MAP0.2 on the COCO image dataset, and at the same time, its corresponding FPGA hardware acceleration scheme is realized. Compared with the original spiking neural network, the storage resource consumption is reduced by 22.5MB.

[0041] Existing research has proposed a BNN network by compressing the CNN network, that is, using single-bit weights or features to transfer parameters or information. The present invention is based on the existing spiking neural network and proposes a BNN neural network detection model with binary time information. Table 1 details the parameter information of each layer of the model. There are two great differences between the spiking neural network and the convolutional neural network. One is the carrying of time information. Therefore, the spiking neural network stores the time information of the recognition of the previous frame of the image, and the target detection is carried out through the joint action of the current frame of the image and the previous frame of the image. The second difference is similar to the binary neural network, and the information transfer between the networks uses single-bit pulse signals. However, in the existing hardware implementation process of the spiking neural network, because the spiking neural network stores the information of the previous frame in each layer, the information to be stored is too large, and the on-chip storage resources of the FPGA are not much, resulting in a great waste of resources. The present invention proposes to transform the neurons of the spiking neural network, remove the accumulation process of the neuron model, and only use the recognition result of a single-frame image for threshold comparison and then firing. At the same time, the time information of the previous frame is accumulated in the last layer of the model, and then normalized and then conventional convolution is performed. Thus, the network model constructed by the present invention can not only retain part of the time information, but also has the characteristic of single-bit signal transmission, forming a unique BNN network model.

[0042] Table 1 Details of each layer of the model

[0043]

[0044]

[0045] The improvements made by the BNN neural network detection method with binary time information provided by the present invention at the algorithm level mainly include: First, the neurons of the spiking neural network adopt a step function, directly comparing the result of single-frame convolution calculation with a reference threshold. If it is greater than the threshold, 1 is emitted; if it is less than the threshold, 0 is emitted. The following formula lists the transformed neurons. Where y is the output of the neurons of the spiking neural network, x*w is the result of the current layer convolution of the spiking neural network, and TH is the threshold. After the above neuron transformation, there is no need to store the previous image information during the hardware implementation process, greatly saving the use of on-chip storage resources. After calculation, 22.5 Mbit of storage resource occupancy can be saved.

[0046]

[0047] After the above replacement, the time characteristics of the spiking neural network are removed, which greatly affects the detection accuracy. The present invention adds a time information fusion layer to make up for it. Refer to Figure 1 , the neural network after removing the time characteristics is a normal YOLO neural network. After changing the RELU activation function of the YOLO neural network to a step function, a BNN network is formed. The time information fusion layer added by the present invention is placed in the front of the yolo layer of the YOLO neural network, that is, the 16th and 24th layers in Table 1. It can be considered that all the layers before the time information fusion layer are the single-frame image feature extraction process. After the feature extraction is completed, feature fusion is performed in the time information fusion layer. Regarding the time information fusion layer, refer to Figure 2 , compared with normal convolution, the time information fusion layer has additional operations of N-frame information accumulation and normalization. At the same time, the activation function of the time information fusion layer uses the RELU function of a normal neural network. We take the time information fusion layer of the 16th layer as an example. Because the signals transmitted from the 15th layer are single-bit 1s and 0s, the cumulative information is actually to judge the number X of 1s in the previous N frames. The normalization is to divide X by the time step N. It should be noted that the above normalization is different from the batch normalization of a normal neural network. The batch normalization is the normalization of channels in space, and the above normalization is the normalization in the time dimension.

[0048] Figure 3 Lists the hardware acceleration scheme of the time information fusion layer of the present invention. The time information fusion layer is mainly divided into four modules: loading, accumulation, calculation, and output. Among them, the loading module refers to Figure 4 , including states 1 to 4. A 3*3 convolution process requires 3 rows of output features for calculation. Therefore, the loading module is mainly divided into 4 rows of buffers. We take the 0th, 1st, 2nd, 3rd, 4th, 5th, and 6th rows of features as an example. Among them, the gray buffer is used for the loading of the next row, and the remaining three rows are used for the current calculation. Through four states, the clock of the loading module can be hidden in the calculation module. Refer to Figure 4, when in state 1, the data of the 0th, 1st, and 2nd rows are stored in the second, third, and fourth buffers for calculation, and the first buffer is used to load the data of the 3rd row. In state 2, since the data of the 0th row stored in the second buffer has been calculated and this row of data is no longer needed for subsequent calculations, the second buffer can then be used to store the data of the 4th row loaded. At this time, the first buffer has already stored the data of the 3rd row. Therefore, the data of the 1st, 2nd, and 3rd rows stored in the first, third, and fourth buffers can be used for calculation. Similarly, in the third state, the data of the 2nd, 3rd, and 4th rows are used for calculation and the data of the 5th row is loaded, and in the fourth state, the data of the 3rd, 4th, and 5th rows are used for calculation and the data of the 6th row is loaded. With these four states, the parallelism of loading and calculation can be achieved. A total of four states can ensure that the storage required for loading and calculation does not conflict with each other.

[0049] The accumulation module of the time information fusion layer constructs a time window through BRAM, continuously adds the accumulated values within the window to the data of the next frame of the newly input data, and at the same time discards the old window data. The calculation module is the core. In the present invention, the data after accumulation is first normalized and then convolution operation is performed. Normalization is a division operation, and division operation will inevitably bring floating-point numbers. Using floating-point numbers for convolution consumes much more hardware resources than using ordinary data for convolution. Therefore, the present invention adopts the method of performing convolution operation first and then normalization in the calculation module. The following formula illustrates the rationality of this method. Convolution is essentially a multiply-accumulate operation. The left side of the following formula represents the operation process of the algorithm, where X1 and X2 are the accumulated values, STEP is the time step, and W1 and W2 are the weights of the convolution kernel. After the normalized accumulated value and the weights are multiplied at the corresponding positions and then accumulated, it is the convolution operation. We extract the common factor N and perform convolution first and then normalization. It can be seen from the following formula that the operation results are the same.

[0050]

[0051] The hardware acceleration method of the time information fusion layer includes that the calculation module first performs convolution operation on the weights and features through N multiply-accumulate units. For each additional multiply-accumulate unit, the calculation speed will be improved. After the multiply-accumulation is completed, the results of convolution of each channel are accumulated and summed. Then it is judged whether the convolution of the input channels is completed. If not, the above operations are repeated. If all the input channels are accumulated, the results are normalized, that is, division, which can be realized by shifting in FPGA. After adding the bias to the normalized result, this convolution calculation is completed. The result of the convolution calculation is compared with 0 through a comparator. If it is greater than 0, the calculation result is output. If it is less than 0, 0 is output, which is the hardware implementation of the RELU activation function of the time information fusion layer.

[0052] It should be noted that the purpose of disclosing the embodiments is to help further understand the present invention. However, those skilled in the art can understand that various substitutions and modifications are possible without departing from the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the embodiments, and the scope of protection claimed by the present invention shall be subject to the scope defined by the claims.

Claims

1. A target detection method for a binary neural network model, characterized in that, Construct a binary neural network detection model with time information based on a spiking neural network. The activation function of the spiking neural network is a step function, and the single-frame convolution result is used for threshold judgment. Then, a time information fusion layer is added to the binary neural network. The model outputs the target detection box information, including the following steps: 1) Obtain multiple frames of image data for target detection, perform binarization on the images to obtain multiple frames of binarized images, which are the images to be detected input into the model; 2) Construct a binary neural network target detection model with binary time information based on a spiking neural network, including a feature extraction layer, a time information fusion layer, and a detection layer. The feature extraction layer is the convolutional layer before the front of the detection layer, used for feature extraction of each frame of image. The time information fusion layer is placed in front of the detection layer of the spiking neural network for feature fusion; 21) Transform the neurons of the spiking neural network; The neurons of the spiking neural network model are transformed to only use the recognition result of the current frame of image for threshold comparison and then judge whether to activate. At the same time, the time information of the previous frames is accumulated in the last layer of the model, and then normalized and convolved. The constructed binary neural network target detection model with binary time information based on the spiking neural network retains part of the time information and has the characteristic of single-bit signal transmission. The neurons of the binary neural network target detection model use the step function. The neurons of the transformed network model are expressed as: where y is the output image of the neurons of the network model, x is the input image to be detected, w is the model weight, x*w is the intermediate result of the convolution of the current layer of the spiking neural network, and TH is the activation threshold. After the neurons are transformed, the previous image information is not stored during the hardware implementation process, thus saving the use of on-chip storage resources; 22) Add a time information fusion layer. The time information fusion layer includes a time information normalization module and a convolution module; The time information fusion layer has N more frames of information accumulation and normalization operations than convolution. At the same time, the activation function of the time information fusion layer uses the RELU function; 3) For the constructed binary neural network target detection model, use the image dataset for training, and obtain the accurate information of the target detection box through the detection layer of the trained target detection model. Thus, the detection of the target is realized.

2. The object detection method of the binary neural network model according to claim 1, characterized in that In step 1), to obtain multiple frames of image data for target detection, specifically, images with detection targets are captured by a camera, and the images are binarized by frequency encoding to obtain multiple frames of binarized images.

3. The object detection method of the binary neural network model according to claim 1, characterized in that, The spiking neural network model specifically uses the YOLO neural network model.

4. The object detection method of the binary neural network model according to claim 1, characterized in that, In step 22), the number of frames with a time step of 1 is the accumulated information; Normalization is normalization in the time dimension, that is, the accumulated information is divided by the time step.

5. The object detection method of the binary neural network model according to claim 1, characterized in that, In step 3), the training dataset specifically uses the MS COCO image dataset.

6. The object detection method of the binary neural network model according to claim 5, characterized in that, The detection box is specifically represented by four coordinates x, y, w, h for a rectangular box, where (x, y) is the center point coordinates of the rectangular box, and (w, h) are the length and width of the rectangular box.

7. The object detection method of the binary neural network model according to claim 1, characterized in that, Design a hardware acceleration module for the time information fusion layer of the constructed model, including: a loading module, an accumulation module, a calculation module, and an output module. The accumulation process of neurons is completed through the loading module and the accumulation module, and the normalization and convolution processes are completed through the calculation module. At the same time, the convolution operation is performed first and then the normalization process is carried out.

8. The object detection method of the binary neural network model according to claim 7, characterized in that, The loading module of the time information fusion layer includes four states, so that the storage required for loading and calculation does not conflict with each other. The accumulation module of the time information fusion layer constructs a time window through a block memory. The results of frame operations are accumulated within the time window, continuously adding the accumulated values within the window to the next frame of image data newly input, and at the same time discarding the old data in the time window. The calculation module includes multiple multiply-accumulate units. The convolution operation is performed first and then the normalization is carried out. It is implemented by shifting in the FPGA. The result after normalization plus the bias completes the calculation of the time information fusion layer. The result of the convolution calculation is compared through a comparator. If it is greater than 0, the calculation result is output. If it is less than 0, 0 is output, that is, the hardware implementation of the RELU activation function of the time information fusion layer is completed, and the hardware acceleration of the object detection of the binary neural network model is realized.

9. The object detection method of the binary neural network model according to claim 8, characterized in that, The loading module of the time information fusion layer includes four-line buffering, where three lines of buffering are used for storage and convolution calculation, and the remaining one line is used for the loading of the third line.

10. The object detection method of the binary neural network model according to claim 8, characterized in that, The calculations of two multiply-accumulate units in the calculation module are expressed as: Among them, the left side of the formula represents the operation process of the algorithm. Among them, X1 and X2 are the values accumulated by different multiply-accumulate units respectively, step is the time step, W1 and W2 are the weights of the convolution kernels respectively. After the normalized cumulative value is multiplied by the weights at the corresponding positions and then accumulated, it is the convolution operation. After optimization, the convolution operation is performed first and then the normalization according to the right side of the formula.