Sparse binary network activation image encoding and decoding hardware accelerator for gesture recognition

By designing a sparse binary network activation image encoding and decoding hardware accelerator, compressing and decoding sparse activation images, the high power consumption problem of convolutional neural network gesture recognition in hardware implementation is solved, and low power consumption and high energy efficiency gesture recognition is achieved.

CN116957023BActive Publication Date: 2025-08-08FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310837180.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-08
Publication Date
2025-08-08
Estimated Expiration
2043-07-08

AI Technical Summary

Technical Problem

The existing convolutional neural network gesture recognition technology consumes high power and has a long delay in hardware implementation, so it cannot be deployed on mobile terminals and IoT devices. There is redundant calculation in the sparse activated images of binary networks, resulting in excessive power consumption and calculation periods.

Method used

The activated image encoding and decoding hardware accelerator for sparse binary networks is designed. By activating the image encoding unit to compress the same vector and store the effective vector, the encoded image decoding unit skips redundant calculations and adopts structures such as sparse vector filters, label generators, effective vector counters and coded image storage units to realize the compression and efficient decoding of sparse activated images.

Benefits of technology

It effectively reduces chip power consumption, reduces data storage and calculation times, improves computing speed and hardware performance, and achieves low-power and high-energy-efficient gesture recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116957023B_ABST
    Figure CN116957023B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of integrated circuit technology, and specifically is a sparse binary network activation image encoding and decoding hardware accelerator for gesture recognition. The hardware accelerator of the present invention includes: an activation image encoding unit, a coded image storage unit, and a coded image decoding unit. The accelerator compresses a large number of identical vectors in the activation image of the sparse binary network. These identical vectors are generated by the binary gesture edges as input. Compressing the activation image can effectively reduce on-chip storage capacity, reduce data handling power consumption, and thus reduce chip power consumption; the accelerator can decode the compressed activation map, skip repeated calculations during the decoding process, thereby reducing the number of calculations, increasing the calculation speed, and effectively improving hardware performance. The present invention has the characteristics of low power consumption and low latency, effectively reducing hardware resources and improving the energy efficiency of hardware.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of integrated circuits, and in particular relates to a sparse binary network activated image encoding and decoding hardware accelerator for gesture recognition. Background Art

[0002] Natural, barrier-free, highly effective, and contactless new intelligent human-computer interaction systems are becoming an inevitable trend in information development. Gestures, as one of the most important channels in human-computer interaction, have the advantages of wide application, simple operation, and high frequency of use. However, currently, gesture recognition mostly uses convolutional neural networks. However, convolutional neural network models are generally parameter-intensive and computationally intensive, resulting in high power consumption and high latency in hardware implementation. Convolutional neural network-based gesture recognition has high resource requirements and correspondingly high power consumption, making it unsuitable for deployment in mobile devices, IoT, and wearable devices. Binary networks, similar to convolutional neural networks, use 1-bit inputs and weights. Using XNOR-PopCount instead of the traditional multiplication-accumulation operations of convolutional neural networks reduces data usage and computational complexity, making them ideal for low-power gesture recognition hardware. Given that the first layer of a binary network typically uses multiplication-accumulation due to RGB input, using sparse gesture edges ensures that this first layer also uses XNOR-PopCount. Furthermore, sparse gesture edges lead to a large number of identical vectors in the binary network's activation image. These vectors are composed of 1-bit activation data from all channels and are present in large quantities in every layer of the binary network. The calculation results of these vectors can be predicted by software, resulting in a large amount of redundant computation, unnecessary power consumption, and wasted computing cycles. Therefore, designing encoding and decoding hardware for the sparse activation images in binary networks to achieve ultra-efficient gesture recognition hardware is urgently needed. Summary of the Invention

[0003] In order to overcome the shortcomings of the prior art, the purpose of the present invention is to propose a sparse binary network activated image encoding and decoding hardware accelerator for gesture recognition with low resource consumption, low power consumption and high energy efficiency.

[0004] The present invention provides a sparse binary network activation image encoding and decoding hardware accelerator for gesture recognition, which includes an activation image encoding unit, an encoded image storage unit, and an encoded image decoding unit. The activation image encoding unit receives the activation image vector of the sparse binary network and its valid signal as input. After encoding, a large number of identical vectors present in the activation image of the sparse binary network are compressed. These identical vectors are generated by the binary gesture edges as input. Compressing the activation image can effectively reduce on-chip storage capacity, reduce data handling power consumption, and thus reduce chip power consumption. The encoded valid vectors are stored in the encoded image storage unit. The encoded image decoding unit reads the encoded image information from the encoded image storage unit and decodes it. During the decoding process, repeated calculations are skipped, thereby reducing the number of calculations, increasing calculation speed, and effectively improving hardware performance.

[0005] The activation image encoding unit includes a sparse vector filter, a label generator, a valid vector counter, and a row counter; when the activation image vector and the activation image vector valid signal are sent to the sparse vector filter and the row counter respectively, the sparse vector filter compares the input activation image vector with the predicted sparse vector provided by the software; if they are equal, the input activation image vector is a sparse vector, and a sparse judgment signal of 0 value is sent to the label generator and the valid vector counter; if they are not equal, the input activation image vector is a valid vector, and the valid vector is sent to the encoding image storage unit, and a sparse judgment signal of 1 value is sent to the label generator and the valid vector counter; wherein the predicted sparse vector provided by the software refers to the sparse vector of each layer directly calculated by the software during the forward reasoning of the software. During the forward reasoning of the software, the weight data remains unchanged. Since there are a large number of sparse 0-value inputs in the input image, it can be considered that there is a part of unchanged data at this time, so the sparse vector of each layer can be predicted by the software.

[0006] The row counter performs statistics based on the activation image valid signal to generate a column count value. When the column count value is less than the image length, it indicates that the current input activation image vector has not reached the number of one row, and the next row signal next_row is kept at 0 and sent to the label generator and the valid vector counter. When the column count value is greater than or equal to the image length, it indicates that the current input activation image vector has reached the number of one row, and the next row signal next_row is pulled up to 1 and sent to the label generator and the valid vector counter, and the column count value is reset to 1.

[0007] The label generator receives the sparse determination signal of the sparse vector filter and the next row signal next_row of the row counter, first determines whether next_row is 1, and if so, sends the label value of the label generator to the coded image storage unit and resets the label value to 0; if not, sets the lowest position of the label value to the value of the sparse determination signal, and shifts the label value one position to the left;

[0008] The valid vector counter receives the sparse determination signal of the sparse vector filter and the next row signal next_row of the row counter, first determines whether next_row is 1, and if so, sends the number of valid vectors in the valid vector counter to the coded image storage unit; if not, adds the number of valid vectors to the value of the sparse determination signal to generate a new number of valid vectors;

[0009] The coded image storage unit includes a valid vector register stack, an output tag register stack, and a first valid vector address register stack; the valid vector register stack receives valid vectors sent by the active image coding unit and stores them in sequence; the output tag register stack receives tag values sent by the active image coding unit and stores them in sequence; the first valid vector address register stack receives the number of valid vectors sent by the active image coding unit and stores them in sequence;

[0010] The coded image decoding unit includes an address generator, a label decoder, a pointer module, a valid vector address module, a 3x3 label comparator, a prediction sparse vector register and an activation image recovery module;

[0011] During decoding, the label decoder reads three rows of label data from the output label register stack, slides from left to right on the three rows of label values in a 3x3 rectangular frame according to the convolution sliding process, takes out one 3x3 label value at a time and sends it to the pointer module, the address generator module, the 3x3 label comparator, and the activation image recovery module until it slides to the rightmost side, indicating that the convolution calculation corresponding to the current label is completed. Then, a new row of label values is read from the output label register stack, the first row of the original three rows of label values is discarded, and three new rows of decoded data are formed until all the output label register stack data are decoded;

[0012] The 3x3 tag comparator receives the 3x3 tag values and makes a judgment; if all 3x3 tag values are equal to 0, it indicates that the subsequent calculation results can be predicted by the software, and the subsequent calculation is skipped; otherwise, the subsequent decoding process is performed;

[0013] The effective vector address module reads the address of the row where the corresponding tag value is located from the first effective vector address register file;

[0014] The predicted sparse vector register stores the sparse activation image vector predicted by the software;

[0015] The pointer module stores the number of valid values from the leftmost position of the 3x3 label value to the leftmost position of the original label value, and uses P1, P2, and P3 to represent the number of valid values of three rows respectively. When the 3x3 label value is received, P1, P2, and P3 are sent to the address generator and P1, P2, and P3 are added to the leftmost value of the 3x3 label value and P1, P2, and P3 are updated;

[0016] The address generator generates an address based on the 3x3 label value, P1, P2, P3 of the pointer module, and the number of valid vectors in the valid vector address module; the address generator adds the number of valid vectors in the corresponding row, the pointer, and the corresponding label to generate the valid vector address corresponding to the label; the address generator determines whether each label value in the 3x3 label value is 0. If it is 0, the corresponding data is a sparse activation image vector and is sent to the activation image recovery module; if it is 1, the valid vector is read from the valid vector register file according to the generated address and sent to the activation image recovery module;

[0017] The activation image recovery module obtains 9 vectors, which are arranged in sequence to form the original 3x3 activation image block and sent to the binary network calculation module for convolution calculation.

[0018] Compared with the prior art, the beneficial technical effects of the present invention are embodied in:

[0019] (1) The present invention introduces new sparsity into the activation image of the binary network. By using sparse gesture edges as the input of the binary network, a large number of identical vectors can be introduced into the activation image of each layer of the binary network. The calculation results of these vectors at each layer can be predicted by software. Therefore, the binary network can compress the activation image and skip a large number of redundant convolution calculations.

[0020] (2) The activation image encoding unit can losslessly compress the sparse activation image of each layer of the binary network and transmit the original activation image to the accelerator in the form of an activation vector. The accelerator can save the activation image as a valid vector, label value, and number of valid vectors based on the known sparse vector, thereby achieving a higher compression ratio, reducing the amount of data stored, and reducing the number of data reads and writes, thereby optimizing hardware power consumption;

[0021] (3) The coded image decoding unit simulates the binary network calculation process on the label value, and can skip redundant calculations based on the results predicted by the software, thereby reducing the power consumption and delay of the binary network calculation and further optimizing the energy efficiency of the binary network. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 Diagram of the sparse activation image encoder structure.

[0023] Figure 2 Sparse gesture edge input image.

[0024] Figure 3 Principles of sparse vector and efficient vector generation.

[0025] Figure 4 Activate the image coding unit structure diagram.

[0026] Figure 5 Activates the image encoding process.

[0027] Figure 6 Structural diagram of the coded image decoding unit.

[0028] Figure 7 The process of decoding coded images.

[0029] Figure 8 Schematic diagram of skipped calculation and effective calculation.

[0030] Figure 9 Compression ratio data.

[0031] Figure 10 Speed up comparison data. DETAILED DESCRIPTION

[0032] The present invention provides a sparse binary network activated image encoding and decoding hardware accelerator for gesture recognition, such as Figure 1 As shown, it is characterized in that the structure includes an activation image encoding unit, a coded image storage unit, and a coded image decoding unit; wherein:

[0033] like Figure 2 As shown, the present invention uses four commonly used gesture recognition datasets, corresponding to RGB, RGBD, and grayscale input images. These datasets are preprocessed to generate sparse gesture edge images as input to the binary network. Each pixel in this sparse gesture edge image is 1 bit and exhibits sparsity, ensuring that the first layer of the binary network performs binary computations and generates sparse vectors within each layer. The weights and inputs involved in the convolution computation in the binary network are 1 bit, with values of 0 and 1 representing -1 and +1, respectively.

[0034] like Figure 3As shown in Figure 1, the generation mechanism of sparse vectors and valid vectors. As shown in (a), the input image is a sparse gesture edge image. Then each 3x3 input block can be divided into two types: one is a 3x3 input block without +1 value, and the other is a 3x3 input block with +1 value. Due to the sparsity of the image, there must be a large number of 3x3 input blocks without +1 value, and the convolution calculation results of this part can be calculated by software. Therefore, there are a large number of repeated vectors in the output activation map of the convolution calculation. Each element of these vectors is 1 bit and is composed of all channels. As shown in (b), when the convolution is not the first layer, the 3x3 input block can also be divided into two parts: one does not contain the valid activation vector of the previous layer, and the other contains the valid activation vector of the previous layer. The result of the 3x3 input block that does not contain the valid activation vector of the previous layer can also be inferred by the software. Therefore, this result is the sparse vector of this layer, and the sparse vector exists in the activation image of each layer.

[0035] Activate the image coding unit, such as Figure 4As shown, it includes: a sparse vector filter, a label generator, a valid vector counter, and a row counter; when the activation image vector and its valid signal are sent to the sparse vector filter and the row counter, the sparse vector filter compares the input activation image vector with the predicted sparse vector provided by the software; if they are equal, the input activation image vector is a sparse vector, and a sparse judgment signal of 0 value is sent to the label generator and the valid vector counter; if they are not equal, the input activation image vector is a valid vector, and the valid vector is sent to the encoded image storage unit, and a sparse judgment signal of 1 value is sent to the label generator and the valid vector counter; the row counter generates a column count value based on the activation image valid signal. When the column count value is less than the image length, it indicates that the current input activation image vector does not reach the number of a row, and the next row signal next_row is kept at 0 and sent to the label generator and the valid vector counter; when the column count value is less than the image length , indicating that the current input activation image vector reaches the number of one row, the next row signal next_row is pulled up to 1 and sent to the label generator and the valid vector counter, and the column count value is reset to 1; the label generator receives the sparse judgment signal of the sparse vector filter and the next row signal next_row of the row counter, and first determines whether next_row is 1. If so, the label value of the label generator is sent to the coded image storage unit and the label value is reset to 0; if not, the lowest position of the label value is set to the value of the sparse judgment signal, and the label value is shifted one bit to the left; the valid vector counter receives the sparse judgment signal of the sparse vector filter and the next row signal next_row of the row counter, and first determines whether next_row is 1. If so, the number of valid vectors of the valid vector counter is sent to the coded image storage unit; if not, the number of valid vectors is added to the value of the sparse judgment signal to generate a new number of valid vectors.

[0036] Figure 5 For the specific encoding process, the original activation image can be regarded as a two-dimensional rectangle represented by the effective activation vector (EAV) and the sparse vector. Then the vector can be regarded as composed of the effective activation vector value and its position. Using CSR compression coding, CSR divides the sparse image into effective values, the first effective value address and the column coordinate. The present invention adopts CSR compression coding deformation to divide the original activation image into effective activation vectors, the first effective value address and coding, and uses coding instead of column coordinates to facilitate the convolution calculation process. Figure 5Taking the original activation image in as an example, the first valid value address is the number of all EAVs before the first EAV data on the leftmost side of each row of original data; therefore, the first row is 0, there are 3 EAVs before the first EAV in the second row, and so on. The label is the representation of the EAV and sparse vector in the original activation image, where each label represents a row of the original activation image, where the 0 value in the label represents the sparse vector, and the 1 value represents the EAV. The valid activation vector value is the set of all EAVs, arranged from left to right and from top to bottom according to the original activation image. The convolution operation is that each convolution kernel slides on the original activation image according to the convolution kernel size and a fixed step size, thereby performing the corresponding multiplication and accumulation operation. The present invention uses coding instead of the original activation image, and the 0 in the coding represents the sparse vector defined previously, and 1 represents an EAV. The position of each EAV can be obtained according to the coding and the first valid value address, and the sparse vector is a fixed value, so the data required for the calculation can be restored. Compression of sparse data can effectively reduce the storage resources required for the original activation image; in addition, calculations performed on the encoded data can skip some repeated calculations, thereby optimizing the calculation speed.

[0037] like Figure 6 As shown, the coded image decoding unit includes: an address generator, a label decoder, a pointer module, a valid vector address module, a 3x3 label comparator, a prediction sparse vector register and an activation image recovery module.

[0038] During decoding, the label decoder reads three rows of label data from the output label register stack, and slides from left to right on the three rows of label values in a 3x3 rectangular box according to the convolution sliding process. Each time, a 3x3 label value (tag) is taken out and sent to the pointer module, address generator module, 3x3 label comparator and activation image recovery module until it slides to the rightmost side, indicating that the convolution calculation corresponding to the current label is completed. A new row of label values is read from the output label register stack, and the first row of the original three rows of label values is discarded to form a new three rows of decoded data until all the output label register stack data is decoded.

[0039] The 3x3 tag comparator receives the 3x3 tag values and makes a judgment. If all 3x3 tag values are equal to 0, it indicates that the subsequent calculation results can be predicted by the software, and the subsequent calculation is skipped. Otherwise, the subsequent decoding process is carried out.

[0040] The effective vector address module reads the address of the row where the corresponding tag value is located from the first effective vector address register file.

[0041] The predicted sparse vector register stores the software predicted sparse activation image vector.

[0042] The pointer module stores the number of valid values from the leftmost position of the 3x3 label value to the leftmost position of the original label value. P1, P2, and P3 represent the number of valid values in three rows respectively. When the 3x3 label value is received, P1, P2, and P3 are sent to the address generator and P1, P2, and P3 are added to the leftmost value of the 3x3 label value and P1, P2, and P3 are updated.

[0043] The address generator generates an address based on the 3x3 label value, P1, P2, P3 of the pointer module, and the number of valid vectors in the valid vector address module; the address generator adds the number of valid vectors in the corresponding row, the pointer, and the corresponding label to generate the valid vector address corresponding to the label; the address generator determines whether each label value in the 3x3 label value is 0. If it is 0, the corresponding data is a sparse activation image vector and is sent to the activation image recovery module; if it is 1, the valid vector is read from the valid vector register stack according to the generated address and sent to the activation image recovery module; the activation image recovery module obtains 9 vectors, arranges them in order to form the original 3x3 activation image block, and sends it to the binary network calculation module for convolution calculation.

[0044] like Figure 7 The figure shows the activation image restoration process. The address is generated based on the first valid address and the label value, and the data is read according to the address to regenerate the original 3x3 input block. Figure 8 It is a classification of effective calculation and skip calculation. According to the 3x3 label value, it can be directly judged whether to skip the convolution.

[0045] The present invention uses four common gesture datasets: RGBDASL, RGBDGES, RGBASL, and GRAYASL to conduct experimental verification on the uncompressed binary neural network accelerator DBA and the energy-efficient sparse binary neural network hardware accelerator SBA of the present invention. In the experiment, the largest edge gesture image LGE and the smallest edge gesture image SME in each dataset are used to represent the worst compression effect (worst acceleration effect) and the best compression effect (best acceleration effect). Figure 9 As shown in , the accelerator achieves a compression ratio of more than 1.72-3.45 times in four commonly used gesture datasets; Figure 10 As shown, the SBA acceleration effect reaches 1.03-1.83 times.

Claims

1. A sparse binary network activated image encoding and decoding hardware accelerator for gesture recognition, characterized in that: The system comprises an activation image encoding unit, an encoding image storage unit, and an encoding image decoding unit. The input of the activation image encoding unit is the activation image vector of the sparse binary network and its valid signal. After encoding, a large number of identical vectors existing in the activation image of the sparse binary network are compressed. These identical vectors are generated by the binary gesture edge as input. The encoded valid vectors are stored in the encoding image storage unit. The encoding image decoding unit reads the encoded image information from the encoding image storage unit and decodes it. Repeated calculations are skipped during the decoding process, thereby reducing the number of calculations, improving the calculation speed, and improving the hardware performance. In particular: The activation image encoding unit includes a sparse vector filter, a label generator, a valid vector counter, and a row counter; when the activation image vector and the activation image vector valid signal are respectively sent to the sparse vector filter and the row counter, the sparse vector filter compares the input activation image vector with the predicted sparse vector provided by the software; if they are equal, the input activation image vector is a sparse vector, and a sparse judgment signal of 0 is sent to the label generator and the valid vector counter; if they are not equal, the input activation image vector is a valid vector, the valid vector is sent to the encoded image storage unit, and a sparse judgment signal of 1 is sent to the label generator and the valid vector counter; The row counter performs statistics based on the activation image valid signal to generate a column count value. When the column count value is less than the image length, it indicates that the current input activation image vector has not reached the number of one row, and the next row signal next_row is kept at 0 and sent to the label generator and the valid vector counter. When the column count value is greater than or equal to the image length, it indicates that the current input activation image vector has reached the number of one row, and the next row signal next_row is pulled up to 1 and sent to the label generator and the valid vector counter, and the column count value is reset to 1. The label generator receives the sparse determination signal of the sparse vector filter and the next row signal next_row of the row counter, first determines whether next_row is 1, and if so, sends the label value of the label generator to the coded image storage unit and resets the label value to 0; if not, sets the lowest position of the label value to the value of the sparse determination signal, and shifts the label value one position to the left; The valid vector counter receives the sparse determination signal of the sparse vector filter and the next row signal next_row of the row counter, first determines whether next_row is 1, and if so, sends the number of valid vectors in the valid vector counter to the coded image storage unit; if not, adds the number of valid vectors to the value of the sparse determination signal to generate a new number of valid vectors; During the activation image encoding process, the original activation image is regarded as a two-dimensional rectangle represented by a valid activation vector and a sparse vector, so the vector is regarded as consisting of the valid activation vector value and its position; using CSR compression coding deformation, CSR divides the sparse image into valid values, the first valid value address and column coordinates; CSR compression coding deformation is to divide the original activation image into valid activation vectors, the first valid vector address and codes, and use codes instead of column coordinates.

2. The sparse binary network activated image encoding and decoding hardware accelerator according to claim 1, characterized in that: The coded image storage unit includes a valid vector register stack, an output tag register stack and a first valid vector address register stack; the valid vector register stack receives the valid vector sent by the activated image coding unit and saves it in sequence; the output tag register stack receives the tag value sent by the activated image coding unit and saves it in sequence; the first valid vector address register stack receives the number of valid vectors sent by the activated image coding unit and saves it in sequence.

3. The sparse binary network activated image encoding and decoding hardware accelerator according to claim 2, characterized in that: The coded image decoding unit includes an address generator, a label decoder, a pointer module, a valid vector address module, a 3x3 label comparator, a prediction sparse vector register and an activation image recovery module; During decoding, the label decoder reads three rows of label data from the output label register stack, slides from left to right on the three rows of label values in a 3x3 rectangular frame according to the convolution sliding process, takes out one 3x3 label value at a time and sends it to the pointer module, the address generator module, the 3x3 label comparator and the activation image recovery module until it slides to the rightmost side, indicating that the convolution calculation corresponding to the current label is completed, reads a new row of label values from the output label register stack, discards the first row of the original three rows of label values, and forms three new rows of decoded data until all the output label register stack data are decoded; The 3x3 tag comparator receives the 3x3 tag values and makes a judgment; if all 3x3 tag values are equal to 0, it indicates that the subsequent calculation results can be predicted by the software, and the subsequent calculation is skipped; otherwise, the subsequent decoding process is performed; The effective vector address module reads the address of the row where the corresponding tag value is located from the first effective vector address register file; The predicted sparse vector register stores the sparse activation image vector predicted by the software; The pointer module stores the number of valid values from the leftmost position of the 3x3 label value to the leftmost position of the original label value, and uses P1, P2, and P3 to represent the number of valid values of three rows respectively. When the 3x3 label value is received, P1, P2, and P3 are sent to the address generator and P1, P2, and P3 are added to the leftmost value of the 3x3 label value and P1, P2, and P3 are updated; The address generator generates an address based on the 3x3 label value, P1, P2, P3 of the pointer module and the number of valid vectors of the valid vector address module; the address generator adds the number of valid vectors of the corresponding row, the pointer and the corresponding label to generate the valid vector address corresponding to the label; the address generator determines whether each label value in the 3x3 label value is 0. If it is 0, the corresponding data is a sparse activation image vector and is sent to the activation image recovery module; if it is 1, the valid vector is read from the valid vector register stack according to the generated address and sent to the activation image recovery module; the activation image recovery module obtains 9 vectors, arranges them in order to form the original 3x3 activation image block, and sends it to the binary network calculation module for convolution calculation.

Citation Information

Patent Citations

  • A non-reducing sparse data encoding and decoding circuit applied to a convolutional neural network and a coding and decoding method thereof

    CN109104197A

  • A fast video inference method based on cyclic residual module

    CN109272035A