Low-energy-consumption pulse network video identification coprocessor and method

By designing a low-energy pulse network video recognition coprocessor, the problems of insufficient utilization of computing resources, excessive power consumption and hardware design complexity in video recognition applications are solved, and efficient video recognition and significant energy efficiency improvement are achieved.

CN120146114APending Publication Date: 2025-06-13XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510212489.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Traditional pulsed neural networks face problems such as insufficient utilization of computing resources, excessive power consumption and hardware design complexity in video recognition applications, especially in scenarios with high real-time requirements.

Method used

Design a low-energy-consuming pulse network video recognition coprocessor, including a data processing module and a pulse neural network computing unit. The data processing module is responsible for preprocessing and feature extraction of video data. The pulse neural network computing unit adopts a step-by-step pruning strategy, and the time step is set to T=1. Combined with 8bit quantization, the network weight and data are optimized, which significantly reduces the power consumption of hardware calculation and data exchange.

Benefits of technology

Efficient video recognition is realized, which significantly reduces the power consumption of hardware computing and data exchange and improves the computing energy efficiency ratio, especially in handwritten digital recognition applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146114A_ABST
    Figure CN120146114A_ABST
Patent Text Reader

Abstract

The invention relates to the field of pulse neural network circuits, in particular to a low-energy-consumption pulse network video recognition coprocessor and method. And the processor acquires external video data through the data processing module, identifies the input video by using the trained pulse neural network calculation unit, and determines the category of the input video. According to the method, a hardware-friendly pulse neural network design is adopted, efficient time processing is realized in combination with an event-driven mechanism, the model complexity is reduced by using a strategy that the time step length T is equal to 1, and the influence of instability of video data on system performance is reduced. Meanwhile, the network weight and data are optimized by adopting 8-bit quantification, and parameters and intermediate results are managed in combination with an on-chip storage mode, so that the power consumption of hardware calculation and data exchange is remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of spiking neural network circuits, and particularly to a low-power spiking network video recognition coprocessor and method. Background Art

[0002] With the rapid development of artificial intelligence and deep learning technologies, neural networks, especially convolutional neural networks, have achieved remarkable applications in many fields such as image processing, speech recognition, and natural language processing. However, traditional neural networks are usually based on continuous signals and digital operations and cannot effectively simulate the working principle of biological nervous systems. To further improve the computational efficiency and energy efficiency of neural networks, spiking neural networks, as a computational model closer to biological nervous systems, have gradually become an important research direction in the field of artificial intelligence.

[0003] Different from traditional neural networks, the core feature of spiking neural networks is that neurons transmit information through discrete spike signals. Spiking neural networks can better simulate the bioelectric activities of the brain and thus have great potential in biological simulation, low-power computing, and real-time processing. At present, the research on the theory of spiking neural networks has been very complete, but the deployment and implementation in daily applications have encountered relatively large limitations. Network acceleration processors based on traditional von Neumann computing architectures such as GPUs and CPUs are not suitable for many actual life scenarios due to their high costs and large power consumption, and the use of dedicated devices to accelerate the inference of spiking neural networks has gradually come into view.

[0004] In terms of hardware accelerators, in recent years, accelerators based on programmable logic devices (FPGAs) have gradually become an important direction to solve this problem. FPGAs have the advantages of high parallel processing capabilities and flexible hardware designs and are thus widely used in fields such as image processing, machine learning, and neural networks. Compared with traditional CPUs and GPUs, FPGAs can provide lower power consumption and higher computational efficiency for specific tasks; however, mapping spiking neural networks onto FPGAs still faces problems such as insufficient utilization of computing resources, excessive power consumption, and high hardware design complexity. Especially in application scenarios with high real-time requirements, these problems are the bottlenecks restricting the widespread application of FPGAs in video recognition. Summary of the Invention

[0005] Aiming at the problems mentioned in the prior art, the present invention proposes a low-power spiking network video recognition coprocessor and method to solve the problems such as insufficient utilization of computing resources, excessive power consumption, and high hardware design complexity in the background art.

[0006] To achieve the above object, the present invention adopts the following technical solutions: The present invention provides a low-power pulsed network video recognition co-processor, comprising: A data processing module, which can convert the acquired video data into a feature map and is used to display the recognition result from the pulsed neural network calculation unit; A pulsed neural network calculation unit, which is connected to the data processing module and is used to receive the feature map from the data processing module, extract features from the feature map, and obtain a recognition result.

[0007] As a further improvement of the present invention, the data processing module includes a connection module and a preprocessing module; the connection module includes an I / O interface for collecting video data and an HDMI interface for feeding back the recognition result; the preprocessing module includes a FIFO, a DDR memory, and an image processing module.

[0008] As a further improvement of the present invention, the pulsed neural network calculation unit includes a first convolutional layer, a second convolutional layer, a batch normalization layer, a neuron activation layer, a max pooling layer, a fully connected layer, and a classifier.

[0009] A method for a low-power pulsed network video recognition co-processor is implemented by using the above device, and includes the following steps: S1. The data processing module converts the acquired video data into a feature map and outputs it to the pulsed neural network calculation unit; S2. The pulsed neural network calculation unit extracts features from the feature map to obtain a recognition result, and feeds the recognition result back to the data processing module for display.

[0010] As a further improvement of the present invention, the process of converting the feature map in S1 is as follows: Crop the video image to meet the required size of the pulsed neural network calculation unit; Perform grayscale processing on the cropped image to obtain a feature map.

[0011] As a further improvement of the present invention, the grayscale processing includes RGB data grayscale and pixel average grayscale.

[0012] As a further improvement of the present invention, the time step of the pulsed neural network in the pulsed neural network calculation unit is 1, and the forward propagation path is unique.

[0013] As a further improvement of the present invention, the weight value of the pulsed neural network is 8-bit weight.

[0014] As a further improvement of the present invention, the specific process of obtaining the recognition result in S2 is as follows: S21. Input the feature map into the first convolutional layer for convolutional operation to obtain a convolved feature map; S22. Input the convolved feature map into the batch normalization layer, and control the data size in the convolved feature map by normalizing the input data of each layer; S23. Input the data after batch normalization into the neuron activation layer for activation; S24. Input the activated data into the max pooling layer for dimensionality reduction. Reduce the data size and retain important feature information by selecting the maximum value in each local region of the feature map; S25. Input the pooled data into the second convolutional layer for a second convolution operation to obtain a feature map after the second convolution; S26. Repeat the process of S23 to S24 for the feature map after the second convolution, and output the final feature map; S27. Convert the final feature map into a one-dimensional array. After the array is activated by the neuron activation layer, connect the array to the classifier through the fully connected layer to judge the image category and output the recognition result.

[0015] As a further improvement of the present invention, the neuron activation layer in S23 is a LIF neuron.

[0016] The present invention has achieved the following technical effects compared with the prior art: The spiking neural network video recognition coprocessor of the present invention collects video data through the data processing module, and uses the trained spiking neural network calculation unit to recognize the input video and determine its category; the present invention adopts a hardware-friendly spiking neural network design, combines the event-driven mechanism to achieve efficient time processing, and uses the strategy of time step T = 1 to reduce the model complexity and mitigate the impact of video data instability on the system performance; at the same time, 8-bit quantization is used to optimize the network weights and data, and the on-chip storage method is combined to manage the parameters and intermediate results, significantly reducing the power consumption of hardware calculation and data exchange. The present invention has a higher computational energy efficiency ratio compared with the CPU implementation scheme in handwritten digit recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a schematic structural diagram of the spiking neural network video recognition coprocessor of the present invention; Figure 2 It is a schematic structural diagram of the data processing module of the present invention; Figure 3 It is a schematic process diagram of the image processing module of the present invention; Figure 4 It is a schematic recognition process diagram of the spiking neural network in the spiking neural network calculation unit of the present invention; Figure 5 It is a schematic structural diagram of the spiking neural network calculation unit of the present invention; Figure 6This is the flowchart of the operation of the spiking neural network video recognition coprocessor of the present invention. Detailed implementation manners

[0018] In the following text, only some exemplary embodiments are simply described. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the spirit or scope of the present invention. Therefore, the drawings and the description are regarded as being exemplary in nature rather than restrictive.

[0019] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.

[0020] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.

[0021] In the present invention, unless otherwise clearly specified and defined, the terms "mounted", "connected", "coupled", "fixed", etc. should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection, an electrical connection, or a communication connection; it may be directly connected, or indirectly connected through an intermediate medium, and it may be the internal communication of two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0022] In the present invention, unless otherwise clearly specified or limited, the first feature being “on” or “under” the second feature may include direct contact between the first and second features, or may include the situation where the first and second features are not in direct contact but in contact through additional features therebetween. Moreover, the first feature being “above”, “over” and “on top of” the second feature includes that the first feature is directly above and obliquely above the second feature, or merely indicates that the horizontal height of the first feature is higher than that of the second feature. The first feature being “under”, “below” and “beneath” the second feature includes that the first feature is directly below and obliquely below the second feature, or merely indicates that the horizontal height of the first feature is less than that of the second feature.

[0023] It should be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups.

[0024] It should also be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms “a”, “an” and “the” are intended to include the plural forms.

[0025] It should be further understood that the term “and / or” used in the specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0026] Schematic diagrams of various structures according to the disclosed embodiments of the present invention are shown in the drawings. These figures are not drawn to scale, where for the purpose of clear expression, some details are enlarged and some details may be omitted. The shapes of various regions and layers shown in the figures and their relative sizes and positional relationships are only exemplary, and may actually deviate due to manufacturing tolerances or technical limitations, and those skilled in the art can design regions / layers with different shapes, sizes and relative positions according to actual needs.

[0027] Embodiments of the present invention will be described in detail below with reference to the drawings.

[0028] As Figure 1 shown, the present invention provides a low-power pulsed network video recognition coprocessor, including: A data processing module, which can convert the acquired video data into a feature map and is used to display the recognition result from the pulsed neural network calculation unit; The spiking neural network computing unit is connected to the data processing module and is used to receive the feature map from the data processing module, extract features from the feature map, and obtain the recognition result.

[0029] As shown in the embodiment Figure 2 The data processing module of the present invention includes a connection module and a preprocessing module. The connection module includes an I / O interface for collecting video data and an HDMI interface for feeding back the recognition result. Specifically, in use, the I / O interface is used to connect to the device for obtaining video data, such as a camera, a camera, etc.; the HDMI interface is used to connect to the device for displaying the recognition result, such as a display screen.

[0030] In the embodiment, the preprocessing module includes a FIFO, a DDR memory, and an image processing module. The FIFO is used to temporarily store data to ensure that the data is read and output in the order of entry; Figure 2 As shown, it includes a DDR controller and a DDR3 (Double Data Rate 3) memory that cooperate to achieve high-speed data transmission.

[0031] When the data processing module of this embodiment is in use, the camera collects video data through the IO interface and converts it into 16-bit data and stores it in the FIFO, thereby ensuring the smooth flow of data; in order to provide fast access for subsequent processing, the data is then written into the DDR memory, and the data flow and status control are coordinated through the AXI bus. After the image processing module processes the image data, the processor sends the processed image data to the spiking neural network computing unit via the AXI bus protocol for image recognition operation of the spiking neural network coprocessor, and finally feeds back the recognition result to the display screen through the HDMI interface.

[0032] As Figure 3 shown, in the embodiment, the process of the image processing module processing the video data is mainly cropping, grayscale conversion of RGB data, and pixel average grayscale conversion.

[0033] In this embodiment, since the size of the image collected by the camera is 640×480, and the data input to the input layer of the spiking neural network is a 28×28 grayscale image, it is necessary to transform the resolution of the input image. The 640×480-sized image is centered and taken as a 476×476-sized image. The cropped 476×476 image can be further divided into 28×28 squares with a size of 17×17.

[0034] The grayscale conversion of RGB data is based on the cropped image. Taking a 17×17 small grid as a unit, the grayscale conversion results of the RGB data corresponding to each grid are calculated, and the 28×28 grids of 17×17 are numbered to facilitate the storage of the collected data. Then, the pixels in each 17×17 grid are grayscale-converted. The weighted grayscale algorithm is adopted in this embodiment:

[0035] In the formula: B is the pixel value of the blue channel, G is the pixel value of the green channel, R is the pixel value of the red channel, is the pixel value of the finally generated grayscale image.

[0036] Among them, B, G, and R are in the same format as the RGB565 data format. After grayscale-converting each pixel, the grayscale values of these 17×17 pixels are added together, and then the average value is taken to obtain the grayscale value of the 17×17 grid image. After calculation, 28×28 grayscale results are obtained for the image data input of the final neural network calculation unit. Among them, the average value calculation method is:

[0037] In the formula, is a certain pixel value of the grayscale image with a size of 28×28 input to the neural network, is the corresponding pixel value of the grayscale image in the feature map with a size of 476×476.

[0038] As Figure 4 shown, the recognition process of the pulse neural network calculation unit in this embodiment is as follows: a. The input image with a size of 28×28 is padded and abstracted into different feature maps through the first convolutional layer. After 8 convolutional kernels with a size of 3×3 perform convolutional operations on the input image, an 8-channel feature vector is output. The data size of each channel is 28×28, and the output dimension is 8×28×28.

[0039] b. The feature map after the convolutional operation is input to the batch normalization layer, and the input data of each layer is standardized to control the data size in the feature map, avoid training instability caused by uneven input data distribution, and prevent overfitting.

[0040] c. The data after batch normalization is input to the pulse neuron. The pulse neuron is an LIF neuron, which determines whether to generate a pulse by calculating whether the membrane potential reaches the threshold; after being activated by the LIF neuron, the eigenvalue will be converted into 0 or 1, indicating that the neuron is activated and emits a pulse.

[0041] d. Perform max pooling (MP) operation on the activated data for dimensionality reduction. By selecting the maximum value in each local region of the feature map, the size of the data is reduced while important feature information is retained, and it can ensure that information is still transmitted by pulses. After max pooling, the size of the feature map is reduced from 28×28 to 14×14, and the number of channels remains unchanged.

[0042] e. Input the pooled data into the second convolutional layer. The number of channels is reduced to half of the original. Through convolutional operations, the network continues to extract high-order features of the image. The dimension of the output feature map is 4×14×14; f. Perform batch normalization and max pooling operations on the data after the second convolution again. The size of the finally output feature data map is 4×7×7; g. Convert the feature data map into a one-dimensional array, and the size of the data tensor is 196×1; h. After activation by LIF neurons and passing through the fully connected layer, connect the data to the classifier to judge the image category and output the recognition result.

[0043] To ensure the light weight of the network, the spiking neural network deployed on the spiking neural network computing unit in the embodiment can gradually compress the time step of the spiking neural network to T = 1 by using the method of gradual pruning.

[0044] Specifically, at the initial stage of training the spiking neural network, the spiking neural network is first pre-trained with a relatively large time step (such as T = 5). This provides enough time for the network to accumulate membrane potential, enabling neurons to effectively emit pulses and transmit information. As the training progresses, the time step is gradually shortened. While ensuring that the network can be effectively trained and converge, the computational delay is gradually reduced. At the same time, the hardware implementation of the gradual pruning method can also reduce the pressure of hardware computing and storage.

[0045] Since the computation of the spiking neural network usually has high parallelism and timing, reducing the time step not only reduces the demand for hardware resources, but also enables the hardware to perform computing and storage operations more efficiently, thereby improving the overall energy efficiency and performance. In addition, since T = 1, the output result of the processor is only related to the input at the current moment, avoiding the instability of the output result caused by inconsistent light input by the camera at different moments.

[0046] The weight value of the spiking neural network in this embodiment adopts 8-bit quantization, and the quantization method is expressed as:

[0047] In the formula: k represents the number of quantization bits, and k takes 8 in the present invention; represents rounding x to the nearest integer; Represents the input value in the range [0, 1] that needs to be quantized; Represents the output result after k-bit quantization, which is a discrete value in the range [0, 1].

[0048] As Figure 5 As shown, the usage process of the pulse neural network computing unit in this embodiment is as follows: The weight data is stored in the BRAM and transmitted to the convolutional layer, batch normalization layer, and fully connected layer through the weight controller. The convolutional layer specifically includes the first convolutional layer and the second convolutional layer. For the first convolutional layer and the batch normalization layer, since multi-bit multiplication operations need to be completed, the computing unit calls the on-chip DSP module to improve the computing efficiency. In the second convolutional layer and the fully connected layer, since the input data is a pulse signal, the multiplication operation between it and the weight is simplified to an addition operation, significantly reducing the computing complexity and power consumption. In addition, the pooling layer adopts the maximum pooling strategy to ensure that the subsequent information is still transmitted in the form of pulses, thereby making full use of the event-driven characteristics of the pulses to further reduce power consumption. Since the time step is set to T = 1, the network model is compressed, and the weight data and intermediate computing results can be fully stored on the chip, significantly reducing the delay and energy consumption of data transmission. Finally, the classifier determines the final classification result by counting the categories of pulse emissions.

[0049] As shown in the embodiment Figure 6 As shown, after the system is powered on, the coprocessor enters the initialization phase, reads in the pre-trained network parameters. After the system acquires the external image to be recognized through the camera and performs preprocessing operations such as data stitching, grayscale conversion, and image size cropping, it calls the pulse neural network computing unit to calculate the data. After the calculation is completed, the display shows the recognition result.

[0050] The above shows and describes the basic principles, main features, and advantages of the present invention. For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic features of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights.

[0051] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art. The above content is only to illustrate the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any modification made on the basis of the technical solution according to the technical idea proposed by the present invention falls within the protection scope of the claims of the present invention.

Claims

1. A low-energy pulse network video recognition coprocessor, comprising: A data processing module, which can convert the acquired video data into a feature map and is used to display the recognition results from the spiking neural network calculation unit; The pulse neural network calculation unit is connected to the data processing module and is used to receive the feature map from the data processing module, perform feature extraction on the feature map, and obtain the recognition result.

2. A low-energy pulse network video recognition coprocessor according to claim 1, characterized in that: The data processing module includes a connection module and a preprocessing module; The connection module includes an I / O interface for collecting video data and an HDMI interface for feeding back recognition results; The pre-processing module includes FIFO, DDR memory and image processing module.

3. A low-energy pulse network video recognition coprocessor according to claim 1, characterized in that: The pulse neural network computing unit includes a first convolution layer, a second convolution layer, a batch normalization layer, a neuron activation layer, a maximum pooling layer, a fully connected layer and a classifier.

4. A method for a low-energy pulse network video recognition coprocessor, characterized in that: The method is implemented by using the device according to any one of claims 1 to 3, comprising the following steps: S1, the data processing module converts the acquired video data into a feature map and outputs it to the pulse neural network calculation unit; S2, the pulse neural network calculation unit extracts features from the feature map to obtain recognition results, and transmits the recognition results back to the data processing module for display.

5. The method of a low-energy pulse network video recognition coprocessor according to claim 4, characterized in that: The process of converting the feature map in S1 is: Crop the video image to a size that meets the requirements of the pulse neural network computing unit; The cropped image is grayed out to obtain the feature map.

6. The method of a low-energy pulse network video recognition coprocessor according to claim 5, characterized in that: Grayscale processing includes RGB data grayscale and pixel average grayscale.

7. The method of a low-energy pulse network video recognition coprocessor according to claim 4, characterized in that: The time step of the spiking neural network in the spiking neural network computing unit is 1, and the forward propagation path is unique.

8. The method of a low-energy pulse network video recognition coprocessor according to claim 7, characterized in that: The weight value of the pulse neural network is 8-bit weight.

9. The method of a low-energy pulse network video recognition coprocessor according to claim 4, characterized in that: The specific process of obtaining the recognition result in S2 is: S21, inputting the feature map into the first convolutional layer to perform convolution operation to obtain a convolved feature map; S22, inputting the convolved feature map to a batch normalization layer, and controlling the data size in the convolved feature map by standardizing the input data of each layer; S23, inputting the batch normalized data into the neuron activation layer for activation; S24, inputting the activated data into the maximum pooling layer for dimensionality reduction processing, reducing the data size and retaining important feature information by selecting the maximum value of each local area in the feature map; S25, input the pooled data into the second convolution layer for secondary convolution operation to obtain a feature map after secondary convolution; S26, repeat the process from S23 to S24 for the feature map after the secondary convolution, and output the final feature map; S27. Convert the final feature map into a one-dimensional array. After the array is activated by the neuron activation layer, it is connected to the classifier after passing through the fully connected layer to judge the image category and output the recognition result.

10. The method of a low-energy pulse network video recognition coprocessor according to claim 9, characterized in that: The neuron activation layer in S23 is LIF neurons.