A Neural Network Image Recognition Method and Electronic Device Based on Pulse Statistics

Through the methods of pulse excitation frequency encoding and pulse count statistics, the computing logic of the convolution layer, maximum pooling layer and Softmax layer of the SNN network is simplified, which reduces the calculation amount and hardware resource consumption, solves the problem of high power consumption of the SNN network, and is suitable for portable low-power devices.

CN115311527BActive Publication Date: 2025-07-29XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210713237.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-22
Publication Date
2025-07-29
Estimated Expiration
2042-06-22

AI Technical Summary

Technical Problem

The existing SNN network image recognition methods have high power consumption, mainly due to the complex logic operations of the convolution layer, the maximum pooling layer and the Softmax layer, which cannot meet the needs of portable low-power devices.

Method used

The pulse excitation frequency encoding method is used to convert the image pixel value into a pulse sequence, and the output of the maximum pooling layer and the Softmax layer is realized through the pulse count statistics, simplifying the operation logic and reducing the calls of the multiplier and adder.

Benefits of technology

It reduces the computing volume and hardware resource consumption, meets the image recognition needs of portable low-power devices, and has a wider application prospect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115311527B_ABST
    Figure CN115311527B_ABST
Patent Text Reader

Abstract

The present invention discloses a neural network image recognition method and an electronic device based on pulse statistics, including: training a CNN network using training image data to obtain a trained CNN network, and extracting network parameters of the convolutional layer and the fully connected layer in the CNN network; converting pixel values of an image to be recognized into a pulse sequence by using a pulse excitation frequency encoding method; using an SNN network to recognize the pulse sequence to obtain an image recognition result; wherein, the convolutional layer and the fully connected layer in the SNN network adopt the same network structure and network parameters as the convolutional layer and the fully connected layer in the CNN network; the max pooling layer and the Softmax layer in the SNN network adopt the same network structure as the max pooling layer and the Softmax layer in the CNN network, and the output corresponding to the max pooling layer and the Softmax layer in the SNN network is realized by using a pulse number statistics method. The present invention can meet the image recognition requirements of existing portable devices with lower power consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image recognition, and particularly relates to a neural network image recognition method based on pulse statistics and an electronic device. Background Art

[0002] With the development of deep neural network image recognition technology, the network structure and training algorithm are relatively mature, and good results have been achieved in image recognition accuracy. However, as the scale of convolutional neural networks continues to increase, the power consumption required for network inference has increased significantly, which limits the further application of convolutional neural networks. Spiking Neural Network (SNN), known as the "third-generation neural network", is more in line with biological characteristics. Different from traditional Convolutional Neural Network (CNN), it transmits data through discrete pulses rather than continuous activation values, and neurons are only activated when receiving input pulses. The SNN network has the characteristic of sparse computation, which can save energy consumption, making it a research hotspot in the field of image recognition. However, the output values of neurons in the SNN network are discrete sequences, making the training method of the CNN network unable to be directly applied to the SNN network. Therefore, converting the trained image recognition CNN network into an existing SNN network through parameter migration has become an effective method to improve the accuracy of the SNN network. The currently commonly used image recognition SNN network was proposed by Bodo Rueckauer et al. from the University of Zurich, Switzerland in 2017. It uses the CNN network for parameter migration to construct the existing SNN network, and the constructed existing SNN network has comparable image recognition accuracy to the CNN network with the same structure.

[0003] However, there are still two problems in the existing SNN network image recognition method: First, the pulse coding process of the pixel values of the image to be recognized needs to be realized by means of the activation values of the convolutional layer in the CNN network. This process requires mixing the convolutional layer in the CNN network into the SNN network, resulting in a large number of multiplication calculations still existing in the SNN network. Second, there are complex logical operations in the max pooling layer and Softmax layer in the SNN network. Its operation logic is comparable to that of the CNN network, and complex dedicated circuits need to be designed when implemented on FPGA, consuming a relatively large amount of hardware resources, resulting in relatively high power consumption. It can be seen that both the existing SNN network image recognition method and the CNN network image recognition method have the problem that due to the complex design of the network structure logic circuit, the hardware resource consumption increases, and then the power consumption increases, making them unable to meet the image recognition requirements of existing portable devices with lower power consumption. Summary of the Invention

[0004] To solve the above problems existing in the prior art, the present invention provides a neural network image recognition method and an electronic device based on pulse statistics. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0005] In a first aspect, an embodiment of the present invention provides a neural network image recognition method based on pulse statistics, including:

[0006] Training a CNN network using training image data to obtain a trained CNN network, and extracting network parameters corresponding to the convolutional layer and the fully connected layer in the trained CNN network;

[0007] Converting the pixel values of the image to be recognized into a pulse sequence using a pulse excitation frequency encoding method;

[0008] Using an SNN network to recognize the pulse sequence to obtain an image recognition result; wherein, the convolutional layer and the fully connected layer in the SNN network adopt the same network structure and network parameters as the convolutional layer and the fully connected layer in the trained CNN network, and the output of the corresponding network layer is obtained based on the network structure and network parameters; the max pooling layer and the Softmax layer in the SNN network adopt the same network structure as the max pooling layer and the Softmax layer in the trained CNN network, and the output corresponding to the max pooling layer and the Softmax layer in the SNN network is realized using a pulse number statistics method, and the output of the Softmax layer is used as the image recognition result.

[0009] In an embodiment of the present invention, the converting the pixel values of the image to be recognized into a pulse sequence using a pulse excitation frequency encoding method includes:

[0010] Calculating the average value of all pixel values of the image to be recognized, and using the average value as the dynamic excitation threshold of all pixel values;

[0011] Using the characteristics of the SNN network structure, inputting the pixel values of the image to be recognized into the pulse neurons corresponding to the SNN network; for each pulse neuron, including:

[0012] For each iteration moment process, including: calculating the membrane potential value of the pulse neuron at the current iteration moment according to each pixel value in the image to be recognized; comparing the membrane potential value of the pulse neuron at the current iteration moment with the dynamic excitation threshold, and outputting the pulse value at the current iteration moment according to the comparison result;

[0013] Until the maximum number of iteration moments is reached, outputting the pulse sequence corresponding to the pulse neuron.

[0014] In an embodiment of the present invention, the formula for calculating the membrane potential value of the pulse neuron at the current iteration moment is expressed as:

[0015] V i V(t)=V(t - 1)+p i (t - 1)+p i ;

[0016] Among them, V i (t) represents the membrane potential value of the i-th spiking neuron at the t-th iteration moment, V i (t - 1) represents the membrane potential value of the i-th spiking neuron at the (t - 1)-th iteration moment, p i represents the gray value of the i-th pixel, and t takes values from 1 to n, where n is the maximum number of iteration moments.

[0017] In an embodiment of the present invention, the method of converting the pixel values of the image to be recognized into a pulse sequence by using the pulse firing frequency encoding method further includes:

[0018] When the membrane potential value of the spiking neuron at the current iteration moment is greater than the dynamic firing threshold, reset the membrane potential value of the spiking neuron at the current iteration moment; recalculate the membrane potential value of the spiking neuron at the next iteration moment according to the reset membrane potential value of the spiking neuron and each pixel value in the image to be recognized; make a comparison according to the membrane potential value of the spiking neuron at the next iteration moment;

[0019] Until the maximum number of iteration moments is reached, output the pulse sequence corresponding to this spiking neuron.

[0020] In an embodiment of the present invention, the formula for resetting the membrane potential value of the spiking neuron at the current iteration moment is expressed as:

[0021] V i '(t)=V i (t)-255;

[0022] Among them, V i '(t) represents the membrane potential value of the i-th spiking neuron after reset at the t-th iteration moment, V i (t) represents the membrane potential value of the i-th spiking neuron at the t-th iteration moment;

[0023] Correspondingly, the formula for recalculating the membrane potential value of the spiking neuron at the next iteration moment is expressed as:

[0024] V i (t + 1)=V i '(t)+p i ;

[0025] Among them, V i (t + 1) represents the membrane potential value of the i-th spiking neuron at the (t + 1)-th iteration moment.

[0026] In one embodiment of the present invention, the CNN network includes a first convolutional layer, a first max pooling layer, a first fully connected layer, and a first Softmax layer connected in sequence;

[0027] Correspondingly, the SNN network includes a second convolutional layer, a second max pooling layer, a second fully connected layer, and a second Softmax layer connected in sequence;

[0028] Among them, the second convolutional layer and the second fully connected layer respectively adopt the same network structure and network parameters as the first convolutional layer and the first fully connected layer, and output pulse sequences of corresponding network layers based on the network structure and network parameters;

[0029] The second max pooling layer and the second Softmax layer respectively adopt the same network structure as the first max pooling layer and the first Softmax layer, and the outputs of their corresponding network layers are respectively formed by counting the number of pulses output by the second convolutional layer and the number of pulses output by the second fully connected layer.

[0030] In one embodiment of the present invention, the process of realizing the output of the second max pooling layer in the SNN network by using the pulse number statistics method includes:

[0031] Collect the pulse sequence output by the second convolutional layer according to the sampling kernel size;

[0032] For the pulse sequence of each sampling kernel, it includes: counting the number of pulses in each pulse sequence in the sampling kernel; finding the maximum value of the number of pulses from all the counted number of pulses corresponding to the sampling kernel; extracting the address of the pulse sequence corresponding to the maximum value of the number of pulses, and retrieving the pulse sequence of the sampling kernel according to the extracted address;

[0033] The pulse sequences of all sampling kernels form the pulse sequence output by the second max pooling layer.

[0034] In one embodiment of the present invention, the process of realizing the output of the second Softmax layer in the SNN network by using the pulse number statistics method includes:

[0035] Count the number of pulses in the pulse sequence output by the second fully connected layer;

[0036] Find the maximum value of the number of pulses from all the counted number of pulses;

[0037] Extract the address of the pulse sequence corresponding to the maximum value of the number of pulses, and retrieve and form the output of the second Softmax layer according to the extracted address, and use this output as the image recognition result.

[0038] In one embodiment of the present invention, the address of the pulse sequence corresponding to the maximum number of pulses is extracted and expressed as:

[0039] D = addr[find(S j :if(M j = M max ))], j ∈ [1, J];

[0040] wherein, addr represents the extraction address, J represents the maximum value of j. For the output formation of the second max - pooling layer, J takes the value of w×h, (w, h) represents the sampling kernel size. For the output formation of the second Softmax layer, J takes the value of K, K represents the number of pulse sequences output by the second fully - connected layer, find(S j :if(M j = M max ) represents finding the maximum value M max of the number of pulses, j and the corresponding pulse sequence S j when M max = max(M j ), j ∈ [1, J], M j represents the number of pulses of the pulse sequence S j statistically counted, represents the t - th pulse value in the pulse sequence S j , n represents the length of the pulse sequence.

[0041] In a second aspect, an embodiment of the present invention provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus;

[0042] The memory is used to store a computer program;

[0043] When the processor executes the program stored in the memory, it implements the steps of any one of the above - mentioned neural network image recognition methods based on pulse statistics.

[0044] Advantages of the present invention:

[0045] The neural network image recognition method based on pulse statistics proposed by the present invention ensures the recognition accuracy of the image to be recognized. During the encoding process of the image to be recognized, the pulse excitation frequency encoding method is used to convert the pulse sequence, and the output of the max-pooling layer and the Softmax layer is formed based on the pulse number statistics. Thus, the introduction of the CNN convolutional layer is avoided, and the operation logic of the max-pooling layer and the Softmax layer in the SNN network is simplified. By using the adder's accumulation calculation instead of complex logic circuits through the pulse data statistics method, the overall computational complexity is further reduced compared with the existing SNN network. In the FPGA implementation, since the calls of multipliers and adders are reduced, the consumed FPGA hardware resources are reduced, and the on-chip power consumption of the FPGA is lowered, which is more conducive to the deployment of the SNN network in the FPGA, and thus has a broader application prospect, such as meeting the image recognition requirements of existing portable devices with lower power consumption.

[0046] The following will further elaborate on the present invention in conjunction with the accompanying drawings and embodiments. Brief Description of the Drawings

[0047] Figure 1 is a schematic flowchart of a neural network image recognition method based on pulse statistics provided by an embodiment of the present invention;

[0048] Figure 2 is a schematic diagram of the structures of a CNN network and an SNN network provided by an embodiment of the present invention, and their corresponding relationship;

[0049] Figure 3 is a schematic diagram of the structure of another SNN network provided by an embodiment of the present invention;

[0050] Figure 4 is a schematic flowchart of converting the pixel values of the image to be recognized into a pulse sequence provided by an embodiment of the present invention;

[0051] Figure 5 is a schematic flowchart of the output formation process of the second max-pooling layer provided by an embodiment of the present invention;

[0052] Figure 6 is a schematic flowchart of the output formation process of the second Softmax layer provided by an embodiment of the present invention;

[0053] Figure 7 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed Embodiments

[0054] The following will further elaborate on the present invention in conjunction with specific embodiments, but the implementation manners of the present invention are not limited thereto.

[0055] Embodiment 1

[0056] To meet the image recognition requirements of portable and lower-power devices while ensuring the accuracy of image recognition, please refer to Figure 1 In an embodiment of the present invention, a neural network image recognition method based on pulse statistics is provided, including the following steps:

[0057] S10. Train a CNN network using training image data to obtain a trained CNN network, and extract the network parameters corresponding to the convolutional layer and the fully connected layer in the trained CNN network.

[0058] Specifically, for the existing SNN network, a method of parameter transfer using a CNN network is used to construct the existing SNN network. Therefore, when constructing the existing SNN network, first train a CNN network using training image data to obtain a trained CNN network. The specific training process is not limited to an algorithm, such as the Adma algorithm. Then, extract the network parameters corresponding to each network layer in the trained CNN network for constructing the network parameters of the corresponding network layer in the existing SNN network. It can be seen that the SNN network has the same network structure as the CNN network, the same arrangement order of network layers, and the same number of network layers.

[0059] The extracted network parameters can be normalized, and the specific normalization method can adopt the existing technology.

[0060] However, the image recognition method of the existing SNN network has the above two problems. Therefore, in an embodiment of the present invention, a neural network image recognition method based on pulse statistics is proposed. The SNN network in the embodiment of the present invention adopts the same network structure as the CNN network. Specifically, the SNN network used is: the convolutional layer and the fully connected layer in the SNN network adopt the same network structure and network parameters as the convolutional layer and the fully connected layer in the trained CNN network, and the output of the corresponding network layer is obtained based on the network structure and network parameters; the max pooling layer and the Softmax layer in the SNN network adopt the same network structure as the max pooling layer and the Softmax in the CNN network, but it is no longer necessary to extract network parameters from the CNN network. Instead, a new idea is proposed to form the output of the corresponding network layer by counting the number of pulses in the output of the previous network layer connected thereto.

[0061] Please refer to Figure 2 In an embodiment of the present invention, an alternative solution is provided. The designed CNN network includes a first convolutional layer, a first max pooling layer, a first fully connected layer, and a first Softmax layer connected in sequence; correspondingly, the SNN network includes a second convolutional layer, a second max pooling layer, a second fully connected layer, and a second Softmax layer connected in sequence.

[0062] Among them, the second convolutional layer and the second fully connected layer respectively adopt the same network structure and network parameters as the first convolutional layer and the first fully connected layer; the second max pooling layer and the second Softmax layer respectively adopt the same network structure as the first max pooling layer and the first Softmax layer.

[0063] It should be noted that in the embodiments of the present invention, the number and arrangement structure of the convolutional layer, the max pooling layer, and the fully connected layer can be set in other forms, but the Softmax layer can only be connected after the last fully connected layer. The number and size of the convolutional kernels can also be set to different values according to actual needs.

[0064] For example, please refer to Figure 3 , the CNN network structure is no longer the first convolutional layer, the first max pooling layer, the first fully connected layer, and the first Softmax layer connected in sequence as shown in FIG. 2. It can be deformed to form a structure as shown in Figure 3 . Specifically:

[0065] The CNN network structure includes multiple groups of convolutional layers and max pooling layers connected in sequence. After the max pooling layer, the first fully connected layer and the first Softmax layer are connected in sequence. It can be seen that Figure 2 the first convolutional layer and the first max pooling layer connected in sequence as shown in can be deformed into multiple groups of convolutional layers and max pooling layers connected in sequence. Among them, Figure 3 in: the number of convolutional kernels in convolutional layers 1-3 is 32 each; the number of convolutional kernels in convolutional layers 4-6 is 64 each; the number of convolutional kernels in convolutional layers 7-9 is 128; the size of the convolutional kernels in convolutional layers 1-9 is (3,3); one max pooling layer is added after every three layers of convolutional layers 1-9; the number of convolutional kernels in convolutional layers 10-11 is 32 and 8 respectively, and the size of the convolutional kernels is (1,1); before inputting into the second fully connected layer, the feature map output by convolutional layer 11 is first unfolded through a Flatten layer and then input into the second fully connected layer; the output result of the second fully connected layer is input into the second Softmax layer.

[0066] Here, the CNN network is trained using the structure shown in Figure 3 . In the SNN network of the embodiments of the present invention: convolutional layers 1-3, convolutional layers 4-6, convolutional layers 7-9, convolutional layers 10-11, and the second fully connected layer adopt the same network structure and network parameters as the corresponding network layers of the trained CNN network; the three max pooling layers and the second Softmax layer adopt the same network structure as the corresponding network layers of the CNN network and form their outputs based on pulse number statistics. Other similar Figure 3 transformed network structures also apply such a corresponding relationship.

[0067] S20. Convert the pixel values of the image to be recognized into a pulse sequence by using the pulse excitation frequency encoding method.

[0068] Specifically, the input of the SNN network is a pulse sequence, and how to accurately convert the pixel values of the image to be recognized into a pulse sequence for the SNN network becomes the key. The existing method is to use the activation values of the convolutional layer in the CNN network to achieve this. However, mixing the convolutional layer of the CNN network into the SNN network will still result in a large number of multiplication calculations in the SNN network, making the power consumption of the SNN network unable to meet the minimum power consumption requirements. To address this existing problem, the embodiment of the present invention proposes a method of converting the pixel values of the image to be recognized into a pulse sequence using the pulse excitation frequency encoding method. During the conversion process, the characteristics of the SNN network structure are utilized. When the number of time steps of the SNN network is set to n, the pixel values of the image to be recognized are converted into a pulse sequence of length n, and then the pixel value p of the image to be recognized i is input into the corresponding pulse neuron i. Please refer to Figure 4 , and the specific pulse sequence conversion process includes:

[0069] Calculate the average value of all pixel values of the image to be recognized, and use the average value as the dynamic excitation threshold for all pixel values;

[0070] Utilize the characteristics of the SNN network structure to input the pixel values of the image to be recognized into the corresponding pulse neurons of the SNN network; for each pulse neuron, it includes:

[0071] For each iteration moment process, it includes: calculating the membrane potential value of the pulse neuron at the current iteration moment according to each pixel value in the image to be recognized; comparing the membrane potential value of the pulse neuron at the current iteration moment with the dynamic excitation threshold, and outputting the pulse value at the current iteration moment according to the comparison result;

[0072] Until the maximum number of iteration moments is reached, output the pulse sequence corresponding to the pulse neuron.

[0073] The embodiment of the present invention first calculates the average value A of all pixel values of the image to be recognized, and then sets the average value A as the dynamic excitation threshold V of each pulse neuron thr , which is expressed by the formula:

[0074]

[0075] where N represents the number of all pixel values of the image to be recognized, and p i represents the gray value of the i-th pixel value.

[0076] The embodiment of the present invention utilizes the characteristics of the SNN network structure to input the pixel value p of the image to be recognized i into the i-th pulse neuron corresponding to the SNN network; for each pulse neuron, where

[0077] When converting the image to be recognized, the length of the pulse sequence for each pixel value conversion is n. Here, the number of iterative moments is designed to be n, and the formula for calculating the membrane potential value of the pulsed neuron at the current iterative moment is expressed as:

[0078] V i (t) = V i (t - 1) + p i (2)

[0079] Among them, V i (t) represents the membrane potential value of the i-th pulsed neuron at the t-th iteration moment, and V i (t - 1) represents the membrane potential value of the i-th pulsed neuron at the (t - 1)-th iteration moment, and p i represents the gray value of the i-th pixel. t takes values from 1 to n, and n is the maximum number of iterative moments.

[0080] It can be seen that V i (t) is the accumulated membrane potential value of the i-th pulsed neuron at the t-th iteration moment. When t = 1, V i (t - 1) = V i (0) = 0, and V i (0) represents the initial value of the membrane potential value during the accumulation process of the i-th pulsed neuron.

[0081] After the neuron fires a pulse, the pixel values in the image to be recognized that exceed the dynamic firing threshold V thr will continuously fire pulses. To prevent this phenomenon, the embodiment of the present invention proposes to reset the membrane potential value after the neuron fires a pulse. Please refer to Figure 4 again. Specifically:

[0082] When the membrane potential value of the pulsed neuron at the current iterative moment is greater than the dynamic firing threshold, reset the membrane potential value of the pulsed neuron at the current iterative moment; recalculate the membrane potential value of the pulsed neuron at the next iterative moment according to the reset membrane potential value of the pulsed neuron and each pixel value in the image to be recognized; make a comparison according to the membrane potential value of the pulsed neuron at the next iterative moment;

[0083] Until the maximum number of iterative moments is reached, output the pulse sequence corresponding to this pulsed neuron.

[0084] For the situation where the membrane potential value of the pulsed neuron at the current iterative moment in the embodiment of the present invention is greater than the dynamic firing threshold, it is necessary to reset the membrane potential value. This reset operation is achieved through the calculation of subtracting a constant from the membrane potential value. This constant is set to the theoretical maximum gray value, that is, 255. The formula for resetting the membrane potential value of the pulsed neuron at the current iterative moment is expressed as:

[0085] V i '(t) = V i(t) - 255 (3)

[0086] Among them, V i '(t) represents the membrane potential value of the i-th spiking neuron after reset at the t-th iteration moment, and V i (t) represents the membrane potential value of the i-th spiking neuron at the t-th iteration moment.

[0087] Correspondingly, the formula for recalculating the membrane potential value of the spiking neuron at the next iteration moment is expressed as:

[0088] V i (t + 1) = V i '(t) + p i (4)

[0089] Among them, V i (t + 1) represents the membrane potential value of the i-th spiking neuron at the (t + 1)-th iteration moment.

[0090] It can be seen that by iterating n times cyclically in time as described above, each pixel value of the image to be recognized can be converted into a pulse sequence, realizing the pulse coding process of the pixel value.

[0091] S30. Use the SNN network to recognize the pulse sequence to obtain the image recognition result; among them, the convolutional layer and the fully connected layer in the SNN network adopt the same network structure and network parameters as the convolutional layer and the fully connected layer in the trained CNN network, and obtain the output of the corresponding network layer based on the network structure and network parameters; the max-pooling layer and the Softmax layer in the SNN network adopt the same network structure as the max-pooling layer and the Softmax layer in the trained CNN network, and use the pulse number statistics method to realize the output corresponding to the max-pooling layer and the Softmax layer in the SNN network, and take the output of the Softmax layer as the image recognition result.

[0092] Specifically, the embodiment of the present invention uses the SNN network to obtain the image recognition result. During the recognition process of the SNN network, due to the complex logical operations in the max-pooling layer and the Softmax layer in the existing SNN network, the power consumption is still relatively large. To solve this problem, the embodiment of the present invention proposes that in the max-pooling layer and the Softmax layer, the output is formed by counting the number of pulses corresponding to the pulse sequence output by the previous network layer connected to it, while the other network layers in the SNN network adopt the same network combination and network parameters as the CNN network to form their outputs. Specifically:

[0093] From Figure 2 it can be known that the SNN network includes a second convolutional layer, a second max-pooling layer, a second fully connected layer, and a second Softmax layer connected in sequence, among which,

[0094] In the second convolutional layer of the embodiment of the present invention, the same network structure and network parameters as those of the first convolutional layer extracted through S10 are adopted. The pulse sequence converted from the image to be recognized is input into the neurons of the second convolutional layer of the SNN network in chronological order in the sequence length direction. The neurons perform cumulative excitation calculations on the input convolution kernels, weight values, and pulse values. The number of calculation cycles is equal to the number of time steps n, and a pulse sequence with a length of n is output. Multiple pulse sequences form the output feature map of this second convolutional layer. This output feature map is a three-dimensional array with a size of (x, y, n), where x and y respectively represent the height and width of the output feature map, and n represents the length of the pulse sequence.

[0095] It should be noted that the neuron excitation threshold V in the second convolutional layer of the SNN network th is different from the dynamic excitation threshold V in the pulse excitation frequency encoding method thr In the second convolutional layer, the neuron excitation threshold V th adopts a fixed value, which is set to 0.06 in the embodiment of the present invention.

[0096] Furthermore, the present invention implements the process of the output of the second max-pooling layer in the SNN network by using the pulse quantity statistics method. Please refer to Figure 5 , including:

[0097] Collect the pulse sequences output by the second convolutional layer according to the sampling kernel size;

[0098] For the pulse sequences of each sampling kernel, including: counting the number of pulses in each pulse sequence in this sampling kernel; finding the maximum value of the number of pulses from all the counted numbers of pulses corresponding to this sampling kernel;

[0099] Extracting the address of the pulse sequence corresponding to the maximum number of pulses, and retrieving the pulse sequence of this sampling kernel according to the extracted address;

[0100] The pulse sequences of all sampling kernels form the pulse sequence output by the second max-pooling layer.

[0101] Assume that the sampling kernel size is (w, h), and the pulse sequence (feature map) output by the second convolutional layer is sampled by this sampling kernel (w, h). Then, each sampling kernel needs to count the number of pulses in w×h pulse sequences. The number of pulses in each pulse sequence output by the second convolutional layer is counted through an accumulation operation. This process can be expressed as:

[0102]

[0103] where M j represents the number of pulses of the counted pulse sequence S j , and S jDenote the j-th pulse sequence collected by the acquisition core, where j ∈ [1, J], and here J takes the value of w × h. Denote the pulse sequence S j The t-th pulse value in it, and n represents the length of the pulse sequence.

[0104] Then, compare the magnitudes of all the counted pulse quantities. If the counted pulse quantity M j is the maximum value, then set M j as the maximum pulse quantity M max It is expressed as:

[0105] M max = max(M j ) (6)

[0106] Extract the address of the pulse sequence corresponding to the maximum pulse quantity M max It is expressed as:

[0107] D = addr[find(S j :if(M j = M max ))] (7)

[0108] Among them, addr represents the extraction address, and find(S j :if(M j = M max )) represents finding the pulse sequence S corresponding to the maximum pulse quantity M max being M j at that time. Retrieve the pulse sequence of the sampling core according to the extracted address. j .

[0109] Complete the statistics of the maximum pulse quantity for all sampling cores, extract the address of the pulse sequence corresponding to the maximum pulse quantity, and the pulse sequence retrieved at this address is used as the pulse sequence of this sampling core. The pulse sequences corresponding to all sampling cores form the feature map output by the second max-pooling layer.

[0110] Furthermore, before inputting the feature map output by the second max-pooling layer into the fully connected layer, first expand the first two dimensions of the feature map output by the second max-pooling layer into 1D through a Flatten layer, and convert the feature map output by the second max-pooling layer into a 2D pulse sequence feature map. For the specific expansion method, refer to the CNN network.

[0111] In the embodiment of the present invention, the second fully connected layer adopts the same network structure and network parameters as the first fully connected layer, and inputs the 2D pulse sequence feature map expanded by the Flatten layer into the neurons of the second fully connected layer in chronological order in the sequence length direction. If the number of neurons in the second fully connected layer is K, then K pulse sequences are output, and the addresses of these K pulse sequences respectively correspond to all possible k categories of images.

[0112] Further, the embodiment of the present invention realizes the output process of the second Softmax layer in the SNN network by using the pulse number statistics method. Please refer to Figure 6 , including:

[0113] Statistical pulse number of the pulse sequence output by the second fully connected layer;

[0114] Find the maximum pulse number from all the statistically counted pulse numbers;

[0115] Extract the address of the pulse sequence corresponding to the maximum pulse number, and retrieve and form the output of the second Softmax layer according to the extracted address, and this output is used as the image recognition result.

[0116] It can be seen that the data output processing of the second Softmax layer is similar to that of the second max-pooling layer. The difference is that the second max-pooling layer divides the feature map (pulse sequence) input by the second convolutional layer according to the sampling kernel size, performs the maximum pulse number statistics on the pulse sequences of all sampling kernel sizes after division, extracts the address of the pulse sequence corresponding to the maximum statistically counted pulse number, and forms the feature output map of the second max-pooling layer according to the pulse sequence retrieved corresponding to the extracted address. While the second Softmax layer directly performs the maximum pulse number statistics on the feature map (pulse sequence) output by the second fully connected layer, extracts the address of the pulse sequence corresponding to the maximum statistically counted pulse number, and determines the image recognition result according to the extracted address. Among them, in formulas (5) to (7), the value of J is K, and K is the number of pulse sequences output by the second fully connected layer.

[0117] Here, the reason why the address extracted in the second Softmax layer can be used as the image recognition result is that each address corresponding to an image category is preset in the second Softmax layer. During the training process of the CNN network, each address of the first Softmax function layer corresponds to an image category, and the second Softmax layer adopts the same network structure as the first Softmax layer. Coupled with the data categories stored at each address known in advance, the image recognition result of the image to be recognized can finally be determined through the address extracted by the second Softmax layer.

[0118] To verify the effectiveness of the neural network image recognition method based on pulse statistics provided by the embodiments of the present invention, the following experiments are carried out for verification.

[0119] 1. Experimental simulation parameters

[0120] The image recognition method provided by the embodiments of the present invention is compared with the existing SNN network image recognition method and the CNN network image recognition method. The image recognition data set used is the CIFAR10 data set. During the experiment, Figure 3 the CNN network and the SNN network shown are used. First, the CNN network is trained through S10, and the network parameters of convolutional layers 1-3, convolutional layers 4-6, convolutional layers 7-9, convolutional layers 10-11, and the second fully connected layer are extracted. Then, according to the steps of S20-S30, the image to be recognized is input into the SNN network for image recognition to obtain the image recognition result. Among them, Figure 3 the three maximum pooling layers and the second Softmax layer in all form their outputs based on pulse number statistics. The specific processing process of the three maximum pooling layers can be seen in Figure 5 and the specific processing process of the second Softmax layer can be seen in Figure 6 .

[0121] During the experiment, the time step number n and the firing threshold V of the image recognition method provided by the embodiments of the present invention and the existing SNN network th are set to be the same. The images to be recognized used are all from the CIFAR10 test set. The image recognition accuracies of the image recognition method provided by the embodiments of the present invention, the existing SNN network, and the CNN network are shown in Table 1.

[0122] Table 1 Accuracy comparison

[0123] Network model name The SNN network of the present invention Existing SNN network CNN network Accuracy rate (%) 86.17 86.36 87.33

[0124] As can be seen from Table 1: The image recognition accuracies of the image recognition method provided by the embodiments of the present invention and the existing SNN network are both close to the image recognition accuracy of the CNN network. Among them, the image recognition accuracy of the embodiments of the present invention is slightly lower than that of the existing SNN network, but the gap between the two is very small. This gap is because the image recognition method provided by the present invention uses the pulse firing frequency coding method to reduce the computational complexity and thus reduces the image recognition accuracy, while all the maximum pooling layers and the second Softmax layer based on pulse number statistics increase the recognition accuracy. Finally, the accuracy gap between the image recognition method provided by the embodiments of the present invention and the existing SNN network is very small.

[0125] The comparison of the computational amounts of the image recognition methods of the image recognition method provided by the embodiments of the present invention and the existing SNN network during the image coding process is shown in Table 2.

[0126] Table 2 Comparison of Computational Amounts in the Image Encoding Process

[0127] Network model name The SNN network of the present invention Existing SNN network Number of multiplications 1 885000 Number of additions 617000 852000

[0128] As can be seen from Table 2, in the image recognition method provided by the embodiments of the present invention, there is only 1 multiplication calculation in the process of calculating the average value of image pixel values during the encoding process of the image to be recognized in the SNN network adopted. However, in the image recognition method of the existing SNN network, since the convolutional layer of the CNN network is used, the amount of multiplication calculation is very large. The amount of addition calculation in the image encoding process of the image recognition method provided by the embodiments of the present invention is reduced by about 30% compared with the image recognition method of the existing SNN network. In the image recognition method provided by the embodiments of the present invention, the number, structure, and neuron calculation processes of the convolutional layer and the fully connected layer are the same, so the computational amounts are equal. However, in the encoding process of the image to be recognized, converting the pulse sequence by using the pulse firing frequency encoding method and forming the outputs of the max-pooling layer and the Softmax layer based on the pulse number statistics both greatly reduce the computational amount. Therefore, compared with the image recognition method of the existing SNN network, the overall computational amount of the image recognition method proposed by the embodiments of the present invention is further reduced.

[0129] In summary, the neural network image recognition method based on pulse statistics proposed by the embodiments of the present invention, on the basis of ensuring the recognition accuracy of the image to be recognized, in the encoding process of the image to be recognized, by using the pulse firing frequency encoding method to convert the pulse sequence and forming the outputs of the max-pooling layer and the Softmax layer based on the pulse number statistics, it avoids the introduction of the CNN convolutional layer and simplifies the operation logic of the max-pooling layer and the Softmax layer in the SNN network. By using the accumulative calculation of the adder instead of the complex logic circuit through the pulse data statistics method, the overall computational amount is further reduced compared with the existing SNN network. And in the FPGA implementation, due to reducing the calls of multipliers and adders, the consumed FPGA hardware resources are reduced, the on-chip power consumption of the FPGA is reduced, which is more conducive to the deployment of the SNN network in the FPGA, and thus has a wider application prospect, such as meeting the image recognition requirements of existing portable devices with lower power consumption.

[0130] Based on the same inventive concept, please refer to Figure 7 , the embodiments of the present invention provide an electronic device, including a processor 701, a communication interface 702, a memory 703, and a communication bus 704. Among them, the processor 701, the communication interface 702, and the memory 703 complete mutual communication through the communication bus 704;

[0131] The memory 703 is used to store a computer program;

[0132] A processor 701, when executing a program stored in a memory 703, implements the steps of the above-mentioned neural network image recognition method based on pulse statistics.

[0133] An embodiment of the present invention provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned neural network image recognition method based on pulse statistics are implemented.

[0134] For the device / electronic device / storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiments.

[0135] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more, unless otherwise specifically defined.

[0136] Although the present application is described herein in conjunction with various embodiments, however, in the process of implementing the claimed present application, those skilled in the art can understand and implement other variations of the disclosed embodiments by viewing the accompanying drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit may implement several functions recited in the claims. Certain measures are recited in mutually different dependent claims, but this does not mean that these measures cannot be combined to produce good results.

[0137] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A neural network image recognition method based on pulse statistics, characterized in that, Including: Training a CNN network using training image data to obtain a trained CNN network, and extracting the network parameters corresponding to the convolutional layer and the fully connected layer in the trained CNN network; Converting the pixel values of the image to be recognized into a pulse sequence using the pulse excitation frequency encoding method; Using an SNN network to recognize the pulse sequence to obtain an image recognition result; wherein, the convolutional layer and the fully connected layer in the SNN network adopt the same network structure and network parameters as the convolutional layer and the fully connected layer in the trained CNN network, and the output of the corresponding network layer is obtained based on the network structure and network parameters; the max pooling layer and the Softmax layer in the SNN network adopt the same network structure as the max pooling layer and the Softmax layer in the trained CNN network, and the output corresponding to the max pooling layer and the Softmax layer in the SNN network is realized by using the pulse number statistics method, and the output of the Softmax layer is used as the image recognition result.

2. The neural network image recognition method based on pulse statistics according to claim 1, wherein The converting the pixel values of the image to be recognized into a pulse sequence using the pulse excitation frequency encoding method includes: Calculating the average value of all pixel values of the image to be recognized, and using the average value as the dynamic excitation threshold of all pixel values; Using the characteristics of the SNN network structure, inputting the pixel values of the image to be recognized into the pulse neurons corresponding to the SNN network; for each pulse neuron, including: For each iteration process, including: calculating the membrane potential value of the pulse neuron at the current iteration moment according to each pixel value in the image to be recognized; comparing the membrane potential value of the pulse neuron at the current iteration moment with the dynamic excitation threshold, and outputting the pulse value at the current iteration moment according to the comparison result; Until the maximum number of iteration moments is reached, outputting the pulse sequence corresponding to the pulse neuron.

3. The neural network image recognition method based on pulse statistics according to claim 2, characterized in that The formula for calculating the membrane potential value of the pulse neuron at the current iteration moment is expressed as: V i V(t) = i V(t - 1)+p i ; Among them, V i (t) represents the membrane potential value of the i-th spiking neuron at the t-th iteration moment, V i (t - 1) represents the membrane potential value of the i-th spiking neuron at the (t - 1)-th iteration moment, p i represents the gray value of the i-th pixel, and t takes values from 1 to n, where n is the maximum number of iteration moments.

4. The neural network image recognition method based on pulse statistics according to claim 3, characterized in that The converting the pixel values of the image to be recognized into a pulse sequence using the pulse excitation frequency encoding method further includes: When the membrane potential value of the pulse neuron at the current iteration moment is greater than the dynamic excitation threshold, resetting the membrane potential value of the pulse neuron at the current iteration moment; recalculating the membrane potential value of the pulse neuron at the next iteration moment according to the reset membrane potential value of the pulse neuron and each pixel value in the image to be recognized; comparing according to the membrane potential value of the pulse neuron at the next iteration moment; Until the maximum number of iteration moments is reached, outputting the pulse sequence corresponding to the pulse neuron.

5. The neural network image recognition method based on pulse statistics according to claim 4, wherein The formula for resetting the membrane potential value of the pulse neuron at the current iteration moment is expressed as: V i '(t) = V i (t) - 255; Among them, V i '(t) represents the membrane potential value of the i-th spiking neuron after reset at the t-th iteration moment, and V i (t) represents the membrane potential value of the i-th spiking neuron at the t-th iteration moment; Correspondingly, the formula for recalculating the membrane potential value of the pulse neuron at the next iteration moment is expressed as: V i (t + 1)=V i '(t)+p i ; Among them, V i (t + 1) represents the membrane potential value of the i-th spiking neuron at the (t + 1)-th iteration time.

6. The method for neural network image recognition based on pulse statistics according to claim 1, wherein The CNN network includes a first convolutional layer, a first max pooling layer, a first fully connected layer, and a first Softmax layer connected in sequence; Correspondingly, the SNN network includes a second convolutional layer, a second max pooling layer, a second fully connected layer, and a second Softmax layer connected in sequence; Among them, the second convolutional layer and the second fully-connected layer respectively adopt the same network structure and network parameters as the first convolutional layer and the first fully-connected layer, and output pulse sequences corresponding to the network layers based on the network structure and network parameters; The second max-pooling layer and the second Softmax layer respectively adopt the same network structure as the first max-pooling layer and the first Softmax layer, and the outputs of their corresponding network layers are respectively formed by counting the number of pulses output by the second convolutional layer and the number of pulses output by the second fully-connected layer.

7. The neural network image recognition method based on pulse statistics according to claim 6, characterized in that, The process of realizing the output of the second max-pooling layer in the SNN network by using the pulse number counting method includes: Collecting the pulse sequence output by the second convolutional layer according to the sampling kernel size; For each pulse sequence of each sampling kernel, including: counting the number of pulses in each pulse sequence in the sampling kernel; finding the maximum value of the number of pulses from all the counted numbers of pulses corresponding to the sampling kernel; extracting the address of the pulse sequence corresponding to the maximum value of the number of pulses, and retrieving the pulse sequence of the sampling kernel according to the extracted address; The pulse sequences of all sampling kernels form the pulse sequence output by the second max-pooling layer.

8. The neural network image recognition method based on pulse statistics according to claim 7, characterized in that The process of realizing the output of the second Softmax layer in the SNN network by using the pulse number counting method includes: Counting the number of pulses in the pulse sequence output by the second fully-connected layer; Finding the maximum value of the number of pulses from all the counted numbers of pulses; Extracting the address of the pulse sequence corresponding to the maximum value of the number of pulses, and retrieving to form the output of the second Softmax layer according to the address, and using the output as the image recognition result.

9. The neural network image recognition method based on pulse statistics according to claim 8, characterized in that The address of the pulse sequence corresponding to the maximum value of the number of pulses is extracted and represented as: D = addr[find(S j : if(M j = M max ))], j ∈ [1, J]; Among them, addr represents the extraction address, J represents the maximum value of j. For the output formation of the second max-pooling layer, J takes the value of w×h, where (w, h) represents the sampling kernel size. For the output formation of the second Softmax layer, J takes the value of K, and K represents the number of pulse sequences output by the second fully connected layer. find(S j :if(M j =M max )) represents finding the maximum value M max of M j and the corresponding pulse sequence S j when it is M max =max(M j ), j∈[1, J], and M j represents the number of pulses in the statistical pulse sequence S j . represents the t-th pulse value in the pulse sequence S j , and n represents the length of the pulse sequence.

10. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; The memory is used to store computer programs; When the processor is used to execute the program stored on the memory, it realizes the steps of the neural network image recognition method based on pulse statistics according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Image classification method based on multi-layer spring convolutional neural network

    CN110119785A

  • Method for learning and recognizing image pulse data space-time information based on Spike cube SNN

    CN110210563A