An Adaptive Convolutional Auto-Encoding Method Based on Spiking Neurons
Through the adaptive convolution automatic coding method, the best encoding time window is automatically found and the parameters are optimized, which solves the problem of encoding time window preset in existing neural coding methods, and improves the accuracy and efficiency of image classification, especially the performance on MNIST and CIFAR10 data sets has been significantly improved.
Patent Information
- Application Number
- CN202211407403.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-09
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2042-11-09
AI Technical Summary
The existing neural coding methods require pre-set encoding time windows and lack effective learning processes, and cannot make full use of time and spatial dimension information for efficient encoding, making it difficult to meet the needs of image classification tasks.
Adaptive convolutional automatic encoding method based on pulse neurons is adopted to automatically find the best encoding time window through adaptive convolutional pulse coding, pulse pixel value mapping decoding and deep convolutional decoding networks, and optimize the encoding parameters to realize pre-training of adaptive convolutional pulse coding.
The accuracy of image classification tasks is improved, especially the accuracy of the static MNIST dataset reaches 99.70%, the accuracy of the static CIFAR10 dataset reaches 91.64%, and the network convergence rate is reduced, improving the efficiency and effect of image classification.
Smart Images

Figure CN115620026B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a pulse coding technology in the field of pulse neural networks, and in particular to an adaptive convolution automatic coding method based on pulse neurons. Background Art
[0002] A core question about sensory systems in systems neuroscience is how neurons represent the input-output relationship between stimuli and their pulses. This problem is expressed as neural coding, which consists of two basic parts: encoding and decoding. The process of information conversion is called pulse coding. After years of development, researchers have proposed a variety of biologically inspired pulse coding schemes for different modalities and different problems. However, there are some problems in terms of information representation efficiency and noise robustness. In particular, the pulse coding process often lacks an effective learning process, making it difficult to fully utilize the multi-dimensionality of pulses for collaborative information carrying. The separation from the back-end Spiking Neural Networks (SNNs) learning algorithm module also makes it difficult to make targeted learning adjustments according to task objectives.
[0003] The encoding time window is the first important time scale for characterizing neural coding. It is defined as the window that encompasses a specific response pattern. A longer encoding time window is not necessarily better. When the network structure is deep, a simulation time that is too short cannot fully represent the information, while a simulation time that is too long will lead to unacceptable computational resource consumption. Currently, there is still no scientifically sound and effective way to predefine the encoding time window in existing neural coding methods.
[0004] The encoding and decoding of spiking neurons can be viewed as the conversion of stimuli and spikes back and forth in different directions. Recently, a new approximate Bayesian method was proposed by Nikhil et al., which uses nonlinear deep neural networks (DNNs) to decode natural images of spike activity from retinal ganglion cell (RGC) populations. This method outperforms the linear reconstruction techniques commonly used to explain neural responses to high-dimensional stimuli. Based on this, Zhang et al. proposed a new DNNs-based decoding framework, Spike-Image Decoder (SID), for reconstructing visual scenes, including static images and dynamic videos obtained from spikes of RGCs populations recorded in experiments. However, both of the above methods only focus on the decoding stage, while the encoding stage of the spikes has not been substantially changed.
[0005] In general, traditional neural coding methods cannot fully utilize time and space dimension information for efficient encoding and representation of information because they require pre-setting of the coding time window and lack an effective learning process. Summary of the Invention
[0006] The purpose of the present invention is to provide an adaptive convolutional auto-encoding method based on spiking neurons, which can effectively improve the performance of spiking neural network image classification tasks.
[0007] To achieve the above object, the present invention proposes the following technical solutions:
[0008] An adaptive convolutional automatic encoding method based on spiking neurons comprises the following steps:
[0009] Step 1: Drawing on the coding mode of photoelectric signal conversion in the biological retina system, and considering that the real value of the static image already preserves the color information related to the photoelectric signal, a fast, efficient and accurate adaptive convolutional pulse coding is proposed. The adaptive convolutional pulse coding includes a first convolutional layer, batch normalization, and spiking neurons connected in sequence. The original image is extracted from the first convolutional layer and the data is batch normalized to obtain the feature map of the original image. The feature map of the original image is then input into the spiking neurons to emit pulses according to the preset coding time window range, resulting in the original image pulse coding sequence for each preset coding time window.
[0010] Step 2: Decode the original image pulse code sequence of each preset coding time window described in step 1 by using the pulse pixel value mapping decoding method to obtain a decoded image, compare the similarity between the decoded image and the original image when the coding time window changes, and automatically find the optimal coding time window for adaptive convolutional pulse coding to process different data sets;
[0011] Step 3: Based on the optimal coding time window described in step 2, the original image is re-input into the adaptive convolution pulse coding to obtain the pulse coding sequence of the original image, and then the pulse coding sequence of the original image is input into the deep convolution decoding network for image reconstruction to obtain the reconstructed image. The adaptive convolution pulse coding learning parameters are optimized according to the error backpropagation between the reconstructed image and the original image to complete the adaptive convolution pulse coding pre-training process.
[0012] Optionally, in step 1, a fast, efficient and accurate adaptive convolutional pulse coding is proposed, specifically including:
[0013] (1) Drawing on the coding mode of the biological retinal system, which consists of a three-layer structure of photoreceptors, bipolar cells, and ganglion cells for photoelectric signal conversion;
[0014] (2) Using standard multi-channel sliding convolution in deep learning to simulate the linear filter operation of bipolar cells, and using spiking neurons to simulate the spontaneous, nonlinear impulse response of ganglion cells to the filtered signal, to generate a pulse code sequence with temporal properties;
[0015] (3) Adaptive convolutional pulse coding consists of the first convolutional layer, batch normalization, and pulse neurons connected in sequence.
[0016] Optionally, in step 1, the feature map of the original image is input into the spiking neuron to emit pulses according to a preset coding time window range, so as to obtain the original image pulse coding sequence of each preset coding time window, specifically including:
[0017] (1.1) Preprocess the original image;
[0018] (1.2) Input the preprocessed original image into the first convolutional layer to extract features, and then perform data batch normalization to obtain the feature map of the preprocessed original image;
[0019] (1.3) The feature map of the preprocessed original image is input into the pulse neuron as an electrical signal, simulating the retina to capture external information and generate action potentials, and generating the original image pulse coding sequence for each pre-set coding time window according to the pre-set coding time window range.
[0020] Optionally, in step (1.3), the spiking neuron model is a leaky integrate-spark model, and its dynamic equation is:
[0021]
[0022]
[0023]
[0024] Where V(t) is the membrane voltage of the neuron at time t, I(t) is the external input current of the neuron at time t, and τ m =R·C is the membrane time decay constant, R and C are the membrane resistance and membrane capacitance of the neuron respectively, ω j is the synaptic weight of the j-th input neuron, is the jth input neuron at T w The arrival time of the kth presynaptic spike within the integration time window, K(·) is the kernel function that describes the time decay effect of the synaptic current, V0 is the normalization factor that makes the kernel function maximum value 1, τ s is the time decay constant of the synaptic current, and H(·) is the Heaviside step function, which only calculates the effect of the pulse input before time t on the postsynaptic current.
[0025] Optionally, in step 2, decoding the original image pulse code sequence of each preset coding time window in step 1 by a pulse pixel value mapping decoding method to obtain a decoded image, comparing the similarity between the decoded image and the original image when the coding time window changes, and automatically finding the optimal coding time window for adaptive convolutional pulse coding to process different data sets, specifically including:
[0026] (2.1) Calculate the pulse emission frequency of each neuron i according to the original image pulse coding sequence of each pre-set coding time window in step 1 Where T is the encoding time window, and n is the number of pulses emitted by each neuron within the encoding time window T;
[0027] (2.2) According to the pulse emission frequency of each neuron i Mapped to the corresponding image pixel value in is the maximum pulse firing frequency of all neurons, which is decoded to obtain the decoded image;
[0028] (2.3) According to the set coding time window range, the similarity between the decoded image and the original image in each coding time window is calculated in turn, and the image similarity of different coding time windows is compared to automatically find the optimal coding time window for adaptive convolutional pulse coding to process the current data set.
[0029] Optionally, in step 3, inputting the pulse code sequence of the original image into a deep convolutional decoding network for image reconstruction to obtain a reconstructed image specifically includes:
[0030] (3.1) The network structure of the deep convolutional decoding network consists of 5 dense blocks and 1 convolution layer. The first to fifth layers are dense blocks, and the sixth layer is a convolution layer. Each dense block consists of 2 convolution layers. The two convolution layers in each dense block use 128 1×1 filters with a stride of 1 and 32 3×3 filters with a stride of 1.
[0031] (3.2) The input of the deep convolutional decoding network is a binary pulse matrix of the pulse code sequence of the original image batch×C1×width×height, where batch is the number of batches input to the deep convolutional decoding network at one time, C1 is the number of output channels of the first convolutional layer, and width and height are the width and height of the original image respectively;
[0032] The first layer is a Dense block. The network input first performs a round of batch normalization, ReLU activation function and convolution operation in sequence, and then repeats the above operation to output a real-valued matrix of batch×32×width×height. The output of this layer is then concatenated with the input of the current layer to generate a real-valued matrix of batch×(C1+32)×width×height as the input of the next layer;
[0033] (3.3) From the second to the fifth layer, each layer performs a Dense block operation in sequence. The output of the second layer is a real-valued matrix of batch×(C1+64)×width×height, the output of the third layer is a real-valued matrix of batch×(C1+96)×width×height, the output of the fourth layer is a real-valued matrix of batch×(C1+128)×width×height, and the output of the fifth layer is a real-valued matrix of batch×32×width×height.
[0034] (3.4) The sixth layer is a convolutional layer, which uses C2 3×3 filters with a step size of 1, where C2 is the number of channels in the original image. The input is the real-valued matrix of batch×32×width×height output by the fifth layer, and the output is a real-valued matrix of batch×C2×width×height;
[0035] (3.5) Finally, the batch×C2×width×height real-valued matrix output by the sixth layer undergoes a Tanh activation function to generate a batch×C2×width×height real-valued matrix, i.e., the reconstructed image.
[0036] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0037] Unlike traditional neural coding methods that require pre-setting encoding time windows and lack an effective learning process, making it impossible to fully utilize time and space dimensional information for efficient encoding, the adaptive convolutional auto-encoding method based on spiking neurons proposed in the present invention can automatically find the optimal encoding time window for processing different data sets. It can also continuously optimize the adaptive convolutional pulse coding parameters during network training to better characterize the captured external stimuli and maximize the retention of original image information. The spiking convolutional neural network constructed based on the adaptive convolutional auto-encoding method proposed in the present invention can achieve an accuracy of 99.70% for classifying the static MNIST dataset and 91.64% for the static CIFAR10 dataset. In addition, when the adaptive convolutional pulse coding pre-training parameters are obtained, the network convergence rate can be greatly reduced, enabling image classification tasks to reach a fitting state more quickly and improving the accuracy of the static MNIST and static CIFAR10 datasets to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0039] Figure 1 A schematic diagram of the process of an adaptive convolutional automatic encoding method based on spiking neurons of the present invention;
[0040] Figure 2 This is a structural diagram of an adaptive convolutional pulse coding model of the present invention that is high-precision and closer to the biological retina;
[0041] Figure 3 This is a flow chart of a pulse pixel value mapping and decoding method of the present invention. DETAILED DESCRIPTION
[0042] Example 1
[0043] The existing neural coding technology needs to pre-set the coding time window and lacks an effective learning process. The embodiment of the present invention provides an adaptive convolution automatic coding method based on pulse neurons, which specifically includes: adaptive convolution pulse coding, pulse pixel value mapping decoding method, and deep convolution decoding network. First, the pulse pixel value mapping decoding method is used to automatically find the optimal coding time window for adaptive convolution pulse coding to process different data sets; next, a deep convolution decoding network is designed to adjust the learning parameters by reconstruction error to realize the adaptive convolution pulse coding pre-training process; finally, the adaptive convolution pulse coding pre-training parameters are naturally obtained, and the adaptive convolution pulse coding is combined with the back-end deep convolutional pulse neural network (Spiking Convolutional Neural Networks, SCNNs) for image classification tasks.
[0044] like Figure 1 As shown, an embodiment of the present invention provides an adaptive convolutional automatic encoding method based on pulse neurons, comprising the following steps:
[0045] Step 1: Based on the coding mode of photoelectric signal conversion in biological retina system, considering that the real value of static image has preserved the color information related to photoelectric signal, a fast, efficient and accurate adaptive convolution pulse coding is proposed, such as Figure 2 shown.
[0046] The embodiment of the present invention utilizes standard multi-channel sliding convolution in deep learning to simulate the linear filter operation of bipolar cells, and uses leaky integrate-and-fire (LIF) neurons to simulate the spontaneous, nonlinear impulse response of ganglion cells to the post-filter signal.
[0047] The embodiment of the present invention converts the LIF neuron dynamics equation into a discrete-time iterative form, and the membrane voltage update rule is as follows:
[0048]
[0049] Among them, V(t) is the membrane voltage of the neuron at time t, I(t) is the external input current of the neuron at time t, is the membrane voltage decay of the neuron per unit time dt. Simplified by the attenuation factor λ, the external input current I(t) is expanded into the weighted sum of the input pulses ∑ j W j s j (t), W j is the synaptic weight of the j-th input neuron, s j(t) is the pulse emission state of the j-th input neuron at time t, represented by a binary value (0 or 1), The decay effect of is fused to the synaptic weight W j Finally, the firing-reset mechanism of LIF neurons is added to the membrane voltage update rule, and the variables are vectorized to finally obtain the discrete-time iterative LIF model.
[0050] V l [t] = λ l (1-S l [t-1])V l [t-1]+W l S l-1 [t]
[0051]
[0052] Among them, l and t are the states of the current l-layer neurons at time t, V l [t] is the neuron's membrane voltage, λ l is the attenuation factor W l is the synaptic weight matrix connecting the presynaptic neuron and the postsynaptic neuron, Indicates that the release threshold B l Regulated normalized membrane voltage, S l [t]∈{0,1} is the pulse emission state controlled by the Heaviside step function H(·), when the normalized membrane voltage When it is greater than or equal to 0, the neuron emits a pulse (S l [t]=1); otherwise it means there is no pulse output (S l [t]=0).
[0053] In this embodiment of the present invention, based on the proposed adaptive convolutional pulse coding, pulse coding is performed on 60,000 training set images of the static MNIST dataset and 50,000 training set images of the static CIFAR10 dataset. The specific steps are as follows:
[0054] (1.1) Preprocess the training set images of the classification dataset. For the static MNIST dataset, first rotate them 30° clockwise, then shift them horizontally and vertically by 0.15× the image aspect ratio, then randomly scale them with a scaling factor in the range of (0.85, 1.11), and finally normalize the pixel values to [0, 1]. For the static CIFAR10 dataset, first randomly crop the images to 32×32 size with 6 steps of padding, then randomly flip them horizontally with probability p=0.5, and finally normalize the pixel values to [0, 1].
[0055] (1.2) The preprocessed training set images are input into the first convolutional layer to extract features and then batch normalize the data to obtain feature maps of the preprocessed training set images. The first convolutional layer uses 128 3×3 filters with a step size of 1.
[0056] (1.3) The feature map of the preprocessed training set image is input into the LIF neuron as an electrical signal, simulating the retina capturing external information to generate action potentials, and generating a pulse coding sequence of the preprocessed training set image for each pre-set coding time window according to the pre-set coding time window range.
[0057] Step 2: Based on the pulse pixel value mapping decoding method, the optimal encoding time window T = 11 for adaptive convolutional pulse coding of the static MNIST dataset (60,000 training set images) is automatically found, with a proportion of up to 99%. The optimal encoding time window T = 12 for the static CIFAR10 dataset (50,000 training set images) is found, with a proportion of over 84%.
[0058] The embodiment of the present invention assumes that a coding time window of varying size can represent image feature information. To balance image information representation and computing resource consumption, the embodiment of the present invention sets the size of the coding time window to be in the range of [10-20].
[0059] When the encoding time window size of the adaptive convolution pulse coding changes, the pulse pixel value mapping decoding method is used to decode the pre-processed training set image pulse coding sequence of each preset encoding time window in step 1 to obtain a decoded image, and the similarity between the decoded image and the original image is compared to automatically find the optimal encoding time window for adaptive convolution pulse coding to process different data sets. The pulse pixel value mapping decoding method process is as follows: Figure 3 shown.
[0060] The pulse-pixel value mapping decoding method uses as input a binary pulse matrix generated by LIF neurons using adaptive convolutional pulse coding. Each element of the matrix represents the number of pulses fired by the LIF neurons within a set encoding time window. The binary pulse matrix is preprocessed and normalized, and a mapping relationship is established between the LIF neuron's pulse firing frequency and pixel value. Decoding generates an n×n decoded image, where n represents the image width and height. Finally, the decoded image is evaluated using three typical image quality evaluation metrics: mean square error (MSE), peak signal-to-noise ratio (PSNR), and structural similarity (SSIM). The encoding time window varies from 10 to 20, meaning each image is encoded and decoded 11 times. This embodiment automatically compares the 11 decoded images of each image in the static MNIST and static CIFAR10 datasets with the original image using the three image quality evaluation metrics. The encoding time window T corresponding to the optimal MSE, PSNR, and SSIM values among the 11 times is output, and the proportion of each T in the corresponding dataset is then calculated. In the static MNIST training dataset, it can be found that T=11 is the optimal value of the encoding time window, with a proportion of up to 99%. In the static CIFAR10 dataset, T=12, the proportion exceeds 84%.
[0061] Step 3: Design a deep convolutional decoding network for image reconstruction and complete the adaptive convolutional pulse coding pre-training process.
[0062] After determining the optimal encoding time window for adaptive convolutional pulse coding to process the static MNIST and static CIFAR10 datasets in step 2, an embodiment of the present invention designs a deep convolutional decoding network, in which a network extended from the DenseNet model is used for the pulse-image reconstruction task for the first time.
[0063] The network structure of the deep convolutional decoding network consists of 5 dense blocks and 1 convolutional layer. Each dense block consists of 2 convolutional layers, which use 128 1×1 filters with a stride of 1 and 32 3×3 filters with a stride of 1.
[0064] The input to the deep convolutional decoding network is a binary pulse matrix of size batch×128×width×height, output after adaptive convolutional pulse coding of the preprocessed training set images. Where "batch" is the number of batches fed into the deep convolutional decoding network, and "width" and "height" are the width and height of the training set images, respectively. The first layer is a Dense block. The network input undergoes a round of batch normalization, ReLU activation, and convolution, followed by another round of these operations. This output is a real-valued matrix of size batch×32×width×height. This layer's output is then concatenated with the input of the current layer to generate a real-valued matrix of size batch×160×width×height, which serves as the input for the next layer.
[0065] From the second to the fifth layers, each layer performs a Dense block operation in sequence. The output of the second layer is a real-valued matrix of batch × 192 × width × height, the output of the third layer is a real-valued matrix of batch × 224 × width × height, the output of the fourth layer is a real-valued matrix of batch × 256 × width × height, and the output of the fifth layer is a real-valued matrix of batch × 32 × width × height. It is worth noting that since there are no Dense blocks following the fifth layer, the output of this layer is not concatenated with the input of the previous layer and is output directly.
[0066] The sixth layer is a convolutional layer, using C 3×3 filters with a stride of 1, where C is the number of channels in the training set images. For the static MNIST dataset, C = 1, and for the static CIFAR10 dataset, C = 3. The input is the real-valued matrix of batch × 32 × width × height output from the fifth layer, and the output is a real-valued matrix of batch × C × width × height.
[0067] Finally, the real-valued matrix of batch×C×width×height output by the sixth layer passes through a Tanh activation function to generate a real-valued matrix of batch×C×width×height, that is, the reconstructed image.
[0068] Given an input image, adaptive convolutional pulse coding (ACP) generates a pulse code sequence within a specified coding time window. This pulse code sequence is then fed into a deep convolutional decoding network for image reconstruction. The reconstructed image is compared for similarity with the original image, and the reconstruction error is backpropagated to optimize the network parameters. Ultimately, a clear and refined reconstructed image is obtained, completing the pre-training process for adaptive convolutional pulse coding.
[0069] Step 4: Build a SCNN network of different depths for the static MNIST dataset and the static CIFAR10 dataset for image classification tasks.
[0070] Read the pre-trained parameters of the adaptive convolutional pulse coding in step 3, which is considered to be a good initialized neural coding layer. The pre-trained adaptive convolutional pulse coding and the back-end deep SCNNs network form an end-to-end image classification network.
[0071] For the static MNIST dataset, the end-to-end SCNN network structure designed in this embodiment of the present invention includes adaptive convolutional pulse coding, 1 convolutional layer, 1 flattening layer and 2 fully connected layers connected in sequence.
[0072] The first layer is adaptive convolutional pulse coding. The input is a real-valued matrix of batch×1×28×28. It passes through the first convolutional layer, batch normalization, and LIF neurons in sequence, and the output is a binary pulse matrix of batch×128×28×28.
[0073] After one average pooling, where the filter size is 2×2 and the stride is 2, a binary impulse matrix of batch×128×14×14 is obtained.
[0074] The second layer is a convolutional layer, which uses 128 3×3 filters with a stride of 1. The input is a binary spike matrix of batch×128×14×14. After one average pooling, where the filter size is 2×2 and the stride is 2, the output is a binary spike matrix of batch×128×7×7.
[0075] The third layer is a flattening layer, which stretches the batch×128×7×7 binary pulse matrix input into a batch×6272 binary pulse vector output. The number of output neurons in this layer is 6272.
[0076] The fourth layer is a fully connected layer. Its input is a binary pulse vector of batch×6272 and its output is a binary pulse vector of batch×2048. The number of output neurons in this layer is 2048.
[0077] The fifth layer is a fully connected layer, whose input is a binary pulse vector of batch×2048 and whose output is a real-valued vector of batch×10, that is, the number of output neurons in this layer is 10.
[0078] For the static CIFAR10 dataset, the end-to-end SCNN network structure designed in the embodiment of the present invention includes adaptive convolutional pulse coding, 4 convolutional layers, 1 flat layer and 3 fully connected layers connected in sequence.
[0079] The first layer is adaptive convolutional pulse coding. The input is a real-valued matrix of batch×3×32×32. It passes through the first convolutional layer, batch normalization, and LIF neurons in sequence, and the output is a binary pulse matrix of batch×128×32×32.
[0080] The second layer is a convolutional layer, which uses 256 3×3 filters with a stride of 1. The input is a binary spike matrix of batch×128×32×32. After one average pooling, where the filter size is 2×2 and the stride is 2, the output is a binary spike matrix of batch×256×16×16.
[0081] The third layer is a convolutional layer, which uses 512 3×3 filters with a stride of 1. The input is a binary spike matrix of batch×256×16×16. After a single average pooling, where the filter size is 2×2 and the stride is 2, the output is a binary spike matrix of batch×512×8×8.
[0082] The fourth layer is the convolutional layer, which uses 1024 3×3 filters with a step size of 1. The input is a binary impulse matrix of batch×512×8×8, and the output is a binary impulse matrix of batch×1024×8×8.
[0083] The fifth layer is the convolutional layer, which uses 512 3×3 filters with a step size of 1. The input is a binary pulse matrix of batch×1024×8×8, and the output is a binary pulse matrix of batch×512×8×8.
[0084] The sixth layer is a flattening layer, which stretches the batch×512×8×8 binary pulse matrix input into a batch×32768 binary pulse vector output. The number of output neurons in this layer is 32768.
[0085] The seventh layer is a fully connected layer. Its input is a binary pulse vector of batch×32768 and its output is a binary pulse vector of batch×1024. The number of output neurons in this layer is 1024.
[0086] The eighth layer is a fully connected layer. Its input is a binary pulse vector of batch×1024 and its output is a binary pulse vector of batch×512. The number of output neurons in this layer is 512.
[0087] The ninth layer is a fully connected layer. Its input is a binary pulse vector of batch×512 and its output is a real-valued vector of batch×10. The number of output neurons in this layer is 10.
[0088] Define the simulated encoding time window as T, the classification task category as C, then the network output O = [o t,i] is a T×C Tensor tensor, the real label Y=[y t,i ], the loss function of the network is defined as The neuron with the highest pulse rate is used as the predicted label
[0089] The network parameter update rule is based on the gradient descent rule, which is similar to the Back Propagation Through Time (BPTT) learning algorithm.
[0090] Technical Effect: The present invention provides an adaptive convolutional automatic encoding method based on spiking neurons that can efficiently and accurately characterize static image information, which has the following advantages:
[0091] 1. Self-determination of the optimal encoding time window: Existing neural coding technologies require a pre-defined encoding time window, but there is currently no scientifically sound method for determining this encoding time window. This present invention utilizes a pulse pixel value mapping decoding method to automatically find the optimal encoding time window for adaptive convolutional pulse coding when processing the static MNIST and static CIFAR10 datasets. The pulse pixel value mapping decoding method decodes the pulse code output by the adaptive convolutional pulse coding into a decoded image. The encoding time window varies from 10 to 20, meaning that each image is encoded and decoded 11 times. Table 1 shows the results of comparing 11 decoded images of a sample test image with the original image using three typical image quality evaluation metrics. It can be seen that T = 12 is the optimal value for the encoding time window for this example image.
[0092] Table 1: PSNR, SSIM, and MSE values of sample images at different encoding time windows
[0093]
[0094] The present invention then compares 11 decoded images of each image in the dataset with the original image using three image quality evaluation metrics. It automatically outputs the encoding time window T corresponding to the optimal values of MSE, PSNR, and SSIM among the 11 results, and then calculates the proportion of each T in the dataset. In the static MNIST dataset, the optimal encoding time window T = 11 has a proportion of up to 99%, and in the static CIFAR10 dataset, the proportion of T = 12 exceeds 84%.
[0095] 2. Determination of Pre-trained Parameters for Adaptive Convolutional Pulse Coding: In this paper, a network derived from the DenseNet model is first used for pulse-to-image decoding and reconstruction tasks. This method can obtain high-quality reconstructed images from a deep convolutional decoding network model, effectively reconstructing both global and detailed image content. Depending on the loss function used for training, intermediate images may differ in color detail processing, but the final reconstructed images are similar with minor differences in texture details. This demonstrates that adaptive convolutional pulse coding can be well trained, thereby determining pre-trained parameters for adaptive convolutional pulse coding.
[0096] 3. Improving the Accuracy of SCNNs: SCNNs primarily consist of stacked convolutional layers for feature extraction, followed by fully connected layers for image classification. This paper designs SCNNs of varying depths for various datasets. For visual classification tasks involving two-dimensional spatial structures, SCNNs are used to directly encode images, extracting effective spatiotemporal features for classification. In experiments on static datasets, the first layer of the SCNNs is replaced by a pre-trained adaptive convolutional pulse coding (ACP) algorithm by default. As shown in Table 2, for the simple static MNIST dataset, the SCNN using the optimized ACP algorithm achieves significantly higher test accuracy than visual classification algorithms using rate coding and other pulse coding methods. Even for the relatively complex static CIFAR10 dataset, the proposed ACP algorithm significantly improves test accuracy and performance. This demonstrates that effective spatiotemporal encoding of static images is crucial for SNN learning for complex classification tasks. By optimizing the encoding process from pixel intensities to spatiotemporal pulse patterns while maximally preserving the original image information, SNNs can achieve better classification performance.
[0097] Table 2: Comparison of classification performance on different visual datasets
[0098]
[0099] Comparative Example 1: Rueckauer B, Lungu IA, Hu YH, et al.Conversion of continuous-valued deep networks to efficient event-driven networks for image classification[J].Frontiers in Neuroscience 11,682,2017.
[0100] Comparative Example 2: Sengupta A, Ye Y, Wang R, et al. Going deeper in spiking neural networks: VGG and residual architectures[J]. Frontiers in neuroscience 13, 95, 2019.
[0101] Comparative Example 3: Lee C, Sarwar S S, Panda P, et al. Enabling spike-based backpropagation for training deep neural network architectures[J]. Frontiers in Neuroscience, 119, 2020.
[0102] Comparative Example 4: Jin Y, Zhang W, Li P. Hybrid macro / micro level backpropagation for training deep spiking neural networks[J]. Advances in neural information processing systems, 31, 2018.
[0103] Comparative Example 5: Severn A, Vineyard C M, Dellana R, et al. Training deep neural networks for binary communication with the whetstone method[J]. Nature Machine Intelligence 1(2), 86 - 94, 2019.
[0104] Comparative Example 6: Gu P, Xiao R, Pan G, et al. STCA: Spatio-temporal credit assignment with delayed feedback in deep spiking neural networks[C]. IJCAI, 1366 - 1372, 2019.
[0105] Comparative Example 7: Wu Y J, Deng L, Li G Q, et al. Direct training for spiking neural networks: Faster, larger, better[J]. Proceedings of the AAAI Conference on Artificial Intelligence 33(01), 1311-1318, 2019。
Claims
1. An adaptive convolutional auto-encoding method based on spiking neurons, characterized in that: The following steps are involved: Step 1: Adaptive convolutional pulse coding includes the first convolutional layer, batch normalization, and spiking neurons connected in sequence. The original image is extracted and batch normalized through the first convolutional layer to obtain a feature map. Then, according to the preset coding time window range, the pulse neurons are input to emit pulses to obtain the original image pulse coding sequence for each preset coding time window. Step 2: Decode the original image pulse code sequence of each pre-set coding time window using the pulse pixel value mapping decoding method to obtain a decoded image. Compare the similarity between the decoded image and the original image when the coding time window changes, and automatically find the optimal coding time window for adaptive convolutional pulse coding to process different data sets. Specifically, it includes: (2.1) According to the original image pulse coding sequence of each pre-set coding time window, calculate the The pulse emission frequency ,in is the encoding time window, is the encoding time window of each neuron The number of pulses released within (2.2) According to each neuron The pulse emission frequency , mapped to the corresponding image pixel value ,in is the maximum pulse firing frequency of all neurons, which is decoded to obtain the decoded image; (2.3) Based on the set coding time window range, the similarity between the decoded image and the original image in each coding time window is calculated in turn, and the image similarities of different coding time windows are compared to automatically find the optimal coding time window for adaptive convolutional pulse coding to process the current data set; Step 3: Based on the optimal coding time window determined in step 2, the original image is re-passed through the first convolutional layer to extract features and batch normalize to obtain a feature map. Then, according to the optimal coding time window, the pulse neurons are input to emit pulses to obtain the pulse coding sequence of the original image. The pulse coding sequence is then input into the deep convolutional decoding network for image reconstruction to obtain the reconstructed image. The adaptive convolution pulse coding learning parameters are optimized based on the error backpropagation between the reconstructed image and the original image to complete the adaptive convolution pulse coding pre-training process.
2. The adaptive convolutional auto-encoding method based on spiking neurons according to claim 1, characterized in that: In step 1, the original image is subjected to the first convolutional layer to extract features and batch normalize to obtain a feature map, which specifically includes: (1.1) Preprocess the original image; (1.2) The preprocessed original image is input into the first convolutional layer to extract features, and then the data is batch normalized to obtain the feature map of the preprocessed original image.
3. The adaptive convolutional auto-encoding method based on spiking neurons according to claim 2, characterized in that: In step 1, the original image pulse coding sequence of each preset coding time window is obtained by inputting pulses into the spiking neurons according to the preset coding time window range, specifically including: (1.3) The feature map of the preprocessed original image obtained in step (1.2) is input into the pulse neuron as an electrical signal, simulating the retina capturing external information to generate action potentials, and generating the original image pulse coding sequence for each pre-set coding time window according to the pre-set coding time window range.
4. The adaptive convolutional auto-encoding method based on spiking neurons according to claim 3, characterized in that: In step (1.3), the spiking neuron model is a leaky integrate-spark model, and its dynamic equation is: ; ; ; in, It is the neurons in The membrane voltage at the moment It is the neurons in The external input current at the moment, is the membrane time decay constant, are the membrane resistance and membrane capacitance of the neuron, It is The synaptic weights of the input neurons, It is The input neurons The first The arrival time of a presynaptic spike, is the kernel function that describes the time decay effect of synaptic current, is the normalization factor, which makes the maximum value of the kernel function 1. is the synaptic current time decay constant, is the Heaviside step function, that is, only the The effect of a pulse input just before the moment on the postsynaptic current.
5. The adaptive convolutional auto-encoding method based on spiking neurons according to claim 1, characterized in that: In step 3, the pulse code sequence of the original image is input into the deep convolutional decoding network for image reconstruction to obtain a reconstructed image, which specifically includes: (3.1) The network structure of the deep convolutional decoding network consists of 5 Block and 1 output convolution layer, the first to fifth layers are block, the sixth layer is the output convolution layer, each Each block consists of two convolutional layers, The two convolutional layers in the block use 128 1×1 filters with a stride of 1 and 32 3×3 filters with a stride of 1 respectively; (3.2) The input of the deep convolutional decoding network is the pulse code sequence batch× of the original image ×width×height binary impulse matrix, where batch is the number of batches input to the deep convolutional decoding network at one time, is the number of output channels of the first convolutional layer, width and height are the width and height of the original image respectively; The first layer is a The network input first performs a round of batch normalization, ReLU activation function and convolution operation in sequence, and then repeats the above operation to output a real-valued matrix of batch×32×width×height. Then, the output of this layer is concatenated with the input of the current layer to generate batch×( +32)×width×height real-valued matrix as the input of the next layer; (3.3) From the second to the fifth layer, each layer executes one Block operation, the output of the second layer is batch×( +64)×width×height real-valued matrix, the output of the third layer is batch×( +96)×width×height real-valued matrix, the output of the fourth layer is batch×( +128)×width×height real-valued matrix, the output of the fifth layer is a batch×32×width×height real-valued matrix; (3.4) The sixth layer is the output convolution layer, using A 3×3 filter with a step size of 1, where is the number of channels of the original image, the input is the real-valued matrix of batch×32×width×height output of the fifth layer, and the output is batch× ×width×height real-valued matrix; (3.5) Finally, the batch× ×width×height real-valued matrix passes through a Tanh activation function to generate batch× ×width×height real-valued matrix, that is, the reconstructed image.