Novel pulse neural network target identification method based on adaptive coding
By introducing a new pulse neural network with adaptive encoding into the artificial intelligence target recognition method, combining the ANN-SNN adaptive encoder and the pulse-driven convolutional neural network model, the problems of large computing overhead, high energy consumption and insufficient model robustness in the existing technology are solved, and more efficient and accurate target recognition is achieved.
Patent Information
- Application Number
- CN202411971437.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
AI Technical Summary
The existing artificial intelligence target recognition methods have high computational overhead and high energy consumption, and the deep neural network model has shortcomings in model robustness, computing efficiency and new task adaptability.
A new pulse neural network target recognition method based on adaptive encoding is proposed. Combined with the ANN-SNN adaptive encoder and the pulse-driven convolutional neural network model, the input image is encoded into a pulse sequence by designing the ANN-SNN adaptive encoder, and feature extraction and recognition are used to use the pulse-driven convolutional neural network model.
It improves the accuracy of pulse coding and identification accuracy, reduces calculation overhead and energy consumption, and enhances the robustness and adaptability of the model.
Smart Images

Figure CN119942069A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a novel pulse neural network target recognition method based on adaptive coding, belonging to the technical field of information perception and recognition. Background Art
[0002] Existing artificial intelligence technology is mainly connectionist artificial intelligence represented by traditional artificial neural networks (ANN), which greatly simplifies the structure and function of biological nervous systems. Traditional artificial neural networks are based on highly interconnected processing between simple computing units. They are information processing models based on nonlinear weighting and statistical data modeling tools, ignoring time information, and therefore cannot efficiently simulate the information processing mechanism of the human brain. On the other hand, although the recognition accuracy of the deep neural network model (DNN) for specific samples can reach or even exceed the recognition level of humans, it has not truly formed an understanding at the cognitive level and is vulnerable to deception attacks from external noise. Therefore, existing DNN models face more and more challenges, such as weak model robustness, high computational overhead, and poor adaptability to new tasks.
[0003] Spiking Neural Network (SNN) is known as the third generation of artificial neural network. It realizes higher-level biological neuron level simulation on the basis of retaining the good characteristics of traditional artificial neural network, and can capture the rich spatiotemporal dynamic characteristics of biological neurons. On the one hand, the input information uses pulse sequence to encode neural information, and contains information of different dimensions such as time, space, frequency and phase. Therefore, compared with the traditional artificial neural network based on frequency coding, the spiking neural network has more biological credibility and computing power. On the other hand, due to the use of binary pulse sequence for information transmission and event-driven information processing, SNN is more suitable for hardware processing and implementation. It can process complex spatiotemporal data in a large-scale parallel, ultra-low power consumption, high performance and strong robustness on the proprietary neuromorphic chip. The present invention focuses on the brain-inspired spiking neural network computing model and pulse coding algorithm, explores the brain-inspired new brain-like neural network theory, proposes a new spiking neural network target recognition method based on adaptive coding, and establishes a brain-like perception model based on spiking neural network, which is expected to innovate a new mode of intelligent perception and provide a technical approach for the development of new intelligent perception technology. Summary of the invention
[0004] The technical problem solved by the present invention is: in view of the problems of high computational overhead and high energy consumption of existing artificial intelligence target recognition methods, a new pulse neural network target recognition method based on adaptive coding is proposed by combining brain-inspired neural network theory with the existing deep neural network framework.
[0005] The technical solution of the present invention is: a novel pulse neural network target recognition method based on adaptive coding, comprising:
[0006] Design an ANN-SNN adaptive encoder to adaptively encode the input image into a pulse train;
[0007] Combined with the typical convolutional neural network architecture, a pulse-driven convolutional neural network model is constructed to extract features from the input pulse sequence;
[0008] The ANN-SNN adaptive encoder, the pulse-driven convolutional neural network model, and the decoding network are connected in series to construct a hierarchical encoding-feature extraction-decoding network architecture, namely, the pulse neural network target recognition model to be trained;
[0009] The data set consisting of input images is input into the pulse neural network target recognition model, and the model training and testing are completed by propagating errors in both time and space dimensions.
[0010] Preferably, each neuron in the ANN-SNN adaptive encoder first receives multiple real-valued values, i.e., image inputs based on the ANN working mode, then processes the information through one or more layers of ANN, and finally generates a pulse sequence as output through the LIF neuron layer based on the SNN working mode.
[0011] Preferably, the pulse-driven convolutional neural network model includes a Spiking LeNet network model, a Spiking AlexNet network model, or a Spiking VGG and Spiking ResNet network model.
[0012] Preferably, the Spiking LeNet network model first obtains the input current value through a layer of convolution structure and batch normalization operation; then, a pulse sequence is generated through a layer of LIF neuron layer, and the spatiotemporal features of the pulse sequence are extracted based on two sets of alternating convolution layers and spatial pooling layers within the simulation period T; finally, the discharge frequency of the output neuron is obtained by decoding through a two-layer fully connected network for image recognition.
[0013] Preferably, the Spiking AlexNet network model first obtains the input current value through a layer of convolution structure and batch normalization operation; then, a pulse sequence is generated through a layer of LIF neuron layer, and the spatiotemporal features of the pulse sequence are extracted based on four sets of alternating convolution layers, batch normalization layers, spatial pooling layers and LIF neuron layers within a simulation period T; finally, the discharge frequency of the output neuron is obtained by decoding through a two-layer fully connected network for image recognition.
[0014] Preferably, the Spiking VGG and Spiking ResNet network models are respectively composed of multiple repeated basic modules; wherein, the basic module of Spiking VGG uses a convolution kernel smaller than 5×5 to extract features, then performs batch normalization processing and reduces the feature dimension based on a spatial pooling layer, and finally generates a pulse sequence through a LIF neuron layer for information transmission; the basic module of Spiking ResNet introduces skip connections for identity mapping, and generates a pulse sequence through a LIF neuron layer for information transmission.
[0015] Preferably, the training of the SAR image target recognition model is realized by propagating errors in two dimensions, time and space, including:
[0016] The surrogate gradient method is used to achieve supervised end-to-end training of SAR image target recognition model by back-propagating errors in both time and space dimensions.
[0017] Alternatively, the STDP method is used to achieve unsupervised training of SAR image target recognition models through local propagation of errors in both time and space dimensions.
[0018] Optimally, alternative functions for backpropagation in both time and space dimensions And using its derivative Complete the back propagation process.
[0019] Among them, α controls the smoothness of the substitution gradient, and V is the membrane potential of the LIF neuron.
[0020] Compared with the prior art, the present invention has the following beneficial effects:
[0021] (1) To address the problem of pulse coding uncertainty, an ANN-SNN adaptive encoder is designed. The ANN layer can mine deeper structural information in the data, and the SNN layer further encodes the data into a pulse sequence based on the dynamic characteristics of the spiking neurons. The encoder can obtain more precise spatiotemporal expression patterns and improve the accuracy of pulse coding.
[0022] (2) In order to address the problems of insufficient learning ability and low recognition accuracy of the SNN model, a high-efficiency and high-performance pulse-driven deep convolutional neural network model was designed by combining the typical convolutional neural network architecture, which achieved high recognition accuracy on multiple data sets. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is an ANN-SNN adaptive encoder;
[0024] Figure 2 It is the Spiking LeNet network model structure;
[0025] Figure 3 Spiking AlexNet network model structure;
[0026] Figure 4 It is the basic module of the Spiking VGG model;
[0027] Figure 5 It is the basic module of the Spiking ResNet model. DETAILED DESCRIPTION
[0028] The specific steps of the method of the present invention are as follows:
[0029] (1) To address the problem of pulse coding uncertainty, an ANN-SNN adaptive encoder is designed, which can adaptively encode the input image into a pulse sequence.
[0030] The pulse neural network simulates biological neuron cells receiving pulse sequences as input. Different input data often require different encoding schemes to encode external stimuli into discrete pulse sequences. Using an encoding method independent of the pulse neural network model will, on the one hand, lose a certain degree of original information characteristics, and on the other hand, will lead to an increase in the simulation cycle, resulting in gradient disappearance. In order to solve the above problems, the present invention designs the following Figure 1 The ANN-SNN adaptive encoder shown is used to generate pulse signals globally. Each neuron in the encoder first receives multiple real-valued images based on the ANN working mode, then processes the information through one or more layers of ANN, and finally generates a pulse sequence as output through the LIF neuron layer based on the SNN working mode.
[0031] For different types of experimental data and learning tasks, ANN-SNN adaptive encoders with different structures can be designed to process input data. In this encoder, the ANN layer mines deeper structural information in the data by performing nonlinear transformation on the input data; the SNN layer further encodes the data into a pulse sequence based on the dynamic characteristics of the pulse neurons to ensure binary nature. At the same time, since the connection structure and number of neurons in the ANN layer can be flexibly adjusted according to the application tasks and data characteristics, and the performance of the overall network architecture can be optimized through learnable model parameters, a more accurate spatiotemporal expression pattern can be obtained.
[0032] (2) In order to address the problems of insufficient learning ability and low recognition accuracy of SNN models, pulse-driven convolutional neural network models of different scales were constructed by combining the typical convolutional neural network architecture.
[0033] In order to improve the learning ability and information processing ability of the SNN model, the present invention combines the typical convolutional neural network architecture to design and implement a deep convolutional neural network model based on pulse driving. The information processing ability of the SNN model is related to the depth of the neural network. For simple image data sets, generally referring to image data sets containing 1 type of target, a shallow Spiking LeNet model is designed; for complex image data sets, generally referring to image data sets containing no less than 3 types of targets, deep Spiking VGG and Spiking ResNet models are constructed in combination with VGG architecture and ResNet architecture respectively; for medium-complex image data sets (other cases), a 7-layer learnable Spiking AlexNet network model is designed.
[0034] 1) Spiking LeNet network model
[0035] For a simple image dataset, the present invention designs a Spiking LeNet network model with the following structure: Figure 2 As shown. First, the original simulation data with Mini-Batch size is passed through a layer of convolution structure and batch normalization operation to obtain a suitable input current value; then, a pulse sequence is generated through a layer of LIF neuron layer, and the spatiotemporal features of the pulse sequence are extracted based on two sets of alternating convolution layers and spatial pooling layers within the simulation period T; finally, the discharge frequency of the output neuron is obtained by decoding through a two-layer fully connected network for image recognition. Compared with the earliest proposed LeNet network structure, the SNN model designed by the present invention only uses 3×3 convolution kernels. The experimental results show that the use of small convolution kernels for feature extraction can achieve better recognition performance while greatly reducing the amount of calculation and the amount of parameters.
[0036] 2) Spiking AlexNet network model
[0037] For medium-complex image datasets, the present invention designs a Spiking AlexNet network model with the following structure: Figure 3 As shown. The model contains a total of 7 learnable weight layers, including 5 convolutional layers and 2 fully connected layers. The ANN-SNN adaptive encoder can generate pulse signals globally to reduce the overall accuracy loss. In order to learn richer feature information and thus improve the information processing capability of the model, the present invention selects the texture features with the largest and strongest response in the features based on the maximum pooling operation in the shallow network, and uses average pooling in the deep network part to retain the overall feature information.
[0038] 3) Spiking VGG and Spiking ResNet network models
[0039] The deep neural network model introduces more nonlinear transformations, which can effectively learn higher-level feature representations and significantly improve the model's expressiveness, which is crucial for recognizing complex input patterns. Therefore, exploring the effectiveness of deep pulse convolutional neural network models is of great research significance. For complex image datasets, the present invention designs Spiking VGG and Spiking ResNet network models based on the VGG network model and the ResNet network model, respectively. The above two models are composed of multiple repeated basic modules, Figure 4 It is the basic module of the Spiking VGG model. Figure 5 It is the basic module of the SpikingResNet model. The basic module of Spiking VGG uses multiple 3×3 small convolution kernels in series to replace large convolution kernels to extract features, which reduces network parameters and deepens the number of network layers while ensuring the same receptive field. Batch normalization is performed after each convolution layer, and then the feature dimension is reduced based on the spatial pooling layer. Finally, a pulse sequence is generated through the LIF neuron layer for information transmission. The basic module of Spiking ResNet introduces skip connections for identity mapping, thereby alleviating the problem of gradient vanishing in deep networks. Each basic module includes two layers of 3×3 convolution kernels, and uses the LIF neuron model for nonlinear activation to replace the ReLu activation function in traditional artificial neural networks.
[0040] (3) To address the problem of non-differentiable pulse activation functions, a gradient substitution method was introduced to achieve end-to-end training of deep pulse neural networks.
[0041] In the SNN model, the pulse discharge process is modeled as a step function, and the activation function of the neuron is zero everywhere except the threshold, and the derivative does not exist at the threshold. Therefore, although the SNN model has a general architecture and time characteristics similar to those of the recurrent neural network, the gradient descent method cannot be directly applied to optimize the pulse neural network. The present invention introduces the approximate derivative of the pulse activation function and proposes a back-propagation algorithm based on pulse learning, which directly trains the SNN model while retaining the dynamic characteristics of the pulse neuron.
[0042] The present invention combines the idea of replacing the gradient, introduces the function f(V), and uses its derivative f′(V) to replace the derivative of the impulse function to complete the back propagation process. The function expression of f(V) and its derivative f′(V) is:
[0043]
[0044] Among them, α can control the smoothness of the alternative gradient and its value is 4; V is the membrane potential of the LIF neuron.
[0045] (4) Using data sets of different complexity, multiple pulse-driven convolutional neural network models constructed in step (2) are tested.
[0046] The present invention uses three datasets of different complexity, Fashion-MNIST, SVHN and CIFAR-10, to evaluate the constructed Spiking LeNet, Spiking AlexNet, Spiking VGG and Spiking ResNet models. Table 1 shows the recognition results of the model designed by the present invention for different datasets. The results show that the three models proposed by the present invention can achieve better experimental results in different datasets. Specifically, for the Fashion-MNIST dataset, the Spiking LeNet model achieved a recognition accuracy of 94.91%; for the SVHN dataset, the Spiking AlexNet achieved a recognition accuracy of 95.78%; for the CIFAR-10 dataset, the present invention constructed VGG-9 and ResNet-11 network models, which achieved recognition accuracies of 93.23% and 93.39% respectively.
[0047] Table 1 Recognition accuracy of different data sets
[0048]
[0049] Parts not described in detail in the present invention belong to common knowledge of those skilled in the art.
Claims
1. A novel pulse neural network target recognition method based on adaptive coding, characterized in that include: Design an ANN-SNN adaptive encoder to adaptively encode the input image into a pulse train; Combined with the typical convolutional neural network architecture, a pulse-driven convolutional neural network model is constructed to extract features from the input pulse sequence; The ANN-SNN adaptive encoder, the pulse-driven convolutional neural network model, and the decoding network are connected in series to construct a hierarchical encoding-feature extraction-decoding network architecture, namely, the pulse neural network target recognition model to be trained; The data set consisting of input images is input into the pulse neural network target recognition model, and the model training and testing are completed by propagating errors in both time and space dimensions.
2. The method according to claim 1, characterized in that: Each neuron in the ANN-SNN adaptive encoder first receives multiple real-valued values, i.e., image inputs based on the ANN working mode, then processes the information through one or more layers of ANN, and finally generates a pulse sequence as output through the LIF neuron layer based on the SNN working mode.
3. The method according to claim 1, characterized in that: The pulse-driven convolutional neural network model includes a Spiking LeNet network model, a Spiking AlexNet network model, or a Spiking VGG and Spiking ResNet network model.
4. The method according to claim 3, characterized in that: The Spiking LeNet network model first obtains the input current value through a layer of convolution structure and batch normalization operation; then, a pulse sequence is generated through a layer of LIF neurons, and the spatiotemporal features of the pulse sequence are extracted based on two sets of alternating convolution layers and spatial pooling layers within the simulation period T; finally, the discharge frequency of the output neuron is obtained by decoding through a two-layer fully connected network for image recognition.
5. The method according to claim 3, characterized in that: The Spiking AlexNet network model first obtains the input current value through a layer of convolution structure and batch normalization operation; then, a pulse sequence is generated through a layer of LIF neuron layer, and the spatiotemporal features of the pulse sequence are extracted based on four sets of alternating convolution layers, batch normalization layers, spatial pooling layers and LIF neuron layers within the simulation period T; finally, the discharge frequency of the output neuron is obtained by decoding through a two-layer fully connected network for image recognition.
6. The method according to claim 3, characterized in that: The Spiking VGG and Spiking ResNet network models are respectively composed of multiple repeated basic modules; the basic module of Spiking VGG uses a convolution kernel smaller than 5×5 to extract features, then performs batch normalization and reduces the feature dimension based on the spatial pooling layer, and finally generates a pulse sequence through the LIF neuron layer for information transmission; the basic module of Spiking ResNet introduces skip connections for identity mapping, and generates a pulse sequence through the LIF neuron layer for information transmission.
7. The method according to claim 1, characterized in that: The training of the SAR image target recognition model is realized by propagating errors in time and space, including: The surrogate gradient method is used to achieve supervised end-to-end training of SAR image target recognition model by back-propagating errors in both time and space dimensions. Alternatively, the STDP method is used to achieve unsupervised training of SAR image target recognition models through local propagation of errors in both time and space dimensions.
8. The method according to claim 7, characterized in that: Alternative functions for backpropagation in time and space And using its derivative Complete the back propagation process. Among them, α controls the smoothness of the substitution gradient, and V is the membrane potential of the LIF neuron.