Image recognition method of pulse neural network based on Fourier and cross-time constraint

By using a finite term Fourier series approximating pulse distribution gradients in the frequency domain and combining cosine similarity constraints, the problem of gradient in pulsed neural network training is solved, and faster convergence and higher accuracy image recognition effect is achieved.

CN120388206AActive Publication Date: 2025-07-29SOUTHWEST UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510345748.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-29
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

In the prior art, the gradient non-differentiation characteristics of pulsed neural networks lead to a decline in training performance, and the pseudo-gradient function may destroy the optimization direction, limiting the direct training performance of pulsed neural networks.

Method used

Using Fourier and cross-time constraints pulse neural network methods, a new loss function is designed to optimize network weights by approximating the gradient of pulse distribution activity using finite term Fourier series in the frequency domain and using cosine similarity to constrain the output distribution of each time step in backpropagation.

Benefits of technology

It effectively alleviates the non-differentiation characteristics of pulse activities, improves the training convergence speed and accuracy of pulse neural networks, and uses the output information of each time step to achieve higher image recognition performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388206A_ABST
    Figure CN120388206A_ABST
Patent Text Reader

Abstract

An image recognition method of a pulse neural network based on Fourier and cross-time constraints is characterized by comprising the following steps: constructing an image recognition system of the pulse neural network based on Fourier and cross-time constraints; the image acquisition module acquires an image sample set, and pre-processes the image sample set to obtain a training set; training the network by using a training set, enabling pulse neurons to emit pulses by using a step function in forward propagation, and adopting a derivative of finite term Fourier series as gradient estimation of the pulses on membrane potential in reverse propagation; measuring the similarity between the time steps by using cosine similarity, calculating the loss of the target function of the network after the rth training, and adjusting network weight parameters according to the loss of the target function; and carrying out image classification identification operation on the preprocessed to-be-identified image data by using the trained network, and outputting an image classification result. The method has the effect of improving model performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and particularly to an image recognition method based on a pulsed neural network with Fourier and cross-time constraints. Background Art

[0002] As a well-known brain-inspired computational model, the spiking neural network uses spiking neurons to emit discrete binary pulse sequences for information transmission similar to biological neurons, thus having biological interpretability and high energy efficiency.

[0003] The special firing mechanism of spiking neurons is a double-edged sword. This all-or-nothing pulse activity inevitably hinders the direct application of the backpropagation algorithm, which has achieved great success in artificial neural networks (ANNs), to train spiking neural networks (SNNs). There are two main methods to alleviate the non-differentiable nature of pulse activity. One is an indirect conversion method called ANN2SNN, which converts a trained ANN into an SNN with the same structure and replaces the non-linear activation in the ANN with spiking neurons, making the pulse frequency of the SNN match the activation in the ANN. Although this method can reduce the performance gap between SNNs and ANNs, a relatively long time step is also required to reduce the error between the activation value and the pulse frequency. Moreover, this method is only applicable to static data sets and is not applicable to neuromorphic data sets, and the conversion method cannot fully utilize the unique spatio-temporal information processing ability of SNNs. Another method is to directly train SNNs through pseudo-gradient learning. The specific approach is to use a smooth curve to replace the original all-or-nothing pulse firing activity during backpropagation. Although pseudo-gradients can be used to directly train SNNs and achieve good performance, there are currently various pseudo-gradient functions, all of which use a differentiable function to approximate the Heaviside function, and there will still be problems of gradient vanishing and gradient explosion, resulting in a decline in the performance of the trained SNNs. Although many efforts have been made to try to solve this problem, the pseudo-gradient functions they use may destroy the main direction of the true gradient.

[0004] Disadvantages of the prior art: Since the gradient of the pulse firing activity is almost everywhere zero, several attempts have been made to alleviate this and optimize SNNs using the backpropagation algorithm with surrogate gradient (SG) functions. However, these pseudo-gradient functions may destroy the main direction of the gradient for optimizing SNNs, thus limiting the performance of directly training SNNs. Summary of the Invention

[0005] An image recognition method based on a pulsed neural network with Fourier and cross-time constraints provided by the present invention can effectively improve the performance of the pulsed neural network model.

[0006] To achieve the above object, an image recognition method based on a pulsed neural network with Fourier and cross-time constraints provided by the present invention is characterized in that it includes the following steps:

[0007] Step 1: Construct an image recognition system based on a pulsed neural network with Fourier and cross-time constraints, where the image recognition system is provided with an image acquisition module, a preprocessing module, and a pulsed neural network with Fourier and cross-time constraints connected in sequence;

[0008] Step 2: The image acquisition module acquires an image sample set and transmits the image sample set to the preprocessing module for preprocessing operations to obtain a training set;

[0009] Step 3: Use the training set to train the pulsed neural network with Fourier and cross-time constraints. In forward propagation, use the step function to make the pulsed neurons emit pulses. In backward propagation, use the derivative of a finite-term Fourier series as the gradient estimate of the pulse on the membrane potential;

[0010] Step 4: Use the cosine similarity CS to measure the similarity between each time step, calculate the loss of the objective function of the pulsed neural network with Fourier and cross-time constraints after the r-th training, and adjust the network weight parameters according to the loss of the objective function;

[0011] Step 5: Repeat steps 3-4. When the set number of training times is reached, end the training to obtain a trained pulsed neural network with Fourier and cross-time constraints;

[0012] Step 6: The image acquisition module acquires the image data to be recognized, transmits the image data to be recognized to the preprocessing module for preprocessing operations, and then transmits it to the trained pulsed neural network with Fourier and cross-time constraints;

[0013] Step 7: The trained pulsed neural network with Fourier and cross-time constraints performs image classification and recognition operations on the preprocessed image data to be recognized and outputs an image classification result.

[0014] Through the above design, using Fourier series in the frequency domain can effectively approximate the gradient of the pulse emission activity. This provides a new perspective for alleviating the non-differentiable characteristics of pulse activity. Using Fourier series can retain the main direction of the true gradient of the pulse activity, so that the trained pulsed neural network SNNs model can converge faster and be more accurate. In addition, by using cosine similarity to make the output distributions of each time step of the SNNs output layer more similar, the information carried by the output of each time step can be better utilized to guide the training of the SNNs, effectively improving the performance of the SNNs model.

[0015] Preferably, in the step 3, the step function is the Heaviside function, and the calculation expression of the Heaviside function is:

[0016]

[0017] where S represents the pulse sequence, t represents the time, V represents the membrane potential across the neuron, V th represents the membrane potential threshold, and Headviside() represents the Heaviside function;

[0018] The calculation expression of the derivative of the finite - term Fourier series is:

[0019]

[0020] where n represents the number of Fourier series; ω represents the angular frequency, and i represents the index variable and is a positive integer.

[0021] Although the pseudo - gradient method enables spiking neural networks (SNNs) to be trained using the error backpropagation algorithm, the currently proposed pseudo - gradient functions all approximate the step function for the original function in the spatial domain. These functions have different shapes and there are certain differences in their performance in different tasks. The present invention approximates the step function using Fourier series in the frequency domain, which can better capture the direction of the true gradient in SNNs.

[0022] Preferably, the derivative h n ′(t) of the finite - term Fourier series is constructed through the following process:

[0023] Any periodic function is expressed as a linear combination of Fourier series composed of trigonometric functions, namely sine and cosine functions, and the expression is as follows:

[0024]

[0025] where f(t) is a function with a period of T = 2π, is the DC component, a n and b n are Fourier coefficients;

[0026] a0, a n and b n The calculation expressions are as follows:

[0027]

[0028]

[0029] The non-periodic step function is extended to a periodic function through periodic extension. The calculation expressions for the coefficients and DC component of the Fourier series of this step function are as follows:

[0030] a0 = 1

[0031] a n = 0

[0032]

[0033] The Fourier series expression of the periodic Heaviside function is as follows:

[0034]

[0035] The derivative expression of the Fourier series is:

[0036]

[0037] Using an infinite-term Fourier series to approximate the Heaviside function is a lossless representation, but the infinite-term Fourier series incurs a huge computational resource overhead. The present invention uses a finite-term Fourier series and ignores the high-frequency part with relatively low energy occupancy. Therefore, the finite-term Fourier series expression is:

[0038]

[0039] The derivative expression of the finite-term Fourier series is:

[0040]

[0041] Use the derivative of the finite-term Fourier series as the derivative of the step function for backpropagation.

[0042] Using a finite-term Fourier series to approximate the Heaviside function has two advantages:

[0043] 1) By using a finite-term Fourier series, the computational overhead of the infinite-term trigonometric function combination is avoided.

[0044] 2) The finite-term Fourier series is more conducive to approximating the Heaviside function. It does not affect the information in the low-frequency part of the Heaviside function, and often the low-frequency information occupies the vast majority of the energy.

[0045] As a preference: In the step 4, the loss function expression of the cosine similarity CS is as follows:

[0046]

[0047] Among them, T represents the total number of time steps, and CS(V[l], V[t]) represents the cosine similarity of the membrane potential at time l and time t.

[0048] In the present invention, by adding a constraint term of cosine similarity to the objective function of SNNs, the output at each time step is made as consistent as possible, reducing the difference between different time steps. The constraint of cosine similarity makes full use of the temporal information contained in the outputs of SNNs at different time steps to improve the generalization ability of SNNs on the test set.

[0049] Preferably: in the step 4, the calculation expression of the objective function is as follows:

[0050] L = L CE + λL CS

[0051]

[0052] Among them, L represents the total loss value, L CE represents the cross-entropy loss, and L CS represents the cosine similarity loss; λ is a hyperparameter used to constrain the magnitude of the cosine similarity loss term; represents the sum of the average membrane potentials of all neurons, k represents the k-th category of image categories, K is the total number of image categories, and y p represents the true label of the p-th neuron; represents the average membrane potential of the p-th neuron,

[0053] L CE is used to ensure that the average membrane potential approaches the true label, and L CS is used to ensure that the similarity of the outputs between each time step is more consistent. The final loss function L can better optimize the update direction of the weights, thereby preventing abnormal total outputs caused by abnormal outputs at specific time steps and reducing the overconfidence of SNNs predictions at specific time points.

[0054] Preferably: when the image sample set is a static data set, the depth residual learning framework SEW-ResNet19 for deep spiking neural networks is used as the spiking neural network based on Fourier and cross-time constraints for training;

[0055] When the image sample set is a neuromorphic data set, a model VGGSNN that combines a VGG network and a spiking neural network is used as the spiking neural network based on Fourier and cross-time constraints for training.

[0056] Preferably: the dynamic equation of the LIF neuron in the spiking neural network based on Fourier and cross-time constraints is:

[0057]

[0058] where τ m is the membrane time constant, and τ m = RC, where R represents resistance and C represents capacitance; V represents the membrane potential across the neuron; I represents the input current, and the input current I is determined by the product of the input pulse train and the synaptic weight, and the expression is as follows:

[0059]

[0060] where W ij represents the weight between the presynaptic neuron i and the postsynaptic neuron j, and S represents the pulse train;

[0061] The continuous differential equation of the LIF neuron is converted into a discrete difference equation, and the expression is as follows:

[0062]

[0063] Preferably: in the step 7, the average membrane potential of the output layer of the spiking neural network based on Fourier and cross-time constraints is used as the classification index, and the expression for updating the membrane potential state of the neurons in the output layer of the spiking neural network based on Fourier and cross-time constraints is:

[0064]

[0065] Each neuron in the output layer corresponds to an image category, and the image category corresponding to the neuron with the highest average membrane potential fired by the output layer is output as the image classification result.

[0066] For example: when the image categories include five image categories of airplane, car, ship, bicycle, and train, the output layer contains five neurons, and the five neurons correspond to the five image categories one by one. If the average membrane potential output by the neuron corresponding to the airplane is the highest, the image classification result of the image data to be recognized is the airplane image.

[0067] In the spiking neural network based on Fourier and cross-time constraints, the input image X is directly encoded to obtain a pulse train input, and then after passing through LIF neurons, a pulse output is obtained. In the last layer of the network, the LIF neurons do not fire pulses (assuming the threshold is infinite), only accumulate the membrane potential, and use the membrane potential as the classification basis.

[0068] Advantages of the present invention:

[0069] 1. The combination of a finite number of Fourier series can be consistent with the Heaviside function in the low-frequency part, with only errors in the high-frequency part. It retains the main direction of the true gradient, enabling the spiking neural network (SNN) to converge faster and achieve high performance.

[0070] 2. Using the temporal information contained in the output distribution of each time step of the output layer of the SNN, a cosine similarity constraint is designed to make the output distribution of each time step more consistent. Brief Description of the Drawings

[0071] Figure 1 It is the overall workflow diagram of the present invention;

[0072] Figure 2 It is the curve of several pseudo-gradient functions in the embodiment after fast Fourier transform and the comparison diagram of the difference between them and the Heaviside in the frequency domain;

[0073] Figure 3 It is the accuracy rate curve diagram of different Fourier series terms n in the embodiment;

[0074] Figure 4 It is the output distribution diagram of the benchmark model and FTC-SNN in the embodiment after the input is DVS-CIFAR10 samples;

[0075] Figure 5 It is the output distribution diagram of the benchmark model and FTC-SNN in the embodiment after the input is CIFAR10 samples;

[0076] Figure 6 It is the network structure schematic diagram of SEW-ResNet19 and VGGSNN in the embodiment. Detailed Embodiments

[0077] The present invention will be further described in detail below with reference to the drawings and specific examples. The following examples or drawings are used to illustrate the present invention, but not to limit the scope of the present invention.

[0078] As Figure 1 shown: An image recognition method for a spiking neural network based on Fourier and cross-time constraints, including the following steps:

[0079] Step 1: Construct an image recognition system for a spiking neural network based on Fourier and cross-time constraints. The image recognition system is provided with an image acquisition module, a preprocessing module, and a spiking neural network based on Fourier and cross-time constraints that are connected in sequence;

[0080] Step 2: The image acquisition module acquires an image sample set and passes the image sample set to the preprocessing module for preprocessing operations to obtain a training set;

[0081] Step 3: Use the training set to train the pulse neural network based on Fourier and cross-time constraints. In forward propagation, use the step function to make the pulse neurons emit pulses, and in backward propagation, use the derivative of the finite-term Fourier series to replace the step function as the gradient estimation of the pulse on the membrane potential;

[0082] Step 4: Use the cosine similarity CS to measure the similarity between each time step, calculate the loss of the objective function of the pulse neural network based on Fourier and cross-time constraints after the r-th training, and adjust the network weight parameters according to the loss of the objective function;

[0083] Step 5: Repeat steps 3 - 4. When the set number of training times is reached, end the training to obtain the trained pulse neural network based on Fourier and cross-time constraints;

[0084] Step 6: The image acquisition module acquires the image data to be recognized, and after passing the image data to be recognized to the preprocessing module for preprocessing operations, passes it to the trained pulse neural network based on Fourier and cross-time constraints;

[0085] Step 7: The trained pulse neural network based on Fourier and cross-time constraints performs image classification and recognition operations on the preprocessed image data to be recognized and outputs the image classification result.

[0086] In step 2, the preprocessing operations include image cropping and image enhancement.

[0087] In step 3, the step function is the Heaviside function, and the calculation expression of the Heaviside function is:

[0088]

[0089] where S represents the pulse sequence, t represents the time, V represents the membrane potential across the neuron, V th represents the membrane potential threshold, and Heaviside() represents the Heaviside function;

[0090] The calculation expression of the derivative of the finite-term Fourier series is:

[0091]

[0092] where n represents the number of Fourier series; w represents the angular frequency, i represents the index variable and is a positive integer.

[0093] The derivative h n ′(t) of the finite-term Fourier series is constructed through the following process:

[0094] Any periodic function is expressed as a linear combination of Fourier series composed of trigonometric functions, namely sine and cosine functions, and the expression is as follows:

[0095]

[0096] where f(t) is a function with a period of T = 2π, is the DC component, and a n and b n are Fourier coefficients;

[0097] The calculation expressions for a0, a n and b n are as follows:

[0098]

[0099] The non-periodic step function is extended to a periodic function through periodic extension, and the calculation expressions for the coefficients and DC component of the Fourier series of this step function are as follows:

[0100] a0 = 1

[0101] a n = 0

[0102]

[0103] The Fourier series expression of the periodic Heaviside function is as follows:

[0104]

[0105] The derivative expression of the Fourier series is:

[0106]

[0107] Therefore, the finite-term Fourier series expression is:

[0108]

[0109] The derivative expression of the finite-term Fourier series is:

[0110]

[0111] There are two advantages in using the finite-term Fourier series to approximate the Heaviside function:

[0112] 1) By using the finite-term Fourier series, the computational overhead of infinite-term trigonometric function combinations can be avoided.

[0113] 2) The finite-term Fourier series is more conducive to approximating the Heaviside function, which does not affect the information in the low-frequency part of the Heaviside function. Often, the low-frequency information occupies the vast majority of the energy.

[0114] To more intuitively illustrate the second advantage, in this embodiment, the fast Fourier transform is respectively performed on h n (t), the differentiable pulse function Dspike(t), and the evolutionary function EvAF(t), and the differences between the Heaviside function and these three functions after FFT are plotted, as Figure 2 shown. It can be seen from Figure 2 the third row of n that the h

[0115] (t) function uses a combination of 20 trigonometric functions to approximate the Heaviside function. In the frequency domain, its low-frequency part is consistent with the Heaviside function, while there are certain differences between Dspike(x) and EvAF(x) and the Heaviside(x) function in both the low-frequency and high-frequency parts. Using the finite-term Fourier series can more accurately maintain the main gradient direction of the original Heaviside function, that is, the low-frequency part that occupies the most energy.

[0116]

[0117] where T represents the total time step, and CS(V[l], V[t]) represents the cosine similarity of the membrane potential at time l and time t.

[0118] In the step 4, the calculation expression of the objective function is as follows:

[0119] L = L CE + λL CS

[0120]

[0121] where L represents the total loss value, L CE represents the cross-entropy loss, and L CS represents the cosine similarity loss; λ is a hyperparameter used to constrain the magnitude of the cosine similarity loss term; represents the sum of the average membrane potentials of all neurons, k represents the k-th image category, K is the total number of image categories, and y p represents the true label of the p-th neuron; represents the average membrane potential of the p-th neuron,

[0122] When the image sample set is a static data set, the deep residual learning framework SEW-ResNet19 for deep spiking neural networks is used for training as the spiking neural network based on Fourier and cross-time constraints;

[0123] When the image sample set is a neuromorphic data set, the model VGGSNN that combines the VGG network and the spiking neural network is used for training as the spiking neural network based on Fourier and cross-time constraints.

[0124] The dynamic equation of the LIF neuron in the spiking neural network based on Fourier and cross-time constraints is as follows:

[0125]

[0126] where τ m is the membrane time constant, and τ m = RC, where R represents resistance and C represents capacitance; V represents the membrane potential across the neuron; I represents the input current, and the input current I is determined by the product of the input spike train and the synaptic weight, and the expression is as follows:

[0127]

[0128] where W ij represents the weight between the presynaptic neuron i and the postsynaptic neuron j, and S represents the spike train;

[0129] The continuous differential equation of the LIF neuron is converted into a discrete difference equation, and the expression is as follows:

[0130]

[0131] In step 7, the average membrane potential of the output layer of the spiking neural network based on Fourier and cross-time constraints is used as the classification index, and the expression for updating the membrane potential state of the neurons in the output layer of the spiking neural network based on Fourier and cross-time constraints is:

[0132]

[0133] Each neuron in the output layer corresponds to an image category, and the image category corresponding to the neuron with the highest average membrane potential fired by the output layer is output as the image classification result.

[0134] Next, through ablation experiments, the spiking neural networks (SNNs) trained using only the Fourier series (FS) as the pseudo-gradient function are compared with the baseline, and then the cosine similarity constraint is imposed on its objective function to verify the effectiveness of the present invention. Then, the present invention is compared with other state-of-the-art methods on two types of datasets. Finally, the influence of the number of Fourier series (FS) is analyzed and the visualization results of temporal similarity are provided.

[0135] To provide more convincing results, this embodiment verifies the superiority of the present invention for training SNNs on two challenging types of datasets, namely the static CIFAR-10 and CIFAR-100 datasets, as well as the neuromorphic datasets DVS-CIFAR10 and N-Caltech101.

[0136] The CIFAR-10 dataset contains 10 classes of a total of 60,000 color images, where 50,000 images are used as training images and the remaining 10,000 are used as test images. The number of images in each class in the training set and the test set is equal, and the size of each image is 32*32. The CIFAR-100 dataset contains 100 classes of a total of 60,000 images, with 600 color images of size 32*32 for each class, where 500 images are used as the training set and 100 images are used as the test set. The DVS-CIFAR10 dataset converts 10,000 images in the CIFAR-10 dataset into 10,000 event stream data through a dynamic vision sensor. 9,000 event stream data are used as the training set, and the remaining 1,000 event stream data are used as the test set. The N-Caltech101 dataset contains 101 classes of event stream data, which are divided into a training set and a test set in a ratio of 9:1.

[0137] For the CIFAR-10 / 100 training set, data augmentation methods such as random crop, RandomHorizontalFlip, and randomly cutting out some regions in the sample and filling with 0 pixel values (cutout) are applied for data augmentation. For the DVS-CIFAR10 training set, since the original resolution of the sensor is 128*128, resulting in the original images being 128*128, first convert its resolution size to 48*48, and then apply data augmentation methods such as random crop and random horizontal flip. For the N-Caltech101 training set, also convert its resolution size to 48*48, and then apply the same data augmentation method.

[0138] For the CIFAR-10 / 100 dataset, the network structure adopted in this embodiment is SEW-ResNet19, and for the neuromorphic dataset, the network structure is VGGSNN(64C3-128C3-AP2-256C3-256C3-AP2-512C3-512C3-AP2-512C3-512C3-AP2-10FC). Here, C3 indicates that the size of the convolutional kernel is 3*3, AP2 indicates using average pooling and the size of the kernel is 2*2, and FC indicates the fully connected layer. The network structures of SEW-ResNet19 and VGGSNN are as Figure 6 shown.

[0139] The hyperparameters of this experiment are slightly different for the static datasets for ANNs and the neuromorphic datasets for SNNs. Table 1 lists the detailed hyperparameter settings including the leakage term of neurons, firing threshold, encoding time window, and learning rate.

[0140] For CIFAR-10 / 100, Direct coding is adopted, that is, directly inputting static images into SNNs, and using its first layer to encode the images into binary pulse sequences. For the remaining two neuromorphic datasets, since each event stream contains a lot of event information, directly inputting it into SNNs will cause a large amount of memory occupation and an increase in training time. Therefore, the original event stream information is integrated into frame data. In addition, in order not to lose the information contained in the membrane potential, the neurons in the output layer only integrate the input from the previous layer, and the membrane potential does not decay and does not fire pulses. For CIFAR-10 / 100, an optimizer of SGD with momentum of 0.9 and weight decay of 5e-4 is used, and the learning rate decays cosine-wise to 0 during training. For DVS-CIFAR10 and N-Caltech101, the Adam optimizer with weight decay of 1e-4 is used, and the learning rate also decays cosine-wise to 0 during training. In addition, in order to save video memory consumption and accelerate the model training time, mixed precision is used to train SNNs. The following accuracies are the results of 3 repeated experiments with different random seeds and are shown in the form of mean and standard deviation.

[0141] Table 1

[0142]

[0143] In the ablation experiments, VGGSNN was used as the backbone, and multiple ablation experiments were conducted on the DVS-CIFAR10 and N-Caltech101 datasets to verify the effectiveness of the two modules proposed in the present invention. For fair comparison, the widely used triangular pseudo-gradient function [ref] was used to calculate the gradient of the spike generation behavior of spiking neurons during backpropagation, and the loss function used the cross-entropy between the mean membrane potential of the output layer and the target label, that is, the cross-entropy loss function L CE , and the SNNs model obtained through such training was used as the benchmark. As shown in Table 2, for the DVS-CIFAR10 dataset, the first row represents the benchmark, and the second row represents replacing the pseudo-gradient function with the Fourier series FS. It can be seen that using FS can benefit the trained model, increasing the accuracy from 81.10% to 81.90%. As shown in the third row, combining FS and the cosine similarity constraint can achieve the best performance, that is, increasing from 81.10% to 83.70%, indicating the superiority of using FS and the cosine similarity constraint simultaneously during training. Similarly, in the N-Caltech101 dataset, it can also be seen that using FS alone can increase the accuracy from 79.87% to 81.62%, and using both modules can increase the accuracy to 82.06%. In summary, a series of experiments in Table 2 show that the two modules proposed in the present invention help to improve the performance of SNNs

[0144] Table 2

[0145]

[0146] In Table 2 One of the most popular among many pseudo-gradient functions is the triangular pseudo-gradient function, and its expression is:

[0147]

[0148] where γ is a constraint factor to control the range and magnitude of the gradient

[0149] To demonstrate the superiority of the method proposed in this invention, a comparison with several existing works on the challenging CIFAR-10 / 100 datasets containing more complex information is shown in Table 3. For the CIFAR-10 dataset, as can be seen from the table, among the methods of training SNNs from scratch, the Fourier and cross-time constraint based spiking neural network FTC-SNN achieves an accuracy 0.17% higher than that of RecDis-SNN which uses 6 time steps and is trained with the membrane potential distribution loss, with only 2 time steps. The TET method encourages the membrane potential distribution of the output layer of SNNs at each time step to approximate the true label and also utilizes the time information processing ability of SNNs, achieving an accuracy of 94.50% at 6 time steps, while FTC-SNN can achieve an accuracy of 95.72% at 2 time steps, 1.22% higher than TET. The Joint A-SNN obtained through the hybrid training method can achieve an identification accuracy of 95.45% at 4 time steps, and FTC-SNN is 0.36% higher than it at the same time step. The SRP trained through the conversion method achieves an accuracy of 95.60% at 8 time steps, while FTC-SNN can achieve a higher recognition effect and low latency at 2 time steps. Additionally, for the CIFAR-100 dataset, the SRP method through ANN2SNN achieves an accuracy of 76.45% at 32 time steps using the VGG16 network, and the SGS-SNN directly trained from scratch achieves an accuracy of 76.74% at 6 time steps, while FTC-SNN can achieve comparable or even better performance with only 2 time steps. Comparative experiments with a series of recent works on these two static datasets fully demonstrate that FTC-SNN can achieve high performance with low latency.

[0150] Table 3

[0151]

[0152] Among them, PLIF is Parametric Leaky Integrate-and-Fire, Dspike is Differentiable Spike, Joint A-SNN is the joint training of artificial neural network and spiking neural network, SRP is the optimization strategy based on residual membrane potential, TET is Time-Efficient Training, RecDis-SNN is to correct the membrane potential distribution to directly train spiking neural network, SGS-SNN is Substitute Gradient Scaling for directly training spiking neural network, LIAF is Leaky Integrate-and-Fire Digital, TCJA-SNN is the Time-Channel Joint Attention of spiking neural network, NDA is Neuromorphic Data Augmentation, Eventmix is the efficient data augmentation strategy based on event learning, and FTC-SNN is the Fourier and cross-time constraint based spiking neural network.

[0153] Compared with the static datasets for ANNs, neuromorphic datasets inherently contain rich temporal information and have a very high temporal correlation, making them more suitable for SNNs to exert their temporal information processing capabilities. Both types of datasets are used to train SNNs with the VGGSNN network. The method proposed in this invention is compared with existing work. As shown in Table 4, FTC-SNN achieves the currently most superior performance on both datasets. On the DVS-CIFAR10 dataset, the method of this invention achieves the best recognition accuracy of 83.70%, which is about 3% better than TCJA-SNN that incorporates an attention mechanism and about 0.53% better than the TET method. On the N-Caltech101 dataset, the method of this invention reaches the best performance of 82.06%, exceeding TCJA-SNN by about 3.56%. The other two neuromorphic data augmentation methods, NDA and EventMix, achieve accuracies of 78.60% and 79.50% respectively using the ResNet-19 / 18 network structure. The method of this invention also uses a data augmentation method and exceeds the NDA method and the EventMix method by about 3.46% and 2.56% respectively under the VGGSNN network structure. In these two neuromorphic datasets, there are certain differences in the inputs at different time steps, so the outputs at each time step are also different. After using the cosine similarity constraint, the outputs at each time step can be made to tend to be consistent, thereby improving its recognition accuracy.

[0154] Table 4

[0155]

[0156] This invention uses a finite-term Fourier series to approximate the pulse activity. The approximation effects are significantly different for different numbers of Fourier series. A larger number of Fourier series can reduce the gap between it and the Heaviside function, but at the same time, it will cause more calculations, ultimately resulting in an increase in the training time of SNNs. Here, using ResNet-20 as the benchmark, a series of experiments are carried out on the CIFAR-10 dataset and the accuracies obtained by SNNs with different terms of Fourier series as the pseudo-gradient function are given. As Figure 3As shown, during training, the Fourier series term n is fixed from 1 to 17, and the change in the test accuracy of SNNs is shown. It can be seen that when n is fixed from 3 to 12, the performance of the network is better than other values. Since the SNNs of the present invention are trained from scratch, when n is fixed to be greater than 12, although the difference between the Fourier series term and the Heaviside function is less than when n is fixed from 3 to 11 at this time, the performance of the network begins to decline. When the Fourier series term n is fixed to 6, 7, 8, the performance of SNNs becomes saturated, so let n start from a relatively mild value and change as the training progresses. This setting enables the network to start training with a mild n in the early stage of training, and a larger n in the later stage can make up for the difference between the Fourier series and the impulse activity function. In the previous experiments 4.2, 4.3, 4.4, let n gradually increase from the mild value ns (such as 7) to 2ns.

[0157] The present invention makes the output of each time step tend to be consistent by imposing a constraint on the loss function, so that SNNs can achieve high performance with lower latency when the output is consistent at each step. This embodiment will visualize the output distribution caused by different types of inputs. First, for visualizing the distribution of the output layer of the FTC-SNN and the benchmark model at different time steps for an input sample in the DVS-CIFAR10 dataset, the visualization results are as Figure 4 shown. It can be seen that the output of the benchmark model at each time step shows great differences. The output of each neuron is very uniform at the first time step, and the true label is the tenth category, resulting in incorrect recognition. In addition, at time steps 3, 6, 7, 9, 10, the output distributions of the benchmark model are inconsistent, and the input samples are respectively recognized as 6, 2, 6, 2, 6. In contrast, the FTC-SNN constrains the output of the originally evenly distributed ten neurons to the tenth neuron with the highest probability at the first time step, thus correcting the recognition result to the true category. In addition, from the second time step to the tenth time step, the output distribution of the FTC-SNN shows strong consistency, and the tenth neuron shows the highest probability, and the probabilities of the remaining neurons are basically 0. Similarly, as Figure 5As shown, a sample from CIFAR-10 is input into the baseline model and FTC-SNN, and the output distributions at different time steps are shown. Finally, the calculated average cosine similarities of the baseline model are 0.6394 and 0.4245 when the input samples are CIFAR-10 and DVS-CIFAR10, respectively, while those of FTC-SNN are 0.9819 and 0.9788. It can be seen from the calculated results that the cosine similarity constraint of the present invention can make the output distributions at each time step more consistent. In summary, FTC-SNN can make the output at each time step tend to be consistent through cosine similarity, and is more likely to achieve high-accuracy performance within a shorter time step.

[0158] The present invention aims to alleviate the difficulty of non-differentiable impulse activation in directly trained SNNs from a new perspective. Different from previous gradient approximations of impulse activities based on spatially differentiable functions, the present invention provides insights from the frequency domain perspective. There are certain differences between the previous dynamic pseudo-gradients and the original Heaviside function when approximating the gradient of impulse activities, whether in the low-frequency part or the high-frequency part. However, the combination of using a finite number of Fourier series proposed in the present invention can be consistent with the Heaviside function in the low-frequency part and only has errors in the high-frequency part, and it retains the main direction of the true gradient, so it can make SNNs converge faster and achieve high performance. Then, the present invention designs a cosine similarity constraint to make the output distributions at each time step more consistent by using the time information contained in the output distribution of each time step of the output layer of SNNs. Combining these two methods, FTC-SNN has achieved a competitive performance in comparison with existing work.

[0159] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An image recognition method for a pulsed neural network based on Fourier and cross-time constraints, characterized in that Including the following steps: Step 1: Construct an image recognition system based on a pulsed neural network with Fourier and cross-time constraints. The image recognition system is provided with an image acquisition module, a preprocessing module, and a pulsed neural network with Fourier and cross-time constraints connected in sequence; Step 2: The image acquisition module acquires an image sample set and passes the image sample set to the preprocessing module for preprocessing operations to obtain a training set; Step 3: Use the training set to train the pulsed neural network with Fourier and cross-time constraints. In forward propagation, use the step function to make the pulsed neurons fire pulses. In backward propagation, use the derivative of a finite-term Fourier series as the gradient estimate of the pulse on the membrane potential; Step 4: Use the cosine similarity CS to measure the similarity between each time step, calculate the loss of the objective function of the pulsed neural network with Fourier and cross-time constraints after the r-th training, and adjust the network weight parameters according to the loss of the objective function; Step 5: Repeat steps 3-4. When the set number of training times is reached, end the training to obtain a trained pulsed neural network with Fourier and cross-time constraints; Step 6: The image acquisition module acquires the image data to be recognized, passes the image data to be recognized to the preprocessing module for preprocessing operations, and then passes it to the trained pulsed neural network with Fourier and cross-time constraints; Step 7: The trained pulsed neural network with Fourier and cross-time constraints performs image classification recognition operations on the preprocessed image data to be recognized and outputs the image classification result.

2. The image recognition method of the pulsed neural network based on Fourier and cross-time constraints according to claim 1, characterized in that: In step 3, the step function is the Heaviside function, and the calculation expression of the Heaviside function is: Among them, S represents the pulse sequence, t represents the moment, V represents the membrane potential across the neuron, and V th represents the membrane potential threshold, and Heaviside() represents the Heaviside function; The calculation expression of the derivative of the finite-term Fourier series is: where n represents the number of Fourier series; w represents the angular frequency, i represents an index variable and is a positive integer.

3. The image recognition method of the pulsed neural network based on Fourier and cross-time constraints according to claim 2, characterized in that: The derivative h n ′(t) of the finite-term Fourier series is constructed through the following process: Represent an arbitrary periodic function as a linear combination of Fourier series composed of trigonometric functions, namely sine and cosine functions. The expression is as follows: where \(f(t)\) is a function with a period of \(T = 2\pi\), is the DC component, \(a\) n and \(b\) n are Fourier coefficients; a0, a n and b n The calculation expressions are as follows: Extend a non-periodic step function into a periodic function through periodic extension. The calculation expressions of the coefficients and DC components of the Fourier series of this step function are as follows: a0=1 a n =0 The Fourier series expression of the periodic Heaviside function is as follows: The derivative expression of the Fourier series is: Therefore, the finite-term Fourier series expression is: The derivative expression of the finite-term Fourier series is:

4. The image recognition method of the pulsed neural network based on Fourier and cross-time constraints according to claim 1, characterized in that: In step 4, the loss function expression of the cosine similarity CS is as follows: Among them, T represents the total number of time steps, and CS(V[l],V[t]) represents the cosine similarity of the membrane potential at time l and time t.

5. The image recognition method of the spiking neural network based on Fourier and cross-time constraints according to claim 4, characterized in that: In step 4, the calculation expression of the objective function is as follows: L = L CE + λL CS Among them, \(L\) represents the total loss value, \(L\) CE represents the cross-entropy loss, \(L\) CS represents the cosine similarity loss; \(\lambda\) is a hyperparameter; represents the sum of the average membrane potentials of all neurons, \(k\) represents the \(k\)-th class of image categories, \(K\) is the total number of image categories, \(y\) p represents the true label of the \(p\)-th neuron; represents the average membrane potential of the \(p\)-th neuron, 6. The image recognition method of the pulsed neural network based on Fourier and cross-time constraints according to claim 1, characterized in that: When the image sample set is a static data set, use the deep residual learning framework SEW-ResNet19 for deep pulsed neural networks as the pulsed neural network with Fourier and cross-time constraints for training; When the image sample set is a neuromorphic data set, use the model VGGSNN that combines the VGG network and the pulsed neural network as the pulsed neural network with Fourier and cross-time constraints for training.

7. The image recognition method of the pulsed neural network based on Fourier and cross-time constraints according to claim 1, characterized in that: The dynamic equation of the LIF neuron in the pulse neural network based on Fourier and cross-time constraints is as follows: where τ m is the membrane time constant, and τ m = RC, where R represents resistance and C represents capacitance; V represents the membrane potential across the neuron; I represents the input current, and the input current I is determined by the product of the input pulse train and the synaptic weight, and the expression is as follows: Among them, W ij represents the weight between the presynaptic neuron i and the postsynaptic neuron j, and S represents the spike train; Convert the continuous differential equation of the LIF neuron into a discrete difference equation, and the expression is as follows:

8. The image recognition method of the pulsed neural network based on Fourier and cross-time constraints according to claim 1, wherein: In the step 7, take the average membrane potential of the output layer of the pulse neural network based on Fourier and cross-time constraints as the classification index, and the expression for updating the membrane potential state of the neurons in the output layer of the pulse neural network based on Fourier and cross-time constraints is: Each neuron in the output layer corresponds to an image category, and the image category corresponding to the neuron with the highest average membrane potential fired by the output layer is output as the image classification result.

Citation Information

Patent Citations

  • Image data classification method for online training spiking neural network model along with time

    CN114998659A

  • Method and related device for training brain-like gesture recognition model and gesture category recognition

    CN117830799A

  • Semantic segmentation method based on event and spiking neural network

    CN117934850A

  • System and method for acoustic based gesture tracking and recognition using spiking neural network

    EP4170382A1

  • Hardware-oriented deep spiking neural network speech recognition method and system

    WO2024152583A1