Pulse neural network image processing method and system based on frequency domain filtering module
By introducing a frequency domain filtering module into the pulsed neural network, the Gaussian frequency selective filter is used to remove redundant information and noise in the image data, and the problem of difficulty in removing redundant information when processing image data is solved, achieving more efficient calculations and lower energy consumption.
Patent Information
- Application Number
- CN202510526281.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Current pulsed neural networks are difficult to effectively remove redundant information when processing image data, resulting in increased computing burden, performance degradation and energy consumption.
The frequency domain filtering module is used to perform frequency domain processing on the input data, and redundant information and noise are removed through a Gaussian frequency selective filter, and key frequency components are enhanced, thereby generating sparse frequency domain data.
Through frequency domain filtering processing, the redundancy of the input data is significantly reduced, the computing efficiency and accuracy of the pulsed neural network are improved, the robustness and stability of the model are enhanced, and energy consumption is reduced.
Smart Images

Figure CN120071022A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of the application of spiking neural networks in the field of image processing, and particularly relates to a method and system for image processing of spiking neural networks based on a frequency-domain filtering module. Background Art
[0002] A spiking neural network (SNN) represents a new paradigm for low-power computing. The working principle of SNN is closer to that of the biological nervous system, which transmits information through discrete spike signals. Compared with traditional artificial neural networks (ANNs), it has significant energy efficiency advantages. SNN shows unique potential in simulating the computational and information processing methods of the brain, and is therefore considered to be one of the important development directions of the next-generation neural networks.
[0003] Although SNN has the potential for low-power computing, however, current research on reducing the redundancy of input (whether it is image data or event data) is still relatively scarce. In practical applications, input data often has redundancy. Especially in the process of processing image data or event data, redundant information will increase the computational burden, resulting in performance degradation and increased energy consumption. Traditional SNN models usually rely on information processing in the time domain when dealing with such redundant data, and this processing method cannot effectively eliminate the redundant information in the data, especially when facing complex and noisy environments.
[0004] To solve this problem, the present invention proposes an innovative method - a frequency-domain filtering module (FF). This frequency-domain filtering module can process the redundancy of input data in the frequency domain, not only can remove noise, enhance the robustness and performance of the model, but also can make the input data sparser, thereby reducing energy consumption. Through frequency-domain filtering processing, the FF module can effectively reduce the redundant part in the input signal, and improve the computational efficiency and accuracy of SNN.
[0005] The core advantage of frequency-domain filtering lies in its ability to process high-frequency noise in input data from the frequency-domain perspective. Through frequency-domain filtering, the network can more easily learn the important features in the data, thereby enhancing its adaptability to different environments, interferences or fault conditions. Especially when facing the noise in image or event data, the FF module can effectively filter out irrelevant frequency components and retain the key low-frequency information, thereby enhancing the robustness and stability of the model. Summary of the Invention
[0006] The object of the present invention is to provide a method and system for processing images of spiking neural networks based on a frequency-domain filtering module in view of the deficiencies of the prior art. The present invention uses a frequency-domain filtering module to remove redundant information and enhance key features in the frequency domain, thereby improving the efficiency of image processing and the performance of the model; at the same time, the redundancy of the input data is reduced through frequency-domain processing, and the energy consumption is reduced.
[0007] The object of the present invention is achieved by the following technical solutions: In the first aspect of the embodiments of the present invention, a method for processing images of spiking neural networks based on a frequency-domain filtering module is provided, including the following steps: (1) Obtain time-domain image data; (2) Convert the time-domain image data into frequency-domain data through discrete Fourier transform; (3) Use a frequency-domain filtering module to process the frequency-domain data to remove redundant information therein and enhance the features corresponding to key frequency components, generating filtered frequency-domain data; wherein, the frequency-domain filtering module is implemented by a Gaussian-type frequency-selective filter; (4) Convert the filtered frequency-domain data back into time-domain data through inverse discrete Fourier transform to obtain the converted time-domain data; (5) Use a spiking neural network to process the converted time-domain data for image feature extraction and classification, and output the final image classification result; wherein, the spiking neural network includes an encoder, a spiking neural network module based on TLIF neurons, and an output layer, and the encoder is used to convert the input converted time-domain data into pulse signals that can be processed by TLIF neurons, and the output of the encoder is then transmitted to the spiking neural network module based on TLIF neurons and the output layer to obtain the image classification result.
[0008] Further, in the step (2), the expression of the frequency-domain data is: ; In the formula, represents the frequency-domain data, f represents the frequency variable, represents the time-domain data, n represents the nth time step, and N is the signal length of the time-domain data.
[0009] Further, the frequency response function of the Gaussian-type frequency-selective filter is: ; In the formula, represents the frequency response function of the Gaussian-type frequency-selective filter, represents the frequency component in the horizontal direction, represents the frequency component in the vertical direction, and Denote the horizontal and vertical coordinates at the central position of the frequency response function in the frequency domain. Denote the bandwidth of the Gaussian frequency selective filter.
[0010] Furthermore, in the step (4), the expression of the converted time-domain data is: ; In the formula, Denote the converted time-domain data, Denote the filtered frequency-domain data.
[0011] Furthermore, the pulsed neural network module based on TLIF neurons includes a downsampling layer and multiple convolutional pulsed neural network modules. Among them, the convolutional pulsed neural network module adopts the Spiking VGG architecture, the MS-ResNet architecture or the SpikiFormer architecture; The downsampling layer is used to unify the image format of the input converted time-domain data x′(n) through downsampling operations to obtain unified image data; then the unified image data is sent into the convolutional pulsed neural network module for feature extraction to obtain image features. This convolutional pulsed neural network module performs feature extraction through the synaptic connections between TLIF neurons and the dynamic changes of their membrane potentials; finally, the extracted image features pass through the output layer to output the final image classification result, and this output layer performs image classification according to the pulse firing mode of TLIF neurons.
[0012] Furthermore, the Spiking VGG architecture includes multiple convolutional layers, multiple batch normalization layers and multiple TLIF neurons. The unified image data is sent into the convolutional pulsed neural network module and passes through the convolutional layer, the batch normalization layer and the TLIF neuron in sequence to process the image data. This Spiking VGG architecture realizes feature learning by simulating the time integration process and the pulse triggering mechanism of TLIF neurons to achieve feature extraction of the image data and obtain image features.
[0013] Furthermore, the MS-ResNet architecture includes multiple TLIF neurons, multiple convolutional layers and multiple batch normalization layers, and is combined with residual connections and multi-scale convolutional layers. The unified image data is sent into the convolutional pulsed neural network module and passes through the TLIF neuron, the convolutional layer, the batch normalization layer, the TLIF neuron, the convolutional layer and the batch normalization layer in sequence to process the image data, extract the image features of the current scale, and fuse the image features of the previous scale with the image features of the current scale to obtain the fused features of the current scale, and finally obtain the multi-scale fused image features.
[0014] Furthermore, the SpikiFormer architecture is based on the self-attention mechanism and uses an adaptive weighting method to extract various features in the image. The SpikiFormer architecture includes multiple separable convolutional layers, multiple channel convolutional layers, a downsampling layer, multiple impulse-driven self-attention layers, and multiple channel multi-layer perceptrons. The uniformly formatted image data is fed into the convolutional spiking neural network module. First, it passes through the separable convolutional layer. After weighting and fusing the features output by the separable convolutional layer with the features input to the separable convolutional layer, it is input into the channel convolutional layer. After weighting and fusing the features output by the channel convolutional layer with the features input to the channel convolutional layer, it is fed into the next separable convolutional layer. After weighting and fusing the features output by the last channel convolutional layer with the features input to this channel convolutional layer, it is input into the downsampling layer. The features output by the downsampling layer are then input into the impulse-driven self-attention layer. After weighting and fusing the features output by the impulse-driven self-attention layer with the features input to this impulse-driven self-attention layer, it is fed into the channel multi-layer perceptron. After weighting and fusing the features output by the channel multi-layer perceptron with the features input to this channel multi-layer perceptron, it is fed into the next impulse-driven self-attention layer. After weighting and fusing the features output by the last channel multi-layer perceptron with the features input to this channel multi-layer perceptron, the final image features can be obtained.
[0015] Furthermore, the training process of the spiking neural network specifically includes: applying perturbations to the transformed time-domain data to obtain perturbed images, and respectively inputting the perturbed images and the transformed time-domain data into the spiking neural network. First, passing through the encoder to obtain a sequence of spiking signals, and then the sequence of spiking signals passes through the spiking neural network module based on TLIF neurons and the output layer to obtain a probability distribution result. Calculating the error of the membrane potential before and after the perturbation through the membrane potential perturbation error, and using this error as an index, combined with the neuron training algorithm, backpropagating to train the spiking neural network.
[0016] The second aspect of the embodiments of the present invention provides a system for implementing the above-mentioned spiking neural network image processing method based on the frequency-domain filtering module, which is characterized by including: A data receiving module, configured to receive time-domain image data; A frequency-domain conversion module, configured to convert the time-domain image data into frequency-domain data through discrete Fourier transform; A frequency-domain filtering module, configured to use the frequency-domain filtering module to process the frequency-domain data to remove redundant information therein and enhance the features corresponding to key frequency components, generating filtered frequency-domain data; wherein, the frequency-domain filtering module is implemented by a Gaussian-type frequency selective filter; A time-domain conversion module, configured to convert the filtered frequency-domain data back into time-domain data through inverse discrete Fourier transform to obtain the transformed time-domain data; A spiking neural network module is used to process the converted time-domain data using a spiking neural network for image feature extraction and classification, and output the final image classification result. Among them, the spiking neural network includes an encoder, a spiking neural network module based on TLIF neurons, and an output layer. The encoder is used to convert the input converted time-domain data into a spike signal that can be processed by TLIF neurons. The output of the encoder is then passed to the spiking neural network module based on TLIF neurons and the output layer to obtain the image classification result.
[0017] The beneficial effects of the present invention are as follows. The present invention converts the input image data from the time domain to the frequency domain through a frequency-domain filtering module, and uses a Gaussian-type frequency-selective filter to remove redundant frequency components and noise, enhance the key frequency components in the signal, and optimize the image feature expression. The frequency-domain data processed by the frequency-domain filtering module of the present invention is subjected to feature extraction and classification through a spiking neural network model, and the final classification result is output. The frequency-domain processing of the present invention can not only improve the classification accuracy of the image, but also reduce the average spike firing rate of the spiking neural network model by reducing redundant inputs, effectively remove noise and redundant components, enabling the spiking neural network model to focus on the key features in the image, thereby effectively reducing energy consumption. By optimizing the bandwidth and selectivity of the frequency-selective filter, the present invention is conducive to enhancing the frequency feature expression ability of the image signal, thereby enhancing the accuracy and efficiency of image processing. By reducing data redundancy and improving the expression ability of signal frequency features, the present invention optimizes the accuracy and efficiency in the image processing process. Through TLIF, the model is more flexible in dealing with subtle perturbations, making the overall model more robust. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a flowchart of the spiking neural network image processing method based on the frequency-domain filtering module of the present invention; Figure 2 is a schematic diagram of the network architecture of the spiking neural network implemented by the Spiking VGG architecture of the present invention; Figure 3 is a schematic diagram of the network architecture of the spiking neural network implemented by the MS-ResNet architecture of the present invention; Figure 4 is a schematic diagram of the network architecture of the spiking neural network implemented by the SpikiFormer architecture of the present invention; Figure 5 is a schematic diagram of the network architecture of the spiking neural network of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0019] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims. It should be understood that the above general description and the following detailed description are exemplary and explanatory only and do not limit the present application.
[0020] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "the", and "said" used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0021] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining". Moreover, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.
[0022] The present invention will be described in detail below with reference to the accompanying drawings. Without conflict, the features in the following embodiments and implementation manners can be combined with each other.
[0023] See Figure 1 , the method for processing images of a pulsed neural network based on a frequency domain filtering module of the present invention specifically includes the following steps: (1) Obtain the time-domain image data x(t), where t represents time.
[0024] It should be understood that the time-domain image data x(t) can be obtained from publicly available datasets, such as the CIFAR10-DVS dataset, the DVS-Gesture dataset, etc. Among them, the CIFAR10-DVS dataset is an event-based neuromorphic dataset, a neuromorphic vision dataset obtained by capturing moving images on a monitor, which contains a total of 10 categories consisting of 10,000 event streams. For each selected static picture, it is made to perform repetitive closed-loop motion, and the events during the motion are recorded by a neuromorphic sensor. The DVS-Gesture dataset is a neuromorphic vision dataset, which is obtained by capturing different gestures of 29 subjects under 3 lighting conditions. The gesture samples in it contain 1,176 event streams from 23 training subjects and 288 event streams from 6 test subjects. These 1,464 event streams together form 11 classes, and the original resolution size is 128×128 for all of them.
[0025] (2) Convert the time-domain image data x(t) to frequency-domain data X(f) through the discrete Fourier transform (DFT).
[0026] Further, the expression of the frequency-domain data X(f) is:
[0027] In the formula, represents the frequency-domain data, f represents the frequency variable, represents the time-domain data, n represents the nth time step, and N is the signal length of the time-domain data.
[0028] It should be noted that when processing the time-domain image data x(t), a two-dimensional Fourier transform is also performed on the time-domain image data x(t), which is expressed as:
[0029] Usually, the basis functions of the two-dimensional (2D) discrete cosine transform (DCT) are:
[0030] Correspondingly, the inverse 2D DCT can be written as:
[0031] In the formula, represents the numerical value of the time-domain image data at the position in the two-dimensional time-domain image data. H and W respectively represent the height and width of the two-dimensional time-domain image data, and i and j respectively represent the horizontal and vertical positions in the two-dimensional time-domain image data; represents in the frequency spectrum after the two-dimensional Fourier transform processing The numerical value of the frequency-domain image data at the position, where h and w respectively represent the height and width positions in the frequency spectrum; represents the two-dimensional Fourier transform.
[0032] (3) Use the frequency-domain filtering module to process the frequency-domain data X(f) to remove redundant information therein and enhance the features corresponding to key frequency components, generating the filtered frequency-domain data X′(f). Among them, the frequency-domain filtering module is implemented by a frequency-selective filter. By the frequency-domain filtering module, noise components (such as low-frequency noise) in the frequency domain can be removed and high-frequency effective components can be enhanced, which helps to improve the signal-to-noise ratio of the spiking neural network in the subsequent steps and enables it to better process the detailed information in the image data.
[0033] Furthermore, the frequency-selective filter is used to remove unnecessary frequency components (such as redundant information like noise) in the frequency-domain data X(f) and enhance the key frequency components in the frequency-domain data X(f). By implementing the function of the frequency-domain filtering module with a frequency-selective filter, the redundancy of the input data can be processed in the frequency domain. It can not only remove noise, enhance the robustness and performance of the model, but also make the input data sparser, thereby reducing energy consumption. Through frequency-domain filtering processing, the frequency-domain filtering module can effectively reduce the redundant part in the input signal and improve the calculation efficiency and accuracy of the SNN.
[0034] Preferably, the frequency-selective filter is a Gaussian-type frequency-selective filter, and its frequency response function is:
[0035] In the formula, represents the frequency response function of the Gaussian-type frequency-selective filter, represents the frequency component in the horizontal direction, represents the frequency component in the vertical direction, and represent the horizontal direction coordinate and vertical direction coordinate at the central position of the frequency response function in the frequency domain, represents the bandwidth of the Gaussian-type frequency-selective filter.
[0036] Specifically, for different frequencies, the Gaussian-type frequency-selective filter gives different coefficients. The coefficients corresponding to high frequencies are smaller, and the coefficients corresponding to low frequencies are larger. Since noise is often located at high frequencies, while information such as image contours is often located at low frequencies, therefore, by screening out high frequencies, the filtered frequency-domain data X′(f) is obtained. Using the above frequency response function, according to the of the frequency component, can be directly calculated, and the numerical value of The coefficient corresponding to a certain time, multiplying the corresponding frequency-domain data by this coefficient (the coefficient is a number between 0 and 1), can achieve frequency-domain filtering and obtain the filtered frequency-domain data X′(f). Based on the characteristic that image noise often distributes in the high frequency, therefore, in addition to the Gaussian frequency-selective filter, the frequency-domain filtering module can also select other filters such as Low-Pass filter, Butterworth filter, Band-Pass filter, Chebyshev filter, etc.
[0037] It should be understood that, according to the noise level and frequency characteristics of the input image, the bandwidth of the Gaussian frequency-selective filter is adaptively and dynamically adjusted , which can adapt to the noise characteristics of different input images, ensure that while removing noise, the key information of the image is retained, thereby improving the noise suppression ability of the frequency-selective filter and optimizing the signal processing effect, and optimizing the filtering effect. By adjusting and optimizing the bandwidth of the Gaussian frequency-selective filter , the frequency-selective filter can retain the important frequency components in the input signal to a great extent, enhance the frequency components in the image that contribute greatly to the classification result, and at the same time remove the frequency components with greater interference, such as noise and redundant frequency components, improving the overall effect of image processing and improving the processing accuracy of the subsequent spiking neural network. The bandwidth Specifically, it can be adjusted according to the relationship between the bandwidth of the frequency-selective filter and the image characteristics and noise level.
[0038] It should be noted that the method and system of the present invention optimize the accuracy and efficiency in the image processing process by reducing data redundancy and improving the expression ability of signal frequency characteristics: ① Specifically, the frequency-domain filtering module effectively removes noise and redundant components, reduces data redundancy, enabling the spiking neural network to focus on the key features in the image. ② By optimizing the bandwidth and selectivity of the frequency-selective filter, the expression ability of signal frequency characteristics is improved, and the expression ability of the frequency characteristics of the image signal is enhanced, thereby enhancing the accuracy and efficiency of image processing.
[0039] (4) Convert the filtered frequency-domain data X′(f) back to the time-domain data x′(n) through the inverse discrete Fourier transform (IDFT) to obtain the converted time-domain data x′(n).
[0040] Furthermore, the expression of the converted time-domain data x′(n) is:
[0041] In the formula, represents the converted time-domain data, represents the filtered frequency-domain data.
[0042] (5) Process the transformed time-domain data x′(n) using a spiking neural network for image feature extraction and classification, and output the final image classification result. Among them, in the spiking neural network, learning and prediction are carried out by simulating the membrane potential dynamic changes and spike generation mechanism of TLIF neurons. The spiking neural network includes an encoder, a spiking neural network module based on TLIF neurons, and an output layer. The output layer includes a fully connected layer. The encoder is used to convert the input transformed time-domain data into spike signals that can be processed by TLIF neurons. The output of the encoder is then transmitted to the spiking neural network module based on TLIF neurons and the output layer to obtain the image classification result.
[0043] It should be understood that the spiking neural network has a specific encoder, and different types of data have suitable encoding methods to generate spike sequences based on time steps, which are the so-called time-domain data. Taking pictures as an example, with Poisson encoding, a certain value will calculate a probability, and then randomly generate spikes within eight time steps according to the probability, that is, a 01 sequence, which is the so-called time-domain image data. Then 800 time steps can represent the original 100 values.
[0044] In this embodiment, the spiking neural network is used to extract image features from the input transformed time-domain data x′(n) and output the image classification result. The spiking neural network module based on TLIF neurons includes a downsampling layer and multiple convolutional spiking neural network module output layers. Among them, the convolutional spiking neural network module adopts the Spiking VGG architecture, the MS-ResNet architecture, or the SpikiFormer architecture. Specifically, the downsampling layer is used to unify the image format of the input transformed time-domain data x′(n) through downsampling operations to obtain unified image data; then the unified image data is sent into the convolutional spiking neural network module for feature extraction to obtain image features. The convolutional spiking neural network module extracts features through the synaptic connections between TLIF neurons and their membrane potential dynamic changes, and effectively extracts features after frequency-domain filtering of the input image signal; finally, the extracted image features pass through the output layer to output the final image classification result, and the output layer classifies the image according to the spike generation pattern of TLIF neurons.
[0045] It should be understood that the spike generation pattern refers to the processing of different data in the model, which can be understood as: data propagates in the spiking neural network, and activating different neurons means different data itself. As the data propagates in the spiking neural network, its features will be extracted, and the image is classified according to the spike generation pattern (the performance of the extracted features on the spike generation of TLIF neurons).
[0046] It should be noted that different mainstream SNN architectures can be connected after the frequency domain filtering module, and the content is not limited. The output layer is the last layer and is a fully connected layer. For example, in this embodiment, a downsampling layer, different SNN architecture modules (such as Spiking VGG architecture, MS-ResNet architecture, or SpikiFormer architecture, etc.) and an output layer are connected after the frequency domain filtering module to implement an image classification task.
[0047] Furthermore, as Figure 2 shown, the Spiking VGG architecture includes multiple convolutional layers, multiple batch normalization layers, and multiple TLIF neurons. The image data after format unification is sent into the convolutional spiking neural network module, and sequentially passes through the convolutional layer, batch normalization layer, and TLIF neurons to process the image data. This Spiking VGG architecture realizes feature learning by simulating the time integration process and pulse triggering mechanism of TLIF neurons, so as to extract the features of the image data and obtain image features. The Spiking VGG architecture can achieve efficient image classification tasks by performing layer-by-layer convolution and batch normalization, and combining the unique pulse response mechanism of SNN for image feature extraction and classification.
[0048] Furthermore, as Figure 3 shown, the MS-ResNet architecture includes multiple TLIF neurons, multiple convolutional layers, and multiple batch normalization layers, and combines residual connections and multi-scale convolutional layers to improve the robustness of the extracted image features through parallel processing and skip connections, and optimize its performance in the spiking neural network. The image data after format unification is sent into the convolutional spiking neural network module, and sequentially passes through the TLIF neurons, convolutional layer, batch normalization layer, TLIF neurons, convolutional layer, and batch normalization layer to process the image data, extract the image features of the current scale, and fuse the image features of the previous scale with the image features of the current scale to obtain the fused features of the current scale, and finally obtain the multi-scale fused image features. The MS-ResNet architecture combines multi-scale feature extraction and residual learning, can adapt to different frequency features of the input image, and enhance the accuracy of image classification.
[0049] Furthermore, as Figure 4As shown, the SpikiFormer architecture is based on the self-attention mechanism and uses an adaptive weighting method to extract various features in the image. Its core operation is to add the original data weighted to the processed data. The SpikiFormer architecture includes multiple separable convolutional layers, multiple channel convolutional layers, downsampling layers, multiple impulse-driven self-attention layers, and multiple channel multi-layer perceptrons. The uniformly formatted image data is fed into the convolutional pulse neural network module. First, it passes through the separable convolutional layer. After the features output by the separable convolutional layer are weighted and fused with the features input to the separable convolutional layer, they are input into the channel convolutional layer. After the features output by the channel convolutional layer are weighted and fused with the features input to the channel convolutional layer, they are sent to the next separable convolutional layer. After the features output by the last channel convolutional layer are weighted and fused with the features input to this channel convolutional layer, they are input into the downsampling layer. The features output by the downsampling layer are then input into the impulse-driven self-attention layer. After the features output by the impulse-driven self-attention layer are weighted and fused with the features input to this impulse-driven self-attention layer, they are sent to the channel multi-layer perceptron. After the features output by the channel multi-layer perceptron are weighted and fused with the features input to this channel multi-layer perceptron, they are sent to the next impulse-driven self-attention layer. After the features output by the last channel multi-layer perceptron are weighted and fused with the features input to this channel multi-layer perceptron, the final image features can be obtained. The SpikiFormer architecture adopts the self-attention mechanism and uses the image features processed in the frequency domain to optimize the classification task.
[0050] Furthermore, the Leaky Integrate and Fire (LIF) neuron describes three basic actions of a neuron: Leaky, Integrate, and Fire. Initially, the LIF neuron maintains a resting potential. , as the current varying with time flows in, the charge in the capacitor C gradually accumulates, causing the membrane potential to gradually rise. Once the membrane potential exceeds the set threshold, the LIF neuron generates an action potential, and then the membrane potential is reset to the resting potential. . During this process, the effects of the membrane resistance and ion channels can be represented by the resistance R. The differential equation of the LIF neuron is:
[0051] In the formula, represents the membrane potential of the LIF neuron at time t, is the resting potential, represents the input of the LIF neuron at time t, is the membrane potential time constant, which affects the amplitude of the input pulse and the decay rate of the membrane potential.
[0052] It should be noted that the TLIF neuron simulates the impulse firing behavior of neurons through a dynamic membrane potential model. Different from the traditional LIF neuron with a fixed membrane potential leakage factor λ, the TLIF neuron has stronger adaptability and can learn to adjust the optimal λ value of the neuron according to the change of the input signal, making the change of the membrane potential more flexible and reducing the error caused by disturbance. This flexibility enables the TLIF neuron to perform more stably and efficiently when processing complex input signals such as designed misleading data. Therefore, the TLIF neuron can better improve the learning ability and robustness of the neural network.
[0053] TLIF is based on the LIF model, and the membrane potential of the neuron is:
[0054] where represents the membrane potential of the i-th neuron at time t in layer represents the leakage factor of the i-th neuron in layer represents the remaining membrane potential of the i-th neuron at time t - 1 in layer represents the j-th neuron in layer for the synaptic weight of the i-th neuron in layer represents the bias, represents the j-th neuron in layer
[0055] When the membrane potential exceeds the threshold, an impulse is fired, and the neuron emits an impulse as:
[0056] where represents the impulse output of the i-th neuron at time t in layer is the step function.
[0057] After emitting an impulse, the electrical position is set to 0, and after simplification, the remaining membrane potential of the neuron is:
[0058] TLIF is optimized and improved on the basis of STBP, and both the spatial domain dimension and the time domain dimension are considered during the backpropagation process. In the spatial domain dimension, two cases and are considered, where L is the neuron output layer, that is, the cases where the neuron is located in the middle layer and the output layer are considered. In the time domain dimension, t = T and t < T are considered, where T is the time window value, that is, the middle time period and the last time period of training are considered. The backpropagation formulas for different cases are as follows, where denotes the pre - synaptic membrane potential of the $i$-th neuron in layer $l$ at time $t$, denotes the firing situation of the $i$-th neuron in layer $l$ at time $t$, the firing value is taken as 1, otherwise 0. denotes the synaptic weight of the $j$-th neuron in layer $l - 1$ with respect to the $i$-th neuron in layer $l$, where $L$ in it represents the loss function.
[0059] When and $t = T$, in this case, the neuron is located in the output layer, the derivative can be directly obtained as:
[0060]
[0061] When and $t = T$, in this case, the neuron is located in the middle layer and in the last time period of training, so it is necessary to back - propagate in the spatial domain to obtain the following formula:
[0062]
[0063] When and $t\lt T$, in this case, in the middle time period of training, it is necessary to back - propagate in the time domain, including two parts, directly taking the derivative of and propagating from the direction to obtain the formula:
[0064]
[0065] When and $t\lt T$, in this case, the neuron is located in the middle layer and in the middle time period of training, it is necessary to decompose the propagation into the time domain and the spatial domain. The neurons in the middle layer accumulate the weighted error signals from the subsequent layer and update the parameters iteratively in different layers:
[0066]
[0067] Based on the above four cases, back - propagation can be carried out in the neural network. Back - propagation is decomposed into two parts: the time domain and the spatial domain. Neurons accumulate the weighted error and update iteratively. Finally, the derivatives of the weight matrix $W$, the bias matrix $b$, and the leakage function matrix $\lambda$ are shown in the formula:
[0068] Among them, the TLIF can be effectively trained by the gradient descent optimization algorithm so far.
[0069] After the output of the TLIF layer, the network continues to transfer the information of the layer before the output layer to the membrane potential perturbation error. The membrane potential perturbation error plays an important role in the training process. By backpropagation, the adjustment parameters include the λ value of the neuron membrane potential to reduce the distance between the attack sample result and the correct result. The membrane potential perturbation error can ensure that the change of the neuron membrane potential will not be overly perturbed by the attack, thus ensuring the stability of the neural network in the face of noise or adversarial attacks. This regularization mechanism effectively improves the robustness of the network, enabling it to maintain good classification performance in complex environments.
[0070] Among them, the formula for the membrane potential perturbation error (MPPE) is:
[0071] Among them, ( ) represents the difference in the membrane potential before and after perturbation, and the symbol with a tilde superscript represents the perturbed version of the original variable. When t = 0, and are both zero. represents the weight from the j-th neuron in the (l - 1)-th layer to the i-th neuron in the l-th layer.
[0072] Finally, the output layer of the network classifies based on the processing results of the neurons. The classification result predicted by the model is compared with the actual label, the task loss Ltask is calculated, and MPPE is introduced to enhance the model's perception ability of perturbations. Finally, a comprehensive loss function is constructed for training optimization. In the image classification task, after the aforementioned processing, the network can successfully identify the category of the image. As Figure 5 shown, taking the "cat" image in the example as an example, even if perturbations designed for the network are applied, the network can finally successfully identify the "cat" category in the image, indicating that this method has high accuracy and robustness in dealing with complex image classification tasks.
[0073] Furthermore, the training process of the spiking neural network specifically includes: As Figure 5As shown, perturbations are applied to the converted time-domain data to obtain perturbed images. The perturbed images and the converted time-domain data are respectively input into the spiking neural network. First, a spike signal sequence is obtained through the encoder, where T represents the time step, B represents the batch size, and N represents the number of neurons. Subsequently, the spike signal sequence passes through the spiking neural network module based on TLIF neurons and the output layer to obtain the probability distribution result; the error of the membrane potential before and after perturbation is calculated through the membrane potential perturbation error, and taking this error as an index, combined with the neuron training algorithm (Spike-Timing-Dependent Backpropagation, STBP), the spiking neural network is trained by backpropagation.
[0074] Among them, during training, the loss function adds a membrane potential perturbation error regularization term on the basis of STBP to minimize the perturbation error. The filtered image and the added noise are together used as the input and passed to the encoder. The role of the encoder is to convert the features of the image into a form of spike signals suitable for processing by the spiking neural network. The model adopts a direct encoding method. Through the encoder, the important features in the image are converted into spike signals that can be processed by TLIF. The output after passing through the encoder is then transmitted to the TLIF neuron layer. The TLIF neuron simulates the spike firing behavior of neurons through a dynamic membrane potential model. Different from the traditional LIF neurons with a fixed membrane potential leakage factor λ, the TLIF neurons have stronger adaptability and can learn to adjust the optimal λ value of the neurons according to the changes of the input signals, making the change of the membrane potential more flexible and reducing the error caused by perturbations. Finally, the classification result is obtained through the output layer.
[0075] TLIF sets the leakage factor λ in LIF as a trainable parameter, and each neuron in different neural network layers has a different λ value. The neural network can find the optimal historical dependence degree between different training stages or tasks, thereby improving the ability to process diverse input data. During training, λ can be dynamically adjusted, so that the neural network can perform more effective learning and adaptation when facing data with different perturbations.
[0076] Based on the analysis of the membrane potential perturbation error, the membrane potential perturbation error calculates the error of the membrane potential of neurons in the neural network between the perturbed input data and the original input data, and quantifies the specific impact of the perturbation on the model. During training, the error value is reduced by adjusting the neural network parameters to minimize the impact of the perturbation on the model, so that the model has a finer-grained perception of the perturbation and has a better robustness optimization effect on the classification result of the neural network.
[0077] For the input data with perturbations, first, the noise is removed through the filter module, and the information density of the input data is optimized. The processed data has better stability against noise, but for carefully designed attacks such as PGD and FGSM, more refined processing by the neural network is required. Subsequently, TLIF is used in the neural network structure to make the leakage parameters trainable, thereby enhancing the flexibility of the neural network against attacks. During the training process, the membrane potential perturbation error is used as a regularization term, noise is added to the input, and the error is quantized to minimize the impact of noise.
[0078] It is worth mentioning that the embodiment of the present invention also provides a spiking neural network image processing system based on a frequency domain filtering module for implementing the spiking neural network image processing method based on the frequency domain filtering module in the above embodiment. The system includes a data receiving module, a frequency domain conversion module, a frequency domain filtering module, a time domain conversion module, and a spiking neural network module.
[0079] In this embodiment, the data receiving module is used to receive the time domain image data x(t).
[0080] In this embodiment, the frequency domain conversion module is used to convert the time domain image data x(t) into frequency domain data X(f) through discrete Fourier transform.
[0081] In this embodiment, the frequency domain filtering module is used to process the frequency domain data X(f) to remove redundant information therein and enhance the features corresponding to key frequency components, generating filtered frequency domain data X′(f). Among them, the frequency domain filtering module is implemented by a Gaussian-type frequency selective filter.
[0082] In this embodiment, the time domain conversion module is used to convert the filtered frequency domain data X′(f) back into time domain data x′(n) through inverse discrete Fourier transform (IDFT), obtaining the converted time domain data x′(n).
[0083] In this embodiment, the spiking neural network module is used to process the converted time domain data x′(n) using a spiking neural network for image feature extraction and classification, and output the final image classification result. Among them, the spiking neural network module learns and predicts by simulating the membrane potential dynamic change and spike generation mechanism of TLIF neurons. The spiking neural network includes an encoder, a spiking neural network module based on TLIF neurons, and an output layer. The encoder is used to convert the input converted time domain data into spike signals that can be processed by TLIF neurons, and the output of the encoder is then transmitted to the spiking neural network module based on TLIF neurons and the output layer to obtain the image classification result.
[0084] It should be noted that the spiking neural network module extracts key feature information from the spike firing patterns of spiking neurons through operations such as max pooling, thereby improving the classification accuracy; selecting the most representative spike pattern according to the firing frequency and intensity of spiking neurons helps to enhance the robustness of classification; the extraction of spike firing patterns selectively activates spiking neurons that contribute to classification by analyzing the change patterns of the membrane potential of spiking neurons, thereby obtaining accurate image classification results.
[0085] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A pulse neural network image processing method based on a frequency domain filtering module, characterized in that: The following steps are involved: (1) Obtaining time domain image data; (2) Convert the time domain image data into frequency domain data through discrete Fourier transform; (3) Processing the frequency domain data using a frequency domain filtering module to remove redundant information and enhance the features corresponding to the key frequency components, thereby generating filtered frequency domain data; wherein the frequency domain filtering module is implemented using a Gaussian frequency selective filter; (4) Converting the filtered frequency domain data back to time domain data through inverse discrete Fourier transform to obtain converted time domain data; (5) Using a spiking neural network to process the converted time domain data to extract and classify image features, and output a final image classification result; wherein the spiking neural network includes an encoder, a spiking neural network module based on TLIF neurons, and an output layer, and the encoder is used to convert the input converted time domain data into a pulse signal that can be processed by the TLIF neurons, and the output of the encoder is then passed to the spiking neural network module based on TLIF neurons and the output layer to obtain an image classification result.
2. The pulse neural network image processing method based on the frequency domain filtering module according to claim 1 is characterized in that: In step (2), the frequency domain data is expressed as: ; In the formula, represents frequency domain data, f represents frequency variable, Represents time domain data, n represents the nth time step, and N is the signal length of the time domain data.
3. The pulse neural network image processing method based on the frequency domain filtering module according to claim 1 is characterized in that: The frequency response function of the Gaussian frequency selective filter is: ; In the formula, represents the frequency response function of the Gaussian frequency selective filter, represents the frequency component in the horizontal direction, represents the frequency component in the vertical direction, and Represents the horizontal and vertical coordinates of the center position of the frequency response function in the frequency domain. Represents the bandwidth of the Gaussian frequency selective filter.
4. The pulse neural network image processing method based on the frequency domain filtering module according to claim 1 is characterized in that: In step (4), the expression of the converted time domain data is: ; In the formula, represents the converted time domain data, Represents the frequency domain data after filtering.
5. The pulse neural network image processing method based on frequency domain filtering module according to claim 1 is characterized in that: The spiking neural network module based on TLIF neurons includes a downsampling layer and a plurality of convolutional spiking neural network modules, wherein the convolutional spiking neural network module adopts a Spiking VGG architecture, an MS-ResNet architecture or a SpikiFormer architecture; The downsampling layer is used to unify the image format of the input converted time domain data x′(n) through a downsampling operation to obtain image data with a unified format; the image data with a unified format is then sent to a convolutional pulse neural network module for feature extraction to obtain image features, and the convolutional pulse neural network module extracts features through the synaptic connections between TLIF neurons and the dynamic changes of their membrane potential; finally, the extracted image features are passed through the output layer and the final image classification result is output, and the output layer performs image classification according to the pulse emission pattern of the TLIF neurons.
6. The pulse neural network image processing method based on frequency domain filtering module according to claim 5 is characterized in that: The Spiking VGG architecture includes multiple convolutional layers, multiple batch normalization layers and multiple TLIF neurons. The image data with unified format is sent to the convolutional spike neural network module, and passes through the convolutional layer, batch normalization layer and TLIF neurons in sequence to process the image data. The Spiking VGG architecture realizes feature learning by simulating the time integration process of TLIF neurons and the pulse triggering mechanism, so as to realize feature extraction of image data and obtain image features.
7. The pulse neural network image processing method based on frequency domain filtering module according to claim 5 is characterized in that: The MS-ResNet architecture includes multiple TLIF neurons, multiple convolutional layers and multiple batch normalization layers, and is combined with residual connections and multi-scale convolutional layers. The image data with unified format is sent to the convolutional pulse neural network module, and passes through TLIF neurons, convolutional layers, batch normalization layers, TLIF neurons, convolutional layers and batch normalization layers in sequence to process the image data, extract the image features of the current scale, and fuse the image features of the previous scale with the image features of the current scale to obtain the fused features of the current scale, and finally obtain the multi-scale fused image features.
8. The pulse neural network image processing method based on frequency domain filtering module according to claim 5 is characterized in that: The SpikiFormer architecture is based on the self-attention mechanism and uses an adaptive weighted method to extract various features in the image. The SpikiFormer architecture includes multiple separable convolutional layers, multiple channel convolutional layers, downsampling layers, multiple pulse-driven self-attention layers, and multiple channel multi-layer perceptrons. The image data with unified format is sent to the convolutional spike neural network module, first passing through the separable convolutional layer, and the features of the separable convolutional layer output are weighted fused with the features of the input separable convolutional layer, and then input to the channel convolutional layer, and the features of the channel convolutional layer output are weighted fused with the features of the input channel convolutional layer, and then sent to the next separable convolutional layer. Separate convolution layer; the features output by the last channel convolution layer are weighted fused with the features input to the channel convolution layer and then input to the downsampling layer; the features output by the downsampling layer are then input to the pulse-driven self-attention layer, the features output by the pulse-driven self-attention layer are weighted fused with the features input to the pulse-driven self-attention layer and then sent to the channel multi-layer perceptron, the features output by the channel multi-layer perceptron are weighted fused with the features input to the channel multi-layer perceptron and then sent to the next pulse-driven self-attention layer; the features output by the last channel multi-layer perceptron are weighted fused with the features input to the channel multi-layer perceptron to obtain the final image features.
9. The pulse neural network image processing method based on frequency domain filtering module according to claim 1 is characterized in that: The training process of the pulse neural network specifically includes: applying perturbation to the converted time domain data to obtain a perturbed image, inputting the perturbed image and the converted time domain data into the pulse neural network respectively, firstly passing through an encoder to obtain a pulse signal sequence, and then passing the pulse signal sequence through a pulse neural network module based on TLIF neurons and an output layer to obtain a probability distribution result; calculating the error of the potential before and after the perturbation through the membrane potential perturbation error, and using the error as an indicator to train the pulse neural network in combination with the back propagation of the neuron training algorithm.
10. A system for implementing the pulse neural network image processing method based on a frequency domain filtering module according to any one of claims 1 to 9, characterized in that: include: A data receiving module, used for receiving time domain image data; A frequency domain conversion module, used for converting time domain image data into frequency domain data through discrete Fourier transform; A frequency domain filtering module is used to process the frequency domain data using the frequency domain filtering module to remove redundant information therein and enhance the features corresponding to the key frequency components, thereby generating filtered frequency domain data; wherein the frequency domain filtering module is implemented using a Gaussian frequency selective filter; A time domain conversion module, used for converting the filtered frequency domain data back to time domain data through inverse discrete Fourier transform to obtain converted time domain data; A pulse neural network module is used to process the converted time domain data using a pulse neural network to extract and classify image features and output the final image classification result; wherein the pulse neural network includes an encoder, a pulse neural network module based on TLIF neurons and an output layer, the encoder is used to convert the input converted time domain data into a pulse signal that can be processed by the TLIF neurons, and the output of the encoder is then passed to the pulse neural network module based on TLIF neurons and the output layer to obtain the image classification result.
Citation Information
Patent Citations
Image edge detection method based on response of cortical neuron in visual direction
CN103345754A
Convolutional neural network-based self-adaptive feature selection target tracking method
CN108288282A
Pulmonary tuberculosis auxiliary decision-making system based on deep convolutional spiking neural network
CN115482230A
Plate surface defect detection system and method based on machine vision
CN119048472A