Image pulse coding method based on multi-scale Census transformation

Through the image pulse coding method based on multi-scale Census transformation, the problems of poor image coding effect and insufficient fusion of multi-scale feature in the prior art are solved, and more efficient image feature extraction and pulse signal generation are achieved, which improves the robustness and recognition accuracy of the pulse neural network.

CN120088344APending Publication Date: 2025-06-03BEIJING INST OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510057670.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art has shortcomings in pulse coding and multi-scale feature fusion of images, and it is difficult to effectively utilize the local and global features of images, especially in the face of factors such as noise and lighting changes, which limits the application of pulsed neural networks in complex image processing tasks.

Method used

An image pulse coding method based on multi-scale Census transformation is used to generate a binary pulse signal through color image preprocessing, Census pulse coding, and weighted fusion of multi-scale Census images, and serve as the input of the pulse neural network.

Benefits of technology

It improves the robustness, accuracy and processing efficiency of pulsed neural networks in image processing tasks, and can more effectively extract image details and global information, which is suitable for object recognition and classification tasks in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088344A_ABST
    Figure CN120088344A_ABST
Patent Text Reader

Abstract

The invention discloses an image pulse coding method based on multi-scale Census transformation. The image pulse coding method comprises the steps that graying, denoising, contrast enhancement and normalization processing are carried out on a color image; on a plurality of scales, performing Census pulse coding on the processed image to obtain Census images of different scales, and respectively converting the Census images into decimal images; and carrying out weighted fusion on the obtained decimal images with different scales, then carrying out normalization, converting decimal pixels into binary pulse signals, and generating a complete pulse image from the pulse signals of all the pixels according to a time sequence. According to the invention, by combining multi-scale feature extraction and an effective pulse coding mechanism, the robustness, precision and processing efficiency of the pulse neural network in an image processing task are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to an image pulse coding method based on multi-scale Census transform. Background Art

[0002] With the continuous development of deep learning and neural network technologies, spiking neural networks (SNNs), as a computational model that simulates biological nervous systems, have been widely applied in the processing of time-series data, image recognition, speech processing, and other fields. Different from traditional artificial neural networks (ANNs), spiking neural networks process information by simulating the pulse generation mechanism of neurons, enabling more efficient processing of time-series data, especially showing stronger adaptability and robustness in dynamic scenarios.

[0003] However, a key challenge of spiking neural networks lies in how to effectively convert image data into pulse signals suitable for neural network processing. Existing image processing methods mainly rely on traditional feature extraction and dimensionality reduction techniques, such as using convolutional layers in convolutional neural networks (CNNs) to extract image features, or performing feature compression through other image preprocessing steps (such as edge detection, histogram equalization, etc.). However, these methods usually generate feature vectors in a fixed format and cannot directly meet the requirements of spiking neural networks for time-series data.

[0004] Currently, some studies have attempted to convert image data into pulse signals and input them into spiking neural networks for processing. For example, some methods use an event-driven approach to encode images, converting each pixel value in the image into a pulse generation pattern related to its brightness change. However, these methods usually lack sufficient utilization of image details and global information, resulting in poor encoding effects and difficulty in achieving efficient object recognition and classification tasks in complex scenarios.

[0005] Meanwhile, in terms of image feature extraction, Census transform is a commonly used non-parametric image processing method that has been widely applied in tasks such as edge detection and texture analysis. Census transform encodes the comparison results of each pixel with its neighbors into a binary string as the local feature of the image. This transform is not affected by illumination changes and is suitable for multi-modal intensity distribution regions such as object boundaries. However, Census transform is mostly used for traditional image feature extraction and image matching problems. When directly applied to pulse coding in spiking neural networks, it lacks adaptability to pulse signals, especially in the application of multi-scale feature fusion, which is still not perfect.

[0006] In summary, although certain progress has been made in spiking neural networks and image feature extraction in the prior art, there are still some problems in the spiking coding of images and multi-scale feature fusion. The existing coding methods are insufficient in extracting local and global features of images and are difficult to effectively handle factors such as noise and illumination changes, which limits the application of spiking neural networks in complex image processing tasks. Therefore, how to improve the image processing ability of spiking neural networks through effective image coding methods remains a technical problem to be solved urgently. Summary of the Invention

[0007] In order to overcome the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide an image spiking coding method based on multi-scale Census transform, which can improve the robustness, accuracy and processing efficiency of spiking neural networks in image processing tasks by combining multi-scale feature extraction and an effective spiking coding mechanism.

[0008] In order to achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0009] An image spiking coding method based on multi-scale Census transform includes the following steps:

[0010] Step 1, preprocessing of color images.

[0011] Image preprocessing is a prerequisite for the efficient operation of spiking neural networks, aiming to eliminate noise, enhance important features and unify image intensity, so that the SNN can focus more on key structural information. Specifically, in this step, the color image is grayscaled, denoised, contrast enhanced and normalized.

[0012] Step 2, Census spiking coding.

[0013] In this step, the image preprocessed in Step 1 is subjected to Census spiking coding at multiple scales to obtain Census images at different scales and convert them into decimal images respectively.

[0014] Step 3, weighted fusion of multi-scale Census images.

[0015] In this step, the decimal images at different scales obtained in the previous step are weighted and fused, then normalized, and then the decimal pixels are converted into binary pulse signals. The pulse signals of all pixels generate a complete pulse image in time sequence.

[0016] In one embodiment, the grayscaling process is to convert the input color image into a grayscale image through grayscaling, which not only reduces redundant color information but also reduces processing complexity. The grayscale value is calculated by the weighted average method from the RGB (red, green, blue) components of the image, and the specific calculation formula is:

[0017] I Gray (x,y) = α 1 ·R(x,y) + α 2 ·G(x,y) + α 3 ·B(x,y)

[0018] where R(x,y), G(x,y), and B(x,y) are the red, green, and blue components of the pixel (x,y) respectively, and α 1 , α 2 , α 3 are weighting coefficients.

[0019] In one embodiment, for contrast enhancement and normalization processing, the grayscale image is denoised, enhanced in contrast by Gaussian filtering, and the pixel value range is unified through a normalization operation to ensure the stability and convergence speed of network training. Denoising makes the image clearer and reduces the risk of misidentification, while contrast enhancement makes the details more prominent, which is in line with the characteristics of SNN being sensitive to brightness changes. The normalization operation avoids adverse effects on the consistency of neuron responses caused by excessive brightness differences.

[0020] In one embodiment, in step two, the pixels at each position (x,y) of the image I are processed by local binary coding. The Census transform is performed using a sliding window of size w×w. The central pixel is compared with its neighboring pixels. If the grayscale value of a neighboring pixel is greater than or equal to the central pixel value, the corresponding binary code is 1, otherwise it is 0. Thus, the output binary image is obtained, where each pixel value is represented as a binary code, reflecting the relative relationship between local pixels. The formula for the Census transform is expressed as:

[0021]

[0022] where is the window radius, C(x,y) represents the Census code at position (x,y), which is a binary code string of dimension w 2 - 1, I(x,y) is the pixel value corresponding to position (x,y) in the image I, and i,j represent the coordinate offsets. For example, for the central pixel I(3,3), when i = 1 and j = 1, I(x + 1,y + 1) = I(4,4). Obviously, i and j are not both 0 at the same time. In this way, the Census transform can effectively capture the local texture features of the image while reducing the influence of illumination changes and noise. For pixels near the image edge, the neighborhood window may exceed the image range, and boundary zero-padding processing is required.

[0023] By selecting different window sizes, Census pulse coding at multiple scales can be achieved.

[0024] The output of the Census transform is a binary image, where each pixel value is represented as a binary code, reflecting the relative relationship between local pixels rather than the absolute brightness value. This makes the subsequent input to the spiking neural network more stable and enhances the robustness of image features.

[0025] In one embodiment, for ease of subsequent processing, the binary image after the Census transform is converted into a decimal image. Specifically, the Census code C(x, y) corresponding to each pixel position (x, y) is converted into a Census image value D(x, y), where D(x, y) is a decimal number, and the formula is expressed as:

[0026]

[0027] where C(x, y) k is the k-th bit of the binary code, and N is the length of the binary code. This process converts the binary image into a decimal image, providing a more intuitive numerical representation for subsequent feature fusion and pulse coding.

[0028] In one embodiment, step three is divided into three parts: feature fusion, normalization processing, and pulse conversion. Among them, for feature fusion, since image features are manifested differently at different scales, a single scale may not be able to capture all important information. Therefore, in step two, Census transforms are performed at multiple scales (such as 3×3, 5×5, 7×7, etc.) to extract image features at different levels. The Census transforms are respectively performed on the image to obtain Census images at different scales, such as D 1 (x, y), D 2 (x, y), D 3 (x, y). By performing weighted fusion on these images, the decimal image D f (x, y) is obtained, which is expressed as:

[0029]

[0030] where M is the number of scales, w m is the weight of the m-th scale, and D m (x, y) is the decimal image obtained by converting the Census pulse code of the m-th scale. Weighted fusion ensures that features at different scales are balanced, enabling full expression of both the details and global information of the image.

[0031] For normalization processing, linear normalization is used to map D f (x, y) to [0, 255], which is expressed as:

[0032]

[0033] Among them, D norm (x, y) is the normalized result, and D min and D max are the minimum and maximum pixel values in the fused feature map.

[0034] For pulse conversion, the decimal pixels after fused normalization are converted into eight-bit binary pulse signals. Each pixel of the image is represented as a pulse vector, and the pulse vectors of all pixels together form a complete pulse image, which is used as the input of the SNN for subsequent feature extraction and pattern recognition tasks.

[0035] In order to make full use of the pulse data generated by Census coding, the present invention designs a classification system based on a spiking neural network. The SNN has the ability to process temporal pulse data and can capture the spatial and temporal features of images. The classification system of the present invention includes an image pulse coding module and a spiking neural network, where:

[0036] The image pulse coding module executes the image pulse coding method based on multi-scale Census transform of the present invention to encode a color image into a pulse image. It can be a separate software module built into the system or implemented in a separate hardware unit.

[0037] The spiking neural network includes an input layer, a hidden layer, a fully connected layer, and an output layer.

[0038] The input layer inputs the encoded pulse image into the network, and the input shape is [T, N, C, H, W], where T is the number of time steps, N is the batch size, C is the number of channels, and H, W are the image height and width. The pulse sequence of each pixel reflects the distribution of its brightness information in the time dimension and provides spatio-temporal feature input for the subsequent layers;

[0039] The hidden layer includes 2 layers of spiking neuron layers for extracting spatio-temporal features; the spiking neuron layer is used to simulate the dynamic behavior of biological neurons, and the model adopts LIF neurons (Leaky Integrate-and-Fire).

[0040] The fully connected layer flattens the output of the hidden layer into a one-dimensional vector, and the number of output channels is the number of classification types.

[0041] In one embodiment, the spiking neuron layer extracts local features of the image. The first spiking neuron layer extracts low-level features, and the second spiking neuron layer extracts high-level features;

[0042] The spiking neuron layer simulates the dynamic behavior of biological neurons;

[0043] The fully connected layer maps the pulse signals output by the upper layer into probability values of each classification type through the softmax function.

[0044] Compared with the prior art, the present invention improves the robustness, stability, and recognition accuracy of image processing using a spiking neural network, and can be used for computer vision tasks such as image recognition and object detection, especially suitable for tasks that need to process dynamic image data and are highly sensitive to spatio-temporal features. Description of the Drawings

[0045] Figure 1 The figure shows a flowchart of the image pulse coding process based on multi-scale Census transform proposed by the present invention.

[0046] Figure 2 The figure shows a schematic diagram of recognizing handwritten digits using the constructed classification system in an embodiment of the present invention.

[0047] Figure 3 The figure shows a comparison diagram of processing the MINIST dataset using the image pulse coding method based on multi-scale Census transform proposed by the present invention.

[0048] Figure 4 The figure shows a comparison diagram of the change trends of the training accuracy and test accuracy in 30 rounds of training in an embodiment of the present invention. Detailed Embodiment

[0049] The following describes the embodiments of the present invention in detail with reference to the drawings and embodiments.

[0050] In view of the deficiencies of traditional image processing and neural network methods in processing complex image data, the present invention proposes an image pulse coding method based on multi-scale Census transform to obtain a pulse image, and further uses the obtained pulse image as the input of a spiking neural network (SNN) for recognition and classification. The image pulse coding method of the present invention aims to extract the details and global information of the image through an accurate coding mechanism, so as to optimize the performance of the SNN in image processing tasks. This method combines steps such as grayscale conversion, denoising, contrast enhancement, normalization, multi-scale feature extraction, pulse coding generation, and feature fusion to ensure the efficient representation of the input data, and significantly improves the robustness, stability, and accuracy of the network.

[0051] The following is the specific implementation process of the image pulse coding method based on multi-scale Census transform of the present invention. Combining the practices of the present invention in actual operations, the specific implementation of each step is described step by step.

[0052] Step 1: Image acquisition and selection.

[0053] 1. Select a dataset.

[0054] In the embodiments of the present invention, the MNIST handwritten digit dataset is selected as the experimental dataset. The MNIST dataset contains 60,000 training images and 10,000 test images. Each image has a size of 28×28 pixels, and the grayscale value ranges from 0 to 255. In the embodiments of the present invention, images containing the digit "1" are selected from this dataset as example data for testing.

[0055] 2. Select the images of the digit "1".

[0056] The present invention selects images of the digit "1" (with a size of 28×28 pixels) from the MNIST training set and uses them as test images for subsequent processing. The grayscale value of each image ranges from 0 to 255, representing the brightness information of the pixels. The digit "1" has clear edges and obvious structural features, which are suitable for verifying the encoding method proposed by the present invention.

[0057] Step 2: Image preprocessing.

[0058] The image preprocessing step is used to ensure the consistency of the images input into the spiking neural network (SNN) and eliminate noise and unnecessary information. Through image preprocessing, the present invention can enhance the key features of the images and unify the input format of the images. The present invention takes the following operations:

[0059] 1. Grayscale processing.

[0060] Since the images in the MNIST dataset are already grayscale images, the present invention omits the grayscale step for color images. If the input image is a color image, it is necessary to convert the RGB image to a grayscale image by methods such as weighted average method to reduce the redundancy of color information.

[0061] 2. Denoising.

[0062] Images may be affected by noise during shooting and transmission, which affects the quality of the images and the accuracy of feature extraction. Therefore, the present invention uses methods such as Gaussian filtering to remove the noise in the images. By smoothing the images and reducing the random noise in the images, the present invention can obtain clearer images, thereby improving the effects of subsequent feature extraction and encoding.

[0063] 3. Contrast enhancement and normalization.

[0064] In order to improve the expressiveness of image details, the present invention performs contrast enhancement operations on the images, making the digit "1" more prominent. Usually, histogram equalization or linear contrast stretching methods are used to enhance the image contrast. Subsequently, the present invention normalizes the pixel values of the images to the range of [0,1] to eliminate the influence caused by brightness differences between different images. This ensures that the neural network can better process and learn the key information in the images.

[0065] Step 3: Census transform.

[0066] The Census transform extracts the texture features of an image through local binary patterns, such that the feature of each pixel is represented as a binary code. Refer to Figure 1 As shown, in the embodiments of the present invention, sliding windows of 3×3, 5×5, and 7×7 are respectively used to process each pixel, and three binary codes are generated to represent the relative relationship between the pixel and its neighboring pixels. The specific steps are as follows:

[0067] 1. Local binary pattern extraction.

[0068] For each pixel at position (x, y) in the image I after the previous preprocessing (corresponding pixel value is I(x, y)), the local binary coding process of the present invention specifically uses a sliding window of size w×w for Census transform processing. By comparing the central pixel with its neighboring pixels, if the gray value of the neighboring pixel is greater than or equal to the central pixel value, it is marked as 1, that is, the corresponding binary code is 1, otherwise it is marked as 0, that is, the corresponding binary code is 0, thus obtaining the output binary image. Each pixel value is represented as an 8-bit binary code, reflecting the relative relationship between local pixels.

[0069] 2. Conversion of binary code to decimal.

[0070] Convert the binary code of each pixel into a decimal value, which is convenient for subsequent feature fusion and pulse code generation. This conversion is achieved by calculating the decimal value of the binary number. This conversion helps to reduce the computational complexity while retaining the necessary feature information.

[0071] For example, assume there is a 3×3 gray image block, and its gray values are as follows:

[0072]

[0073] (1) Select the central pixel

[0074] The central pixel is located at coordinates (x, y) = (2, 2), and the pixel value I(x, y) = 100.

[0075] (2) Construct the neighborhood window

[0076] For a 3×3 window, the neighboring pixels include the 8 pixels around the central pixel.

[0077] (3) Compare the neighboring pixels with the central pixel

[0078] For each pixel I(x + i, y + j) in the neighborhood, compare its gray value with the gray value of the central pixel I(x, y).

[0079] (4) Neighborhood pixels and comparison results:

[0080] Upper pixel: I(1,1) = 100 ≥ 100 => 1; I(2,1) = 102 ≥ 100 => 1; I(3,1) = 101 ≥ 100 => 1;

[0081] Left and right pixels in the same row: I(1,2) = 98 ≥ 100 => 0; I(3,2) = 99 ≥ 100 => 0;

[0082] Lower pixel: I(1,3) = 97 ≥ 100 => 0; I(2,3) = 96 ≥ 100 => 0; I(3,3) = 95 ≥ 100 => 0.

[0083] (5) Generate binary coding string

[0084] Arrange the comparison results (0 or 1) into a binary string in a predetermined order. Arrangement order (starting from the upper left corner, row first):

[0085]

[0086] Generated binary coding string:

[0087] C(2,2) = [1,1,1,0,0,0,0,0]

[0088] Convert the binary coding into a decimal number:

[0089] D(2,2) = 224

[0090] Step Four: Multi-scale Census feature extraction and weighted fusion.

[0091] Since the features of an image may vary at different scales, for the same image, the present invention uses multi-scale Census transform to extract multi-level features of the image, obtains Census images at different scales, and converts them into decimal images respectively; subsequently, the information at different scales is combined through a weighted fusion method to obtain a more comprehensive feature representation. The specific operations are as follows:

[0092] 1. Multi-scale feature extraction.

[0093] The present invention performs multi-scale Census transform by applying sliding windows of different sizes, such as windows of sizes 3×3, 5×5, 7×7, etc. Windows of different scales can capture local and global features in the image. Smaller windows are suitable for capturing detailed features, while larger windows are helpful for capturing global structural information.

[0094] 2. Weighted fusion.

[0095] Perform weighted fusion on the features extracted at multiple scales, that is, perform weighted fusion on the decimal images of different scales. The present invention assigns different weights to the feature maps of each scale for weighted calculation according to the importance of each scale. The final fused feature map is the weighted average of the feature maps of all scales, which can integrate the information of different scales and avoid the deficiencies caused by relying only on the features of a single scale.

[0096] 3. Normalization processing.

[0097] Perform normalization processing on the fused image to ensure that each pixel value is within the standardized range (for example, [0, 255]). The normalized image can better meet the input requirements of the spiking neural network and avoid adverse effects on the consistency of neuron responses caused by excessive differences in pixel values.

[0098] Step Five: Spiking signal generation and encoding.

[0099] 1. Spiking signal generation.

[0100] Convert each pixel value in the fused feature map from decimal pixels to binary spiking signals to obtain a complete spiking image. Specifically, each decimal pixel of each pixel is represented by 8-bit binary. For example, assuming the use of census images with three different scale radii, at pixel (x, y), the census image values of different scales are:

[0101] D 1 (x, y) = 150; D 2 (x, y) = 184; D 3 (x, y) = 200

[0102] Calculate the fused value: D f (x, y) = 0.5×150 + 0.3×184 + 0.2×200 = 75 + 55.40 = 170.2 (minimum D min = 100, maximum D max = 200)

[0103] The normalized value:

[0104] Convert the fused and normalized decimal pixel 179 into an eight-bit binary spiking signal through spiking conversion:

[0105]

[0106] Finally, each pixel of the image is represented as an 8-dimensional spike vector, which contains the spatio-temporal activity information of the pixel. The spike vectors of all pixels together constitute a complete spike image. Each bit of binary represents the spike emission state of the pixel at a certain moment, reflecting the change of the image in the time dimension.

[0107] 2. Spike image construction.

[0108] The spike signals of all pixels generate a spike image in time sequence to form a complete spike image. The spike image is used as the input of the spiking neural network for further training and pattern recognition.

[0109] Executing the image spike coding method based on multi-scale Census transform constitutes an image spike coding module, which is used to encode a color image into a spike image. A spiking neural network is built using an input layer, a hidden layer, a fully connected layer, and an output layer, and forms a classification system with this image spike coding module. The spiking neural network in the system is trained as follows:

[0110] In the training stage, the spiking neural network gradually learns the spatio-temporal features in the image by simulating the spike emission mechanism of neurons. The present invention uses the "Spike-Timing-Dependent Plasticity (STDP)" learning rule to optimize the spike emission timing of the neural network, enabling the network to correctly identify the spatio-temporal features of handwritten digits through the spike responses of neurons.

[0111] Through the backpropagation algorithm, the present invention optimizes the connection weights and spike emission timing of the network. In each round of training, the network gradually adjusts the spike emission time to learn the unique spatio-temporal patterns of the digits in the image. The ultimate goal of training is to enable the network to accurately identify the spike coding patterns of the digits "0" to "9". Referring Figure 2 As shown, the present invention performs ten-class classification on the MINIST dataset, and the output channels of the fully connected layer are 10. Among them, the first spike neuron layer of the spike neuron layer has 784 neurons for extracting low-level features, and the second spike neuron layer has 784×2 neurons for extracting high-level features; the fully connected layer maps the 784×2 spike signals output by the upper layer into probability values of each classification type of the digits 0–9 through the softmax function.

[0112] During the training process, the cross-entropy loss function is used to measure the error between the model prediction and the true category, and the spike-based backpropagation algorithm is used to adjust the network weights, enabling the system to gradually learn and optimize the classification performance. Finally, through multiple rounds of iterative training, the system can accurately identify the corresponding digit categories in the spike data.

[0113] The handwritten digit recognition using the above system is as follows:

[0114] The trained spiking neural network can receive the input spiking image signal. In the test phase, each input image is passed to the spiking neural network for processing through the generated spiking signal. Each pixel in the image is represented by a spiking sequence, reflecting the spatio-temporal characteristics of the image.

[0115] During the digit recognition process, the spiking neural network gradually extracts the spatio-temporal characteristics of the image based on the input spiking signal, and finally determines the digit represented in the image. By analyzing the spiking signal, the network can recognize the digit in the image and output the corresponding digit category according to the spiking response of the neurons.

[0116] The spiking intensity and timing of each neuron reflect the characteristics of the digit in the image. The trained spiking neural network determines the most matching digit category (from 0 to 9) by analyzing these spiking signals. Although the digit positions in the MNIST dataset are fixed, the spiking neural network of the present invention can locate the central position of the digit according to the spatio-temporal characteristics in the image to ensure accurate classification.

[0117] Figure 3 Shows the comparison diagram of processing handwritten digits 0 - 9 in the MINIST dataset by the image spiking coding method based on multi-scale Census transform proposed by the present invention. It can be seen that the images processed by the present invention have excellent recognition effects.

[0118] Figure 4 Shows the change trends of the training accuracy and test accuracy during 30 rounds of iteration. As the number of training rounds increases, the accuracy of the training set gradually improves and tends to saturation, while the test set accuracy increases rapidly in the early stage of training and gradually stabilizes in the later stage, reflecting the convergence process of the network.

Claims

1. An image pulse coding method based on multi-scale Census transform, characterized in that: The steps include: Step 1: grayscale, denoise, contrast enhance and normalize the color image; Step 2: Perform Census pulse encoding on the image processed in step 1 at multiple scales to obtain Census images of different scales, and convert them into decimal images respectively; Step three, weighted fusion of decimal images of different scales in step two, and then normalization, and then converting decimal pixels into binary pulse signals, and generating a complete pulse image by aligning the pulse signals of all pixels in time sequence.

2. The image pulse coding method based on multi-scale Census transform according to claim 1, characterized in that: In the step 2, pixels at each position (x, y) of the image I are processed by local binary coding, and a sliding window is used for Census transformation, where the window size is w×w, and the central pixel is compared with its neighborhood pixels. If the grayscale value of the neighborhood pixel is greater than or equal to the central pixel value, the corresponding binary code is 1, otherwise it is 0, thereby obtaining an output binary image, in which each pixel value is represented by a binary code, reflecting the relative relationship between local pixels; by selecting different window sizes, Census pulse coding at different scales is achieved.

3. The image pulse coding method based on multi-scale Census transform according to claim 2, characterized in that: The formula of the Census transformation is expressed as: in is the window radius, C(x,y) represents the Census code at position (x,y), which is a w 2 -1-dimensional binary code string, I(x,y) is the pixel value corresponding to the position (x,y) in image I, i,j represent the offset of the coordinates, i,j are not 0 at the same time.

4. The image pulse coding method based on multi-scale Census transform according to claim 2, characterized in that: For pixels close to the edge of the image, zero padding is performed on the boundaries.

5. The image pulse coding method based on multi-scale Census transform according to claim 2, characterized in that: The binary image is converted into a decimal image, that is, the Census code C(x,y) corresponding to each pixel position (x,y) is converted into a Census image value D(x,y), where D(x,y) is a decimal number, and the formula is expressed as follows: Where C(x,y) k is the kth bit of the binary code, and N is the length of the binary code.

6. The image pulse coding method based on multi-scale Census transform according to claim 5, characterized in that: In step 3, the weighted fused decimal image D f (x,y) is represented as: Where M is the scale number, w m is the weight of the mth scale, D m (x,y) is the decimal image obtained by Census pulse coding conversion at the mth scale.

7. The image pulse coding method based on multi-scale Census transform according to claim 6, characterized in that: The step 3 is to normalize D using linear normalization. f (x,y) is mapped to [0,255], expressed as: Where D norm (x,y) is the normalized result, D min and D max are the minimum and maximum pixel values ​​in the fused feature map.

8. The image pulse coding method based on multi-scale Census transform according to claim 6, characterized in that: In the step three, the fused and normalized decimal pixels are converted into eight-bit binary pulse signals. Each pixel of the image is represented as a pulse vector, and the pulse vectors of all pixels together constitute a complete pulse image, which is used as the input of the SNN for subsequent feature extraction and pattern recognition tasks.

9. A classification system based on a pulse neural network, comprising an image pulse coding module and a pulse neural network; The image pulse encoding module executes the image pulse encoding method based on multi-scale Census transform according to any one of claims 1 to 8, encoding the color image into a pulse image; The spiking neural network includes an input layer, a hidden layer, a fully connected layer and an output layer; The input layer inputs the encoded pulse image into the network, and the input shape is [T, N, C, H, W], where: T is the number of time steps, N is the batch size, C is the number of channels, H and W are the image height and width. The pulse sequence of each pixel reflects the distribution of its brightness information in the time dimension, providing spatiotemporal feature input for subsequent layers; The hidden layer includes two layers of pulse neuron layers, which are used to extract spatiotemporal features; The fully connected layer flattens the output of the hidden layer into a one-dimensional vector, and the number of output channels is the number of classification types.

10. The classification system based on pulse neural network according to claim 9, characterized in that: The spiking neuron layer extracts local features of the image, the first spiking neuron layer extracts low-level features, and the second spiking neuron layer extracts high-level features; The spiking neuron layer simulates the dynamic behavior of biological neurons; The fully connected layer maps the pulse signal output by the upper layer into the probability value of each classification type through the softmax function.