Image defogging algorithm based on multi-scale pulse neural network
By introducing multi-scale pulse neural networks into the image defog removal algorithm, the pulse distribution mechanism of biological neurons is simulated, and the problem of poor results in existing algorithms in complex scenarios is solved, and efficient and low-resource image defog treatment is achieved.
Patent Information
- Application Number
- CN202510216218.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-26
AI Technical Summary
The existing image defog removal algorithm is not effective in complex scenarios and has high computational complexity, making it difficult to effectively process in low-resource environments.
The image defog algorithm based on multi-scale pulsed neural network is adopted, and the multi-scale LIF module and fully connected neural network layer are combined to simulate the pulse distribution mechanism of biological neurons, and feature extraction and defog treatment are performed.
It realizes efficient defog processing of images in low computing resource environments, improves image clarity and detail recovery effect, and is suitable for a variety of image processing scenarios.
Smart Images

Figure CN120147184A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and more particularly, to an image dehazing algorithm based on a multi-scale spiking neural network. Background Art
[0002] Due to the influence of factors such as fog, haze, and air pollution, image details and contrast may be weakened, and colors may also be distorted. Therefore, image dehazing and restoration are extremely important.
[0003] Methods based on image enhancement, methods based on physical models, and methods based on deep learning are traditional image processing methods. The method based on the physical model starts from the atmospheric scattering model, establishes a mathematical model of the image degradation process, and then inversely restores the clear image. The most famous of these methods is the Dark Channel Prior (DCP) algorithm. The DCP algorithm is based on an important observation: in natural scenes, there is at least one pixel with a very low pixel value in one or more color channels of the vast majority of non-sky regions. Using this prior knowledge, DCP can effectively estimate hazy images and thus restore clear images. Although DCP performs well in many cases, it has a high complexity and poor effects in some scenarios.
[0004] Therefore, it is necessary to address the defect that the comprehensive use effect of traditional image processing methods is poor, and propose an image dehazing algorithm based on a multi-scale spiking neural network. Summary of the Invention
[0005] The content of this application is used to briefly introduce concepts, which will be described in detail in the following detailed implementation section. The content of this application is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0006] To solve the technical problems mentioned in the above background art section, some embodiments of this application provide an image dehazing algorithm based on a multi-scale spiking neural network, including:
[0007] Preprocess the input foggy image; the preprocessing includes image normalization and data augmentation;
[0008] Execute an embedding program for extracting preliminary features on the preprocessed image;
[0009] Construct a basic layer composed of multi-scale LIFs; the basic layer includes multi-scale convolution, LIF spike generation, and fully connected; the fully connected is composed of two-dimensional convolutions connected by LeakyReLU;
[0010] Input the image into the spiking structure of the LIF to form a spiking mechanism that simulates biological neurons;
[0011] Perform a drop path regularization operation on the output image of the LIF Module;
[0012] Input the image into the fully connected layer for regularization processing;
[0013] Use SK-fusion to perform feature fusion on the image;
[0014] Return the embedding program that extracts preliminary features from the preprocessed image until the number of loops is the numerical value N; N is a positive integer;
[0015] Perform de-embedding on the output image;
[0016] Restore the embedded feature map to achieve image dehazing.
[0017] Furthermore, scale the data proportionally so that its pixel values fall within the range of [-1, 1].
[0018] Furthermore, extract image features based on the convolutional layer;
[0019] Use the image features for channel mixing;
[0020] Simulate the spiking mechanism of biological neurons based on the LIF neuron layer;
[0021] Perform token mixing;
[0022] Based on the activation function layer, convert the output of the LIF neuron layer into a pulse signal.
[0023] Furthermore, based on the linear transformation layer, fuse the input feature map and the features of the original image;
[0024] Based on the activation function layer, form the output port of the neural network;
[0025] Incorporate the linear transformation layer and the activation function layer into the fully connected neural network layer;
[0026] Connect each fully connected neural network layer to form a full connection.
[0027] Furthermore, based on the LIF, form an MLP neural network layer;
[0028] Use the MLP neural network layer to convert the input features into an embedded vector;
[0029] Based on the embedded vector, form the input parameters of the pulse neuron module in the spiking structure.
[0030] Furthermore, add a dwconv to the first MLP layer;
[0031] Replace the AxialShift of the MLP layer based on LIF;
[0032] Based on the horizontal LIF of the MLP layer, accumulate and propagate information in the horizontal direction of the image features;
[0033] Based on the vertical LIF of the MLP layer, accumulate and propagate information in the vertical direction of the image features.
[0034] Furthermore, utilize the information accumulated and propagated in the horizontal direction of the image features and the information accumulated and propagated in the vertical direction of the image features to simulate the spike generation mechanism of biological neurons for feature mixing.
[0035] Furthermore, when u < V th the membrane potential u decays over time and receives an external input I:
[0036]
[0037] When the membrane potential u reaches the threshold V th the neuron fires a spike and the membrane potential is reset to u reset :
[0038] o = 1, u = u reset , u ≥ V th
[0039] where u is the membrane potential, I is the input from the upper layer, τ is the time coefficient, o is the output, and V th is the firing threshold of this neuron. When a spike is triggered, the membrane potential u is reset to u reset , where t in the formula represents the index of the group.
[0040] Furthermore,
[0041]
[0042] o = u u = u reset , u ≥ V th
[0043] We replace the output 1 with u to change the original binary output to a linear output, thus retaining the full-precision information. Applying the full-precision LIF model to iterative LIF, we get:
[0044]
[0045] W T the transpose of the weight matrix, x is the input feature, r t+1 nis the final full-precision output at step t+1, o t+1 n is a temporary variable recording the output state at step t+1, and the coefficients τ and V th are neural network learning parameters.
[0046] Furthermore,
[0047]
[0048] In summary: The method of the present invention can be widely applied to various scenarios in reality that require image dehazing and restoration, including but not limited to: improving the clarity and recognizability of images by dehazing the haze images collected by traffic cameras, thereby enhancing traffic management and accident prevention capabilities. Applied to the processing of satellite images and drone images to remove the blurring effects brought by the atmospheric environment and improve the accuracy of ground object recognition. In the security monitoring system, dehazing the monitoring videos under bad weather conditions to improve the monitoring effect and ensure public safety. Used for the processing of medical images to remove noise and blurring, thereby improving the clarity and diagnostic accuracy of the images. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The drawings forming a part of this application are used to provide a further understanding of this application, making other features, objectives, and advantages of this application more obvious. The schematic embodiments of the drawings of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application.
[0050] In addition, throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the elements and elements are not necessarily drawn to scale.
[0051] In the drawings:
[0052] Figure 1 is the overall flowchart of the method of the present invention.
[0053] Figure 2 is the schematic diagram of the overall structure of the hybrid dehazing model.
[0054] Figure 3 is the schematic diagram of the SK-Fusion module in the fog model. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] Embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present application. It should be understood that the drawings and embodiments of the present application are only for exemplary purposes and are not used to limit the protection scope of the present application.
[0056] In addition, it should be noted that for ease of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0057] The present application will be described in detail below with reference to the drawings and in combination with embodiments.
[0058] S1. Preprocess the input fog image, including image normalization (to make the pixel values of the image within a unified range) and data augmentation (including image cropping, flipping, etc.);
[0059] S2. Perform an embedding operation on the preprocessed image (implemented through two-dimensional convolution) to extract preliminary features;
[0060] S3. Construct a basic layer composed of multi-scale LIF (Leaky Integrate and Fire), including multi-scale convolution, LIF pulse emission, and fully connected (Multi-Layer Perceptron, MLP);
[0061] S4. The feature map is normalized through the multi-scale LIF module, and convolution operations are performed respectively using three different scales of convolution kernels (3x3, 5x5, 7x7) to extract features of different scales, and then processed respectively;
[0062] S5. Input the processed feature map into the LIF Module (LIF pulse emission structure) to simulate the pulse emission mechanism of biological neurons;
[0063] S6. Perform a drop path regularization operation on the output of the LIF Module to introduce randomness to prevent overfitting;
[0064] S7. Input two fully connected layers composed of two-dimensional convolutions connected by LeakyReLU and perform regularization processing;
[0065] S8. Use the SK-fusion (Selective Kernel fusion) shown as Figure 3 to perform feature fusion and integrate multi-scale feature information;
[0066] S9. Repeat the above basic layer operations five times to extract higher-level features layer by layer;
[0067] S10. Perform an unembedding operation (implemented by two-dimensional convolution) on the final output of the basic layer to restore the embedded feature map and achieve defogging.
[0068] The present invention adopts a novel hybrid model architecture that combines a spiking neural network and a fully connected neural network. Specifically, the model uses a spiking neuron (Leaky Integrate-and-Fire, LIF) module for simulating biological neurons after being processed by three different scales of convolutional kernels as the feature extraction layer, enabling the neural network to transmit and process information through discrete spike signals, thereby improving the computational efficiency and energy efficiency ratio. The output features are further processed by a fully connected neural network respectively to enhance the accuracy of image detail restoration. Each LIF module includes a convolutional layer, a LIF neuron layer, and an activation function layer; the convolutional layer is used to extract image features and perform channel mixing; the LIF neuron layer is used to simulate the spike generation mechanism of biological neurons and perform token mixing; the activation function layer is used to convert the output of the LIF neuron layer into a spike signal. The LIF Module (LIF spike generation structure) includes: an MLP layer that converts the input features into an embedded vector and provides a suitable input for the subsequent spiking neuron module. In this part, a dwconv (Depthwise Convolution) is added after the first MLP layer, and the shift operation adopted in AxialShift (axial shift, a convolutional mechanism based on axial displacement to more effectively capture local and global information in the image) is replaced with a LIF neuron. Since the iterative LIF neuron is essentially an activation function, it is necessary to add a dwconv layer in front of the iterative LIF neuron. Next, we use horizontal LIF (HLIF) and vertical LIF (VLIF) to accumulate and propagate information in the horizontal and vertical directions of the image features respectively. These two modules perform feature mixing by simulating the spike generation mechanism of biological neurons, retaining local and global information in the feature map. The full-precision LIF neurons we introduced are different from traditional LIF neurons. The behavior of the classical LIF model can be modeled as follows:
[0069] (1) When u < V th the membrane potential u decays over time and receives an external input I:
[0070]
[0071] (2) When the membrane potential u reaches the threshold V thWhen a neuron fires a pulse, the membrane potential is reset to u reset :
[0072] o = 1, u = u reset , u ≥ V th
[0073] where u is the membrane potential, I is the input from the upper layer, τ is the time coefficient, o is the output, and V th is the firing threshold of this neuron. When a pulse is triggered, the membrane potential u is reset to u reset . Different from the accumulation of traditional LIF neurons in the time domain, the neurons used here accumulate and fire in the spatial domain. In this design, the LIF module first divides the picture into several groups. The t in the formula represents the index of the group, rather than the time step. Since the input features are in full precision, we prefer to obtain a full-precision output to preserve the information in the group. To meet our requirements, we use the following full-precision LIF function:
[0074]
[0075] o = u u = u reset , u ≥ V th
[0076] We replace the output 1 with u to change the original binary output to a linear output, thus preserving the full-precision information. Applying the full-precision LIF model to iterative LIF, we get:
[0077]
[0078] Here, W T is the transpose of the weight matrix, x is the input feature, r t+1 n is the final full-precision output at step t + 1, and o t+1 n is just a temporary variable recording the output state at step t + 1. The coefficients τ and V th are learnable. Using this explicit iterative LIF neuron, the backpropagation process can be completed with the chain rule:
[0079]
[0080]
[0081] The output passes through the activation function layer to further process the output signal to complete the generation of the pulse signal. Finally, SK-fusion is used for fusion. Overall, 5 consecutive basic layers are cascaded to extract higher-level features layer by layer to output more excellent results.
[0082] The image dehazing and restoration method based on multi-scale spiking neural network proposed by the present invention has high efficiency. The introduction of spiking neural network greatly reduces the computational complexity and energy consumption of the model and improves the processing efficiency. High performance. The model has multi-scale feature extraction and restoration capabilities, and can achieve high-quality image dehazing and restoration in various haze environments. Wide applicability. The method can be applied to a variety of image processing scenarios and has broad application prospects. Low resource consumption. Compared with traditional deep learning models, this method reduces the consumption of computing resources while ensuring high performance.
[0083] The dehazing dataset (Realistic Single Image Dehazing, RESIDE) is a commonly used dataset for image dehazing algorithm research. It provides real-world images in various different environments and scenarios to help us evaluate and compare the performance of dehazing algorithms. The RESIDE dataset contains multiple sub-datasets, and each sub-dataset has its own different characteristics and uses. In our experiments, three of its sub-datasets were used, namely RESIDE-IN (Indoor Training Set), RESIDE-OUT (Outdoor Training Set), and RESIDE-6K (Real-world 6K Dehazing Dataset). RESIDE-IN is the dehazing training set in the RESIDE dataset specifically for indoor scenes, mainly containing synthetic fog images in indoor scenes. When generating this dataset, indoor images were used and haze was simulated through a physical model. Indoor scenes usually have more stable light, which helps the algorithm optimize its performance in specific environments. RESIDE-OUT is the dehazing training set in the RESIDE dataset for outdoor scenes, mainly containing synthetic fog images of outdoor scenes, simulating the haze effect in the natural environment. Due to the greater light changes and complexity in outdoor scenes, its scene diversity is higher, and it can better simulate complex situations in real life. RESIDE-6K is a dataset in the RESIDE dataset containing 6,000 real-world haze images, mainly containing real haze images from real life. These images were obtained under natural conditions without using any synthetic fog generation technology.
[0084] In fog and haze weather, a large number of tiny suspended particles in the air will refract and scatter light. Then, the light after mixing with the light reflected by the target to be observed results in a foggy image. In the foggy image, the image clarity and contrast of the observed target are reduced, and even phenomena such as image color deviation and a large amount of detail loss occur, thereby making it impossible to obtain the true image information of the target to be observed.
[0085] In the existing related technologies, methods such as enhancing contrast and color restoration are used to convert foggy images into defogged images. However, such implementation methods require the foggy images to be processed to be obtained under sufficient light sources. In complex light source scenarios, such as in night environments, the defogging methods used in related technologies are prone to causing the loss of a lot of valuable information, ultimately resulting in poor quality of the obtained target defogged images.
[0086] Based on the current situation, the present application provides a brand-new image defogging method aimed at improving the quality of defogged images.
[0087] Next, the image defogging method in the present application will be described through the following specific implementation.
[0088] Data preprocessing: First, the algorithm normalizes the input foggy image so that its pixel value range is between [-1, 1]. Subsequently, in some cases, data augmentation techniques such as random cropping, flipping, and rotation are used to generate more training samples.
[0089] Model construction: The hybrid model is mainly constructed based on multi-scale spiking neural network layers. The spiking neural network part uses horizontal and vertical LIF neuron models to simulate the spiking activities of biological neural networks, and at the same time uses multi-scale convolution for feature processing and then fusion. This part of the network can effectively capture and process sparse data, greatly improving the processing efficiency. The model uses L1 as the loss function and optimizes the model parameters through the backpropagation algorithm. To improve the training efficiency and model performance, the model uses the AdamW optimizer, and gradually reduces the learning rate during training through the cosine annealing method, reducing the chance of the model falling into local minima and making large learning rate updates, helping the model to converge better, thereby improving the prediction accuracy. In addition, by evaluating and tuning the model multiple times, it is ensured that it performs excellently on both the training set and the validation set.
[0090] Model training: The model training adopts the supervised learning method, and uses the labeled clear images and haze images to train the model. The loss function adopts the L1 loss method to take into account the global and local detail restoration effects of the image. And different numbers of trainings are performed according to different data sets.
[0091] Model Optimization: To improve the training efficiency and performance of the model, the present invention introduces a mixed-precision training technique, which realizes the dynamic switching between half-precision and full-precision through Automatic Mixed Precision (AMP), reducing the video memory occupancy and computational overhead. At the same time, the AdamW optimizer and the AdamW optimizer with weight decay are used to optimize the main model parameters and the spiking neuron parameters respectively, so as to further enhance the generalization ability of the model. And the method of cosine decay is used to gradually reduce the learning rate during the training process, reducing the chance of the model falling into local minima and making large learning rate updates, helping the model to converge better, thereby improving the prediction accuracy.
[0092] Model Verification and Testing: After the training is completed, the model is evaluated through a preset validation set to select the best model parameters. Finally, it is tested in the test set and compared with existing defogging methods to verify the superiority of the method of the present invention. The experimental results show that the method of the present invention significantly improves the clarity and detail restoration effect of the defogged image while maintaining low power consumption and high efficiency.
[0093] Model Deployment: Finally, the method of the present invention can be deployed into practical applications, such as scenarios that require real-time defogging like video surveillance and driverless driving. Through model lightweight and acceleration technologies, it ensures efficient image defogging and restoration on edge devices.
[0094] Performance Evaluation and Optimization: To objectively evaluate the experimental results, we use Peak Signal-to-Noise Ratio (PSNR) and structural similarity index (SSIM) to evaluate the defogging effect, and optimize the parameters of the defogging network based on the evaluation results. PSNR is used to evaluate the distortion degree of an image or video. It is calculated by comparing the error between the original signal (usually the original foggy image or video) and the reconstructed signal (usually the image or video after defogging), and the formula is as follows:
[0095]
[0096] Among them, MAX I is the maximum value of the image pixels, and MSE is the mean square error. And SSIM is an index to measure the similarity between two images. Among the two images used, one is an undistorted image without compression, and the other is a distorted image. The calculation formula is as follows:
[0097]
[0098] Among them, SSIM(x,y): the structural similarity index, indicating the structural similarity degree between image x and image y.
[0099] μx : The average brightness of image x. μ y : The average brightness of image y. The variance of image x, representing the degree of dispersion of the brightness distribution of image x. The variance of image y, representing the degree of dispersion of the brightness distribution of image y. σ xy : The covariance of image x and image y, representing the correlation of the brightness distributions of the two images. c 1 and c 2 : Two constants introduced to avoid a zero denominator, usually taking the value of c 1 =(K 1 L)2, c 2 =(K 2 L)2, where L is the dynamic range of the pixel values (e.g., for an 8-bit image, L = 255), and K 1 and K 2 are constants less than 1.
[0100] The above description is only some preferred embodiments of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the embodiments of the present application.
Claims
1. An image dehazing algorithm based on a multi-scale spiking neural network, comprising: Preprocess the input fog image; The preprocessing includes image normalization and data enhancement; Performing an embedding procedure to extract preliminary features on the preprocessed image; Construct a base layer composed of multi-scale LIF; the base layer includes multi-scale convolution, LIF pulse emission, and full connection; the full connection is composed of two-dimensional convolution connected by LeakyReLU; The image is input into the spiking structure of LIF to form a spiking mechanism that simulates biological neurons; Perform drop path regularization on the output image of LIF Module; Input the image to the fully connected layer for regularization; Use SK-fusion to fuse image features; Returning to the embedding procedure for extracting preliminary features from the preprocessed image until the number of cycles reaches a value N, where N is a positive integer; De-embed the output image; The embedded feature map is restored to achieve image dehazing.
2. The image dehazing algorithm based on multi-scale spiking neural network according to claim 1 is characterized in that: The preprocessing of the input fog image includes: Scale the data so that its pixel values fall within the range [-1, 1].
3. The image dehazing algorithm based on multi-scale spiking neural network according to claim 2 is characterized in that: The method of constructing a base layer composed of multi-scale LIFs includes: Extract image features based on convolutional layers; Use image features to perform channel mixing; Simulate the pulse firing mechanism of biological neurons based on LIF neuron layer; Perform token mixing work; Based on the activation function layer, the output of the LIF neuron layer is converted into a spike signal.
4. The image dehazing algorithm based on multi-scale spiking neural network according to claim 3 is characterized in that: The constructing of a base layer composed of multi-scale LIFs also includes: Based on the linear transformation layer, the input feature map and the features of the original image are fused; Based on the activation function layer, the output port of the neural network is formed; Incorporate linear transformation layers and activation function layers into fully connected neural network layers; Connect each fully connected neural network layer to form a fully connected network.
5. The image dehazing algorithm based on multi-scale spiking neural network according to claim 4 is characterized in that: The step of inputting the image into the pulse emission structure of the LIF to form a pulse emission mechanism simulating biological neurons includes: Based on LIF, an MLP neural network layer is formed; Use the MLP neural network layer to convert the input features into embedding vectors; Based on the embedding vector, the input parameters of the spiking neuron module in the spiking structure are formed.
6. The image dehazing algorithm based on multi-scale spiking neural network according to claim 5 is characterized in that: The LIF-based MLP neural network layer is formed, including: Add a dwconv in the first MLP layer; AxialShift based on replacing the MLP layer with LIF; Based on the horizontal LIF of the MLP layer, the image features are accumulated and propagated in the horizontal direction; Based on the vertical LIF of the MLP layer, the information of the image features is accumulated and propagated in the vertical direction.
7. The image dehazing algorithm based on multi-scale spiking neural network according to claim 6 is characterized in that: The step of inputting the image into the pulse emission structure of the LIF to form a pulse emission mechanism simulating biological neurons also includes: By accumulating and propagating information in the horizontal direction and in the vertical direction of image features, the pulse emission mechanism of biological neurons is simulated to perform feature mixing.
8. The image dehazing algorithm based on multi-scale spiking neural network according to claim 5 is characterized in that: The image dehazing algorithm based on the multi-scale spiking neural network also includes the behavior modeling of the LIF model: When u <V th When , the membrane potential u decays over time and receives external input I: When the membrane potential u reaches the threshold V th When the neuron fires a pulse, the membrane potential is reset to u reset : o=1,u=u reset ,u≥V th where u is the membrane potential, I is the input from the upper layer, τ is the time coefficient, o is the output, V th is the firing threshold of this neuron. When a spike is triggered, the membrane potential u is reset to u reset , where t represents the index of the group.
9. The image dehazing algorithm based on multi-scale spiking neural network according to claim 8 is characterized in that: The image dehazing algorithm based on the multi-scale spiking neural network also includes full-precision LIF modeling: o=0, u <V th o=uu=u reset ,u≥V th We replace the output 1 with u, so that the original binary output becomes a linear output, thus retaining the full-precision information. Applying the full-precision LIF model to iterative LIF, we get: W T The transpose of the weight matrix, x is the input feature, r t+1 n is the final full-precision output of step t+1, o t+1 n is a temporary variable that records the output state of step t+1, the coefficients τ and V th Learn parameters for neural networks.
10. The image dehazing algorithm based on multi-scale spiking neural network according to claim 9 is characterized in that: The image dehazing algorithm based on the multi-scale spiking neural network also includes the back propagation modeling of the chain rule of explicit iterative LIF neurons:
Citation Information
Patent Citations
SAR image ship target identification method based on pulse neural network
CN113111758A
Steel surface defect image classification method based on pulse convolutional neural network
CN116188870A
Image classification learning method based on multi-scale pulse convolutional neural network
CN118447317A
Time-space domain feature dynamic target identification method based on spiking neural network
CN118823484A
Device and method for recognizing image using brain-inspired spiking neural network and computer readable program for the same
US20230117659A1
Cited By
Image restoration method and system and storage medium
CN122492500A
An image restoration method, system and storage medium
CN122492500B