Algorithm for enhancing night concrete image based on state space coding

Through a deep learning algorithm based on state space encoding, combined with state space module and convolution module, a night concrete image enhancement model is built, which solves the problems of noise and color distortion in night images, and realizes high-quality image enhancement and real-time detection.

CN120198310AActive Publication Date: 2025-06-24SOUTHWEAT UNIV OF SCI & TECH
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202510689359.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-06-24
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with the severe noise and color distortion problems in night concrete images, and the traditional methods have problems of local over-enhancement, under-enhancement and poor robustness.

Method used

A deep learning algorithm based on state space encoding is adopted, combined with state space module and convolution module, a night concrete image enhancement model is built, and images are processed through multi-level encoder and fusion decoder, and the loss function is optimized to reduce artifacts and overexposure.

Benefits of technology

Effectively capture global and local information of the image, reduce artifacts, improve image quality, reduce computing power requirements, and is suitable for edge devices to achieve real-time image enhancement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198310A_ABST
    Figure CN120198310A_ABST
Patent Text Reader

Abstract

The invention discloses an algorithm for enhancing a night concrete image based on state space coding, and relates to the field of digital image processing, and the algorithm comprises the steps: S1, constructing a night concrete image enhancement algorithm based on a state space module and a convolution module based on a state space convolution deep network technology; s2, constructing a training model based on the night concrete image enhancement algorithm in the S1; and S3, a test sample is input into the training model constructed in the S2 to obtain an enhanced concrete image, and the training model evaluates an enhancement effect based on a loss function. According to the night concrete image enhancement method based on state space coding provided by the invention, the proposed night concrete image enhancement algorithm is respectively used for feature extraction and multi-level feature fusion of images under different levels, and compared with a low-illumination enhancement algorithm based on Transform, the performance is improved, the parameter quantity of a model is reduced, and the accuracy of the image enhancement algorithm is improved. And the reasoning speed is accelerated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of digital image processing. More specifically, the present invention relates to an algorithm for enhancing night-time concrete images based on state space coding. Background Art

[0002] With the popularization of artificial intelligence algorithms, deep learning methods have received great attention in both the industrial and academic fields. During the process of concrete quality inspection, the slump of concrete can be judged by taking pictures of the flow of concrete with a camera. However, during the night-time concrete production process, due to reasons such as insufficient exposure at night, the concrete images taken at night are prone to problems such as the overall or partial areas being too dark, it is difficult to capture detailed information, color distortion, and severe noise. In an environment with insufficient light, low-quality images with high noise, low contrast, and poor visibility will be obtained due to the excessive exposure time or insufficient sensitivity of the photographic machine, which will in turn affect the subsequent detection of the concrete slump. Therefore, it is necessary to improve the image quality through low-light image enhancement technology to provide a reliable basis for the subsequent concrete slump detection.

[0003] Currently, low-light image enhancement algorithms can generally be divided into image enhancement algorithms based on traditional technologies and image enhancement algorithms based on deep learning, among which: Traditional methods include gray-scale transformation methods, histogram equalization methods, and Retinex methods. Traditional image enhancement methods. Gray-scale transformation belongs to the spatial domain processing method, and its essence is to modify the gray scale of each pixel point in the image according to certain rules to change the gray scale range of the image. Typical gray-scale transformations can be divided into linear transformations and non-linear transformations. Histogram equalization is to use the cumulative distribution function to adjust the output gray scale so that the gray scale distribution of the image becomes more uniform. The Retinex theory is based on the view that the color of an object depends on its ability to reflect long-wave (red), medium-wave (green), and short-wave (blue) illumination, rather than the absolute value of the reflected light. By estimating the incident component from the original image and processing it, image enhancement is achieved. The most classic algorithms currently are the single-level Retinex algorithm, the multi-level Retinex algorithm, and the MSR algorithm with color restoration. However, these methods have the problem of excessive enhancement resulting in too large color deviation.

[0004] Deep learning-based methods include Jain et al. using a variant of stacked sparse denoising autoencoders to brighten and denoise low-light images. Lu et al. proposed an end-to-end multi-branch enhancement network that extracts effective features through a feature extraction module, an enhancement module, and a fusion module to improve the performance of low-light enhancement models. To reduce the computational burden of deep learning models based on the Retinex theory, Li et al. proposed a lightweight LightenNet network for low-light image enhancement. Some teams proposed using adversarial networks (GANs) to enhance low-light images. Through the adversarial process, the generator can learn to generate more natural and closer-to-real illumination condition images. Jiang et al. proposed EnlightenGAN, which can provide effective enhancement without relying on paired training data. Yu et al. segmented the input image into sub-images according to the exposure by strengthening adversarial learning of exposure photos. For each sub-image, the policy network sequentially learns local exposure based on reinforcement learning. Yang et al. proposed a semi-supervised deep recursive band network (DRBN). DRBN first restores the linear band representation of the enhanced image under supervised learning, and then reorganizes the given band through a learnable linear transformation based on unsupervised adversarial learning to obtain an improved band representation. However, these methods cannot effectively handle problems such as severe noise and color distortion existing in the enhanced images.

[0005] Therefore, although some methods in the prior art can achieve good performance, they still have some limitations: for unsupervised and semi-supervised learning methods, there are still certain challenges in achieving stable training, avoiding color deviation, and establishing cross-domain information relationships. For reinforcement learning, it is rather difficult to design an effective reward mechanism and implement it effectively and stably. Most of the existing night single-frame image enhancement algorithms only have a certain effect on some low-light images without noise, while the enhancement results of some particularly dark areas often have severe noise and color distortion, making it difficult to meet the actual needs.

[0006] Furthermore, for example, a low-illumination image enhancement method based on feature fusion and attention embedding with the application number 202310828756.4 introduced attention and feature fusion, which well solved the problem that convolutional networks are difficult to capture global information of images. However, the introduction of the attention mechanism led to a huge increase in the number of model parameters, making it easy for the model to be difficult to deploy on some low-computing-power devices.

[0007] Another example is a low-light image enhancement method based on a convolutional neural network with the application number 201810246429.7, which uses convolution to enhance the image. However, convolution captures poor global information between image contents and is prone to artifacts and overexposure during image enhancement.

[0008] For another example, in a multi-exposure image fusion method based on a multi-layer autoencoder with the application number 202211424921.1, it only enhances low-light images in grayscale and does not solve the problem of enhancing low-light images.

[0009] As can be seen from the above, various existing methods are prone to problems such as local over-enhancement, under-enhancement, and color distortion, resulting in poor robustness of traditional methods, as well as problems in existing methods based on Transformer and convolutional neural networks, such as large model parameter quantities, high computing power requirements, and convolutional blocks can only capture local image information. In summary, there is currently no effective enhancement method for nighttime concrete images. Summary of the Invention

[0010] An object of the present invention is to solve at least the above problems and / or deficiencies and provide at least the advantages described later.

[0011] To achieve these objects and other advantages of the present invention, an algorithm for enhancing nighttime concrete images based on state space coding is provided, including: S1. Using state space convolutional depth network technology, construct a nighttime concrete image enhancement algorithm combining a state space module and a convolutional module; S2. According to the nighttime concrete image enhancement algorithm in S1, construct and train a neural network model; S3. Input the test samples into the trained neural network model to obtain the enhanced images, and evaluate the enhancement effect through the overall loss function in the model. If the enhancement effect meets the expected requirements, it is considered that the neural network model training is completed; if not, return to the S2 stage for further adjustment; S4. Deploy the trained neural network model to edge devices to enhance the nighttime concrete falling images through the neural network model and feedback them to the host computer for real-time display; Among them, in S2, the nighttime single-frame image enhancement algorithm is based on the Pytorch platform of the deep learning framework for model training, and the overall loss function used during model training is L total Expressed by the following formula: In the above formula, , , respectively represent L 2, L SSIM , L color weight coefficients; L 2 is the Euclidean distance loss,L SSIM is the structural similarity loss, L color is the color loss, and L 2. L SSIM , L color are respectively represented by the following formulas: In the above formula, x i , y i respectively represent the RGB values of the pixel points output by the model and the RGB values of the pixel points of the normally exposed image, X , Y respectively represent the enhanced image output by the model and the normally exposed image, n is the number of image pixels, SSIM is the structural similarity formula, G high is the color distribution of daytime pictures that are trained to capture the image color distribution and have normal exposure, G pred is the color distribution of pictures that are trained to capture the image color distribution, have normal exposure, and are enhanced for night images, H is the height of a single-frame image, W is the width of a single-frame image.

[0012] Preferably, in S1, it further includes: using a camera to capture the concrete falling process during the day and night, forming a video dataset, and performing cropping processing on the images in the video dataset to generate an image dataset for training: In S2, using the daytime concrete image dataset in the training dataset as the ground truth, using the nighttime concrete image dataset as the data to be enhanced, after initializing the neural network parameters, through the overall loss function L total perform iterative training on the neural network.

[0013] Preferably, the nighttime concrete image enhancement algorithm includes a multi-level downsampling module, a multi-level encoder, and a multi-level fusion decoder, and its processing flow includes: S10. Obtain images at three different levels of detail, local, and global from the nighttime single-frame image through the multi-level downsampling module; S11. For the images at the three different levels of detail, local, and global obtained in S10, respectively process them through the corresponding detail encoder module, local encoder module, and global encoder module in the multi-level encoder to obtain a preprocessed concrete image; S12. For the preprocessed images obtained by each hierarchical encoder module in S11, a multi-level fusion decoder is used for splicing processing, and upsampling is performed after the splicing processing, thereby obtaining an enhanced image.

[0014] Preferably, in S11, a normalization convolutional layer and a state space image processing layer are defined as a standard state space module, the convolutional kernel size of the normalization convolutional layer is 1×1, and the convolutional stride is 1; Among them, the detail encoder module includes a standard state space module; The local encoder module includes three standard state space modules and a residual mapping; The global encoder module includes a standard state space module and a state space module, and the state space module includes two state space image processing layers.

[0015] Preferably, in S12, the splicing processing flow of the multi-level fusion decoder is as follows: S120. In the global encoder module, the output passing through a standard state space module is defined as feature image I, the output passing through the global encoder module is defined as feature image II, and the output not passing through the global encoder module is defined as feature image III. Feature images I, II, and III are spliced according to the number of channels, and after the splicing result is quadruple upsampled, the number of channels is reduced through a convolutional layer to obtain the global feature image extraction result F high ; S121. In the local encoder module, the output passing through a standard state space module is defined as feature image V, the output passing through two standard state space modules is defined as feature image VI, the output passing through three standard state space modules is defined as feature image VII, and the output not passing through the local encoder module is defined as feature image VIII. Feature images V, VI, VII, and VIII are spliced according to the number of channels, and the splicing result is passed through two convolutional layers to expand the convolutional channels to obtain the feature image extraction result F high ; F middle ; S122. After splicing the result obtained in S121 F middle with the feature image X output by the detail encoder module, the enhanced image is obtained through double upsampling.

[0016] Preferably, the processing flow of the multi-level fusion decoder includes: S130. Use Pixel-Shuffle to upsample the output of the global encoder module by 4 times and connect it to the output of the local encoder module in depth, so as to fuse the feature information extracted from the global encoder module and the local encoder module through convolution operations; S131. Upsample the convolution result in S130 by four times, connect it to the output of the detail fusion module in depth, and then generate the restored RGB image through a convolution layer.

[0017] Preferably, in S2, the calculation formula of SSIM is: Where, , , ; In the above formula, μ X is X the average value of, μ Y is Y the average value of, is X and Y the standard deviation of, C 1, C 2 is a constant, is X the standard deviation of, is Y the standard deviation of, is Y the standard deviation of, is the weighting coefficient, N is the number of image pixels, is X the standard deviation.

[0018] The present invention at least includes the following beneficial effects: The present invention adopts a method for enhancing night concrete images based on the state space equation, which can well capture the global information and local information of the images, ensure the quality of image enhancement, further improve the quality of the enhanced images, and has low computing power requirements. It can be applied to a color enhancement system based on state space coding, enabling the system to be used for night concrete image enhancement based on deep learning, greatly improving the enhancement quality of night concrete images, and bringing an innovative solution to the field of night image enhancement. Specifically, it has the following effects: (1) Reduce artifacts after night image enhancement: Introduce deep learning and state space coding to reduce artifacts generated by the model after enhancement; (2) Adopt an optimized loss function: By designing a color loss function, reduce the overexposure situation during enhancement.

[0019] (3)Multi-scenario support: Compared with traditional image enhancement methods, this method can be trained and improved for different scenarios, enhancing the generalization ability of the model and enabling optimized training for different scenarios.

[0020] (4)Continuous optimization and improvement: By collecting concrete image datasets during the day and night in real time, in-depth data analysis can be conducted, and continuous training on night images can be carried out to improve the night enhancement effect.

[0021] (5)Using state space model encoding with low computational complexity: Compared with the night enhancement model based on the attention mechanism of Transformer, the model has fewer parameters and lower computational complexity, and the model inference time is short, capable of achieving real-time effects.

[0022] (6)Real-time detection: This device can enhance night concrete images in real time, avoiding the timeliness issues in traditional methods.

[0023] Other advantages, objectives, and features of the present invention will be partially reflected by the following description and partially understood by those skilled in the art through the research and practice of the present invention. Description of the Drawings

[0024] Figure 1 Schematic diagram of the processing flow after building the night concrete image enhancement system based on state space encoding in the present invention; Figure 2 Schematic diagram of the processing flow of the multi-level night single-frame image enhancement algorithm of the present invention; Figure 3 Schematic diagram of the basic architecture of the state space model in the present invention; Figure 4 Schematic diagram of the final architecture of the state space model in the present invention. Detailed Embodiment

[0025] The following further elaborates on the present invention in conjunction with the drawings, enabling those skilled in the art to implement it with reference to the description in the specification.

[0026] A real-time enhancement method for night surveillance videos based on state space encoding, the technical solution mainly includes: 1. The camera captures concrete images: The camera captures the day and night videos of concrete falling as video datasets.

[0027] 2. Dataset processing: Crop the collected dataset to an appropriate image size.

[0028] 3. Train the neural network: Use the daytime video dataset as the ground truth, and the nighttime video as the data to be enhanced. Initialize the neural network parameters, and then iteratively train the neural network through the designed loss function.

[0029] 4. Deploy the detection device: Deploy the trained neural network model into the device to achieve real-time detection function.

[0030] 5. Real-time detection and feedback: The camera device transmits video data to the neural network model. The model performs real-time enhancement on the nighttime video stream data and feeds it back to the host computer for real-time display.

[0031] 6. Maintenance and calibration: Regularly maintain the camera device and the neural network model to ensure normal operation, and calibrate the device to adapt to different nighttime situations.

[0032] Specifically, a method for enhancing nighttime concrete images based on state space encoding enhances nighttime concrete images by introducing a method that fuses the state space model and the convolutional network. It can better capture the global and local information of the image, and uses multiple hierarchical encoders for parallel computing, enabling the model to obtain the enhanced image faster on some devices capable of parallel computing. The model has relatively fewer parameters and faster inference speed. It specifically includes the following steps: S1. Device setup Install the camera device to ensure that the camera device can capture the video stream and display it on the host computer through transmission; after passing the captured nighttime concrete image through the state space image enhancement network and obtaining the enhanced single-frame image, display it on the host computer as shown in Figure 1 shown.

[0033] S2. Design an image enhancement algorithm composed of a multi-level encoder and a single-frame image multi-level fusion decoder module Based on the state space convolutional deep network technology, design a nighttime single-frame image enhancement algorithm composed of a state space module and a convolutional module. In the image multi-level downsampling module in S1: The pooling kernel sizes used in the multi-level downsampling module are 2×2, 4×4, and 16×16 respectively. When the input is a single-frame image of the original size, the downsampled images by 4 times, 16 times, and 256 times can be obtained respectively, which are the detail, local, and global images. Specifically, the nighttime single-frame image enhancement algorithm specifically includes the following content: 1. The multi-level downsampling module of the single-frame image: The module is composed of different pooling kernel sizes, aiming to obtain low-light images at different levels. Among them, the pooling method used is max pooling.

[0034] 2. Standard State Space Module: The state space module consists of a standardized convolutional layer and a state space image processing layer. The convolutional kernel size of the standardized convolutional layer is 1×1, and the convolutional stride is 1.

[0035] 3. Detail Encoder Module: The detail encoder module consists of a standard state space module.

[0036] 4. Local Encoder Module: The local encoder module consists of three standard state space modules and is concatenated at the last layer.

[0037] 5. Global Encoder Module: The global encoder module consists of a standard state space module and a state space module composed of two state space image processing layers.

[0038] 6. Multi-level Fusion Decoder Module: In the global encoder module, the output after passing through a standard state space module, the output of the global encoder module, and the output not passing through the global encoder module are concatenated by the number of channels. Then, the concatenated result is upsampled by a factor of four. Finally, the number of channels is reduced through a convolutional layer. Then, the result is concatenated with the output of the local encoder module by the number of channels. The concatenated result is passed through two convolutional layers to expand the convolutional channels, and then concatenated with the output of the detail encoder module. Finally, it is upsampled by a factor of two to obtain the enhanced image.

[0039] Furthermore, it should be noted that the multi-level encoder described in S1 consists of a standard state space module, a detail encoder module, a local encoder module, and a global encoder module; Among them, the multi-level encoder structure adopts a pyramid-parallelizable encoder structure. The length and width of the input original image are scaled according to the levels of 4 times, 8 times, and 16 times to obtain the detail, local, and global images. When processing the image at each level, each level is an independent encoder structure; The data processing flow of the standard state space module is as follows: The input image is divided into non-overlapping patches of 1×1, and then the dimension of the image is mapped to C′。The generated embedded image is normalized using hierarchical normalization and then fed back to the state space module for feature extraction. The state space module is divided into two branches: the first branch is processed through a linear layer and an activation function, and the second branch is processed through a linear layer, depthwise separable convolution, and an activation function, and then enters the 2D-Selective-Scan (SS2D) for processing. SS2D consists of three steps: (1) Flatten a two-dimensional feature into a one-dimensional vector along four different directions (upper left, lower right, lower left, upper right). (2) Feed the four one-dimensional vectors obtained in the previous step into the selective state space image processing layer for operation. (3) Fuse the four one-dimensional vectors into a two-dimensional feature as the output. After feature normalization, it is merged with the output of the first branch through element-wise multiplication, and then a linear layer is used to mix the features and added to the residual connection to form the output of the VSS block. LeakyReLU is used as the activation function.

[0040] The multi-level fusion decoder includes the following: First, fuse the feature information extracted from the global and local modules. Use Pixel-Shuffle to upsample the output of the global encoder module by 4 times and connect it depthwise to the output of the local encoder module, and then perform a convolution operation. Finally, upsample the convolution result by four times and connect it depthwise to the output of the detail fusion module. Then, generate the restored RGB image through a convolution layer.

[0041] S3. Build a neural network model and train it Based on the image enhancement algorithm designed in S2, build a multi-level state space convolutional depth network low-light enhancement model, and use the deep learning framework Pytorch platform to train the model. Iterate 100,000 times on the LOL-v2 training set. The learning rate starts from 1e-4 and decreases by 0.1 times after every 1,000 iterations.

[0042] It should be noted that the loss function used by the deep learning framework Pytorch platform in training the model in S2 includes the Euclidean distance loss L 2. Structural similarity loss L SSIM and color loss L color : The overall loss function described is L total : Among them, the calculation formula of SSIM is: Among them, , , ; In the above formula, μ X is X the average value of, μ Y is Y the average value of, is X and Y the standard deviation of, C 1 , C 2 are constants, is X the standard deviation of, is Y the standard deviation of, is Y the standard deviation of, is the weighting coefficient, N is the number of image pixels, is X the standard deviation.

[0043] For G high is the color distribution of a normally exposed daytime picture obtained by the trained UNet network for capturing the color distribution of an image, G pred is the color distribution of a picture with normal exposure and enhanced night image obtained by the trained UNet network for capturing the color distribution of an image, H is the height of a single-frame image, W is the width of a single-frame image.

[0044] S4. Output result Input the low-light single-frame image in the night video of the test set of the data set into the low-light enhancement model trained in S3, and output the corresponding enhanced image.

[0045] Example: As Figure 2 , the present invention proposes a multi-level night single-frame image enhancement algorithm based on a state space and convolution hybrid model, with fewer parameters and faster model inference speed compared to other deep learning methods, while hardly losing the image enhancement effect, mainly including the following steps: S1. Downsample a low-light image H with a length of W and a width of P in the RGB three channels by 2 times, 4 times, and 16 times respectively in terms of length and width, and output the 2-times downsampled picture P 2x , the 4-times downsampled picture P4x 、 16x downsampled image P 16x 。 The role of downsampling is to reduce the size of the image, reduce the computational load, and can extract features from images at three different levels, obtaining feature information at different levels. Improve the generalization ability of the model.

[0046] S2. Feature extraction is performed on images at different levels through different encoder structures, and the parameters of each level encoder are shown in Table 1: Table 1 And the specific content of the encoders at different levels is as follows: S20. 2x downsampled image P 2x First, a detail encoder is used to extract the feature information of the image. The structure of this detail encoder is a standard state space module, which contains two main parts: the first part is a 1×1 convolutional layer responsible for extracting the feature information of the image; the second part is a state space image processing layer for further processing and optimizing the extracted features. Such a structure design not only effectively reduces the number of model parameters, but also ensures the efficiency and quality of feature extraction. Through such processing, the obtained image features can be used in subsequent image analysis or image restoration tasks, improving the performance and effect of image processing.

[0047] S21. 4x downsampled image P 4x Image features are obtained through a local encoder. The local encoder consists of three standard state space modules and a residual mapping. In this encoder, the input image is first processed by the first standard state space module, and then the output is sequentially passed through the other two standard state space modules. After being processed by each standard state space module, the outputs at each stage are concatenated according to their number of channels.

[0048] S22. 16x downsampled image P 16x Image features are obtained through a global encoder.

[0049] The global encoder module consists of a standard state space module and a state space module composed of two state space image processing layers.

[0050] The specific details of the above-mentioned state space image processing layer are as follows: As Figure 3 shown in the state space image processing layer, considering the input , x as the input sequence, Represents a tensor with C′ channels, where the height of a single-frame image for each channel is H , and the width is W .

[0051] Establish the following continuous state-space model : , ; where, in the continuous state space, is the intermediate state, is the input sequence, is the output sequence, A , B , C are learnable parameter matrices, A with size ( D , N ), B with size ( b, l, N ), C with size ( b, l, N ), where D is the dimension of the input vector, N is the dimension of the hidden layer, b is the batch size, l is the sequence length.

[0052] Update the parameters continuously with the number of iterations, and discretize the above continuous state space, which can be expressed as: , , where , are the discretized learnable parameter matrices, h k is the intermediate state at time k , h k-1 is the intermediate state at time k - 1, x k is the input sequence at time k ; the discretization rule is: , where ∆ is the step size, with size ( b, l, N ), and , h 0 is the initial intermediate state, x 0 is the initial input sequence, is the identity matrix, and exp(·) is the exponential function.

[0053] Through the derivation of multiple temporal inputs, the convolutional representation can be obtained: Among them, , is used to represent a learnable parameter matrix. Usually is a HiPPO matrix, where , , .

[0054] In order to make the model have selectivity and input dependence, the final model is: , Among them, , , are rule functions for making the discrete learnable matrix , , dynamically change with input parameters, that is: , , , x t is t the input at time h t-1 is t the intermediate state at time -1, h t is t the intermediate state at time . The matrices A , B , C and ∆ are all learnable parameters and are continuously iterated during training. The input image is divided into non-overlapping 1×1 patches, and then the dimensions of the image are mapped to C′ . The generated embedded image is normalized using layer normalization and then fed back to the state space module for feature extraction. As shown in Figure 4 the final architecture diagram ( Figure 4 in, the multiplication sign represents element-wise matrix multiplication, and σ represents the SiLu activation function), the state space processing module is divided into two branches: the first branch is processed by a linear mapping layer and the SiLu activation function, and the second branch is processed by a linear mapping layer, depthwise separable convolution, and the SiLu activation function, and then enters the state space model for processing.

[0055] S3. Concatenate the feature map of the image after passing through the global encoder module and the output without passing through the global encoder module according to the number of channels. The specific content is as follows: Assume the input is , divide the input into non-overlapping 4×4 patches, and then map the dimensions of the picture x to C′, the embedded image is generated during this process , and finally, Layer Normalization is used for x′ normalization, and then it is sent to the SSM encoder for feature extraction. The SSM encoder consists of four stages. At the end of the first three stages, patch merging operations are applied to reduce the height and width of the input features while increasing the number of channels. [2, 2, 2, 2] VSS blocks are used in the four stages, and the number of channels in each stage is C′′ , 2 C′′ , 4 C′′ , 8 C′′ . Suppose the number of channels after encoding is C 1, and the number of channels before passing through the encoder is C 2, then the final output number of channels is ( C 1 + C 2). Then, the number of channels is reduced through a convolutional layer. The concatenated result is upsampled by a factor of four to obtain the global encoder feature extraction result F high . The upsampling method is the PixelShuffle method provided by Pytorch.

[0056] S4. Concatenate the number of channels of the output F high after passing through the local encoder module, then expand the convolutional channels of the concatenated result through two convolutional layers, then concatenate it with the output after passing through the detail encoder module, and finally perform upsampling by a factor of two to obtain the enhanced image.

[0057] In the low-light image enhancement method of the present invention, the weight parameters of each designed network are continuously updated according to the training data and the loss function during the training and learning process until the parameters of each network model are saved when the training process converges. In actual application, only a low-light image and a light adjustment parameter need to be input, and the enhanced effect can be obtained at the output end.

[0058] Specifically, for a dataset containing hundreds of pairs of images with different exposure times, the normally exposed images are labeled as P h , and the low-light images are labeled as P l . The paired images P h , P l are input into the network to obtain the finally enhanced image. In the present invention, a new loss function is constructed, and this loss function includes three constraint relationships, which are specifically expressed as follows: Among them, , , , x i 、 y i respectively represent the RGB values of the pixel points output by the model and the RGB values of the pixel points of the normally exposed image. X 、 Y respectively represent the enhanced image output by the model and the normally exposed image. G high 、 G pred are respectively the color distribution of the normally exposed daytime pictures and the color distribution of the pictures enhanced for night images obtained by the UNet network trained to capture the color distribution of images. L The function of item 2 is to make the result of the model as consistent as possible with the result of the normally exposed image. L SSIM functions to not lose the structural information of the image while enhancing it. L color functions to let the model learn the color distribution of the image.

[0059] S5. Input the paired training set of the normally exposed and low-light images into the enhanced model network designed in steps 2 - 4 for training. The specific parameters for the training of the model in this example are shown in Table 2 as follows: Table 2 S6. Input the test set into the model trained in steps 2 - 4 for testing to verify the effect of the model; the training set uses the comprehensive data mixed with the original data and the enhanced data, while the data on the test set is the original data to ensure the credibility of the test effect of the model on the original data. On the LOL - v2 dataset, the training set and the test set of all models use the same set of data.

[0060] It can be seen from the above example that the present invention has the following effects: 1. The number of model parameters is low. The method of the present invention adopts a method of parallel encoding at multiple different levels. Compared with the linear encoding method of U - Net, it reduces the number of model parameters and improves the model calculation speed. Finally, the number of model parameters is about 1.40M, and the number of floating - point operations of the model is about 4.16G.

[0061] In the prior art, for example, in the KinD model with relatively good image restoration at present, the number of model parameters is 8.02M, and the number of floating - point operations of the model is about 34.99G. For the EnlightGAN model, the number of model parameters is 114.35M, and the number of floating - point operations is 61.01G.

[0062] 2. More effective loss functions The method of the present invention uses a method in which three types of loss functions interact with each other to enable the model to learn the information in the image, namely the L2 loss function, the ms-SSIM loss function, and the color loss function, improving the enhancement effect of the model.

[0063] In the prior art, in traditional methods, usually only the ms-SSIM loss and the L2 norm loss function are used as the total loss function, lacking a loss function involving image color information.

[0064] 3. Data parallel processing The present invention performs parallel processing of data on different downsampling modules through multi-level downsampling, and finally fuses the data processed in parallel by the multi-level downsampling modules through a first-level feature fusion module, improving the running speed of the model.

[0065] Prior art: Currently, a large number of image enhancement models adopt model structures such as U-Net for linear feature extraction and feature processing. This structure has also been proven to achieve good image enhancement effects, but due to linear processing, the speed of image enhancement is usually slow.

[0066] 4. Support for optimization training of multiple complex scenarios When the present invention needs to be transferred to other scenarios, it can collect the daytime and nighttime concrete videos taken in other scenarios, and on the basis of the originally trained model, fine-tune and retrain the enhancement model to adapt to the current scenario.

[0067] In the prior art, currently, the nighttime concrete image enhancement model enhances the current complex nighttime scenario, and when migrating to different scenarios, it is necessary to re-initialize the model parameters and retrain.

[0068] 5. Image global feature capture ability The present invention improves the global capture ability of the image by designing a multi-level feature extraction module and a state space model.

[0069] In the prior art, traditional image enhancement methods based on convolutional networks can only capture local features of images, and have weak global capture ability for images.

[0070] 6. Fast model inference speed Through reasonable network structure design, the model of the present method only takes 34 ms to process 1024×1024 low-light images.

[0071] In the prior art, traditional methods may have problems such as large model size and slow inference speed.

[0072] 7. Continuous optimization and improvement The present invention can conduct in-depth data analysis by collecting datasets of concrete falling during the day and night in real time, and can continuously train the night monitoring videos to improve the night enhancement effect.

[0073] In the prior art, traditional methods usually cannot provide large-scale data for analysis and improvement.

[0074] The above solution is only an illustration of a preferred example, but is not limited thereto. When implementing the present invention, appropriate substitutions and / or modifications can be made according to the needs of users.

[0075] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those skilled in the art, additional modifications can be easily achieved. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to the specific details and the illustrated and described examples here.

Claims

1. An algorithm for enhancing nocturnal concrete images based on state space encoding, characterized in that, Including: S1. Using the state-space convolutional deep network technology, construct a night concrete image enhancement algorithm that combines a state-space module and a convolutional module; S2. According to the night concrete image enhancement algorithm in S1, construct and train a neural network model; S3. Input the test samples into the trained neural network model to obtain the enhanced images, and evaluate the enhancement effect through the overall loss function in the model. If the enhancement effect meets the expected requirements, it is considered that the neural network model training is completed; if not, return to the S2 stage for further adjustment; S4. Deploy the trained neural network model to the edge device to enhance the night concrete falling images through the neural network model and feedback them to the host computer for real-time display; Among them, in S2, the single-frame night image enhancement algorithm is based on the Pytorch platform of the deep learning framework for model training, and the overall loss function used during model training is L total represented by the following formula: In the above formula, , , respectively represent L 2, L SSIM , L color 's weight coefficients; L 2 is the Euclidean distance loss, L SSIM is the structural similarity loss, L color is the color loss, and L 2, L SSIM , L color are respectively represented by the following formulas: In the above formula, x i , y i respectively represent the RGB values of the pixel points of the model output and the RGB values of the pixel points of the normally exposed image, X , Y respectively represent the enhanced image output by the model and the normally exposed image, n is the number of image pixels, SSIM is the structural similarity formula, G high is the color distribution of daytime pictures that are trained to capture the color distribution of images and have normal exposure, G pred is the color distribution of pictures that are trained to capture the color distribution of images, have normal exposure, and are enhanced for night images, H is the height of a single-frame image, W is the width of a single-frame image.

2. The algorithm for enhancing nocturnal concrete images based on state space encoding as claimed in claim 1, wherein In S1, it also includes: using a camera to capture the concrete falling process during the day and night to form a video dataset, and cropping the images in this video dataset to generate an image dataset for training: In S2, the daytime concrete image dataset in the training dataset is used as the ground truth, and the nighttime concrete image dataset is used as the data to be enhanced. After initializing the neural network parameters, the neural network is iteratively trained through the overall loss function L total [[ID=3 3. The algorithm for enhancing night-time concrete images based on state space coding according to claim 1, characterized in that, The night concrete image enhancement algorithm includes a multi-level downsampling module, a multi-level encoder, and a multi-level fusion decoder, and its processing process includes: S10. Obtain images at three different levels of detail, local, and global from a single night image through the multi-level downsampling module; S11. Process the images at the three different levels of detail, local, and global obtained in S10 through the corresponding detail encoder module, local encoder module, and global encoder module in the multi-level encoder respectively to obtain preprocessed concrete images; S12. Perform splicing processing on the preprocessed images obtained by each level encoder module in S11, and perform upsampling after the splicing processing to further obtain the enhanced images.

4. The algorithm for enhancing night concrete images based on state space coding according to claim 3, characterized in that In S11, a standard convolutional layer and a state-space image processing layer are defined as a standard state-space module, and the convolutional kernel size of the standard convolutional layer is 1×1, and the convolutional stride is 1; Among them, the detail encoder module includes a standard state-space module; The local encoder module includes three standard state-space modules and a residual mapping; The global encoder module includes a standard state-space module and a state-space module, and the state-space module includes two state-space image processing layers.

5. The algorithm for enhancing night-time concrete images based on state-space coding as claimed in claim 3, wherein In S12, the splicing processing process of the multi-level fusion decoder is: S120. In the global encoder module, define the output passing through a standard state space module as Feature Image I, the output passing through the global encoder module as Feature Image II, and the output without passing through the global encoder module as Feature Image III. Concatenate Feature Image I, Feature Image II, and Feature Image III according to the number of channels, and then perform four-fold upsampling on the concatenated result. After that, reduce the number of channels through a convolutional layer to obtain the global feature image extraction result. F high ; S121. In the local encoder module, the output after passing through one standard state space module is defined as feature image Ⅴ, the output after passing through two standard state space modules is defined as feature image Ⅵ, the output after passing through three standard state space modules is defined as feature image Ⅶ, and the output without passing through the local encoder module is defined as feature image Ⅷ. Concatenate feature images Ⅴ, Ⅵ, Ⅶ, and Ⅷ according to the number of channels, and then expand the number of convolution channels of the concatenation result through two convolutional layers to obtain the feature image extraction result F high ; F middle ; S122. After splicing the F obtained in S121 with the feature image Ⅹ output by the detail encoder module, an enhanced image is obtained through two-fold upsampling. F middle After splicing the F obtained in S121 with the feature image Ⅹ output by the detail encoder module, an enhanced image is obtained through two-fold upsampling.

6. The algorithm for enhancing night-time concrete images based on state space encoding according to claim 3, wherein, The processing process of the multi-level fusion decoder includes: S130. Use Pixel-Shuffle to upsample the output of the global encoder module by 4 times and connect it to the output of the local encoder module in depth to fuse the feature information extracted from the global encoder module and the local encoder module through convolution operations; S131. Upsample the convolution result in S130 by four times, connect it to the output of the detail fusion module in depth, and then generate a restored RGB image through a convolutional layer.

7. The algorithm for enhancing night-time concrete images based on state-space coding as claimed in claim 3, wherein In S2, the calculation formula of SSIM is: Among them, , , ; In the above formula, μ X is X 's average value, μ Y is Y 's average value, is X and Y 's standard deviation, C 1. C 2 is a constant, is X 's standard deviation, is Y 's standard deviation, is Y 's standard deviation, is the weighting coefficient, N is the number of image pixels, is X standard deviation.

Citation Information

Patent Citations

  • Low-illumination image enhancement method based on convolutional neural network

    CN108447036A

  • Multi-exposure image fusion method based on multi-scale auto-encoder

    CN115689962A

  • Low-illumination image enhancement method based on feature fusion and attention embedding

    CN116797488A

  • Time domain response analysis method of floating structure based on state space model

    CN109033025A

  • Fast low-light image enhancement method based on global frequency domain filtering

    CN116777776A