A multi-scale image defogging method applicable to a camera device
By generating a unique identifier in the camera device and combining the atmospheric scattering model and fog inversion model, the monitoring video is processed frame by frame and decomposed on multiple scales, solving the problems of low fog feature removal efficiency and high error rate in the monitoring video, achieving a more efficient and accurate fog removal effect.
Patent Information
- Application Number
- CN202411672516.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2044-11-21
AI Technical Summary
The existing fogging removal method has low efficiency and high error rate for fogging features in a small amount of fogging images in surveillance video, and the recovered fogging-free image is too different from the real scene details.
Through built-in rules, unique identifiers are generated, combined with atmospheric scattering model and fog inversion model, the surveillance video of the camera device is decomposed frame by frame, preprocessed, multi-scale decomposition and mist feature removal, improving fog removal efficiency and reducing error rate.
Improve the removal efficiency of fog features in the surveillance video, reduce the removal error rate, and ensure that the restored foggy-free image is closer to the details of the real scene.
Smart Images

Figure CN119599911B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a multi-scale image defogging method suitable for a camera device. Background Art
[0002] In recent years, with the continuous advancement and upgrading of digital devices such as cameras and sensors, image processing capabilities have also been greatly improved. Multi-scale image dehazing technology is one of the more advanced technologies, which uses image processing technology to remove haze and blur.
[0003] With advances in digital technology, image processing and computer vision are increasingly being applied across a wide range of fields, from industry and medicine to entertainment. In many scenarios, clear, realistic, and usable images are required to support decision-making and judgment. Haze and blur can affect image quality, necessitating multi-scale dehazing of captured images.
[0004] Patent No. CN2019100159472 discloses an image defogging method based on a multi-scale residual network. The method includes: obtaining fog-free images in different scenes to form a fog-free image dataset; extracting the depth information of the fog-free image, applying fog interference of different concentrations to the fog-free image according to the depth information of the fog-free image, obtaining a foggy image, and forming a training dataset of all foggy images obtained based on the fog-free image; constructing a multi-scale residual network, inputting the training dataset into the multi-scale residual network, training the multi-scale residual network, and obtaining a trained image defogging model; inputting the foggy image to be processed into the trained image defogging model, and the image defogging model outputs a fog-free image corresponding to the foggy image to be processed. The method of the above invention can better handle fog images of different concentrations and different scales, solve the problem of less training data, achieve better results with less training data, and is suitable for fog images of different concentrations and different scales.
[0005] Patent No. CN2021102514847 discloses a multi-scale connected image dehazing algorithm based on UNet3+. The haze image dataset is input into the dehazing network, and the residual network and non-local block operation are used in each level of the encoder to extract the feature information of the current scale. Then, the downsampling operation is performed to reduce the scale of the haze image, so that the next level encoder can extract feature information of different scales. After passing through the three-level encoder in sequence, the feature information of different scales of the haze image is extracted respectively. The feature information of different scales encoded from different levels and the feature information output by the previous layer decoder are aggregated, and then the current level decoding part is completed through channel adjustment, residual network, non-local block and other operations. After the four-level decoder stage, the obtained feature map is added pixel by pixel to the original input haze map to obtain the image after removing the haze, ensuring that the dehazed image is closer to the information in the original haze-free scene.
[0006] Although the above patents can all perform defogging on images captured by the camera device, the existing defogging methods cannot accurately remove the fog features of a small number of foggy images in the surveillance video; and the restored fog-free images differ too much from the details of the real scene. Summary of the Invention
[0007] The purpose of the present invention is to provide a multi-scale image defogging method suitable for a camera device, which can generate a unique identifier by calculating the device serial number and the video reception time through built-in rules, and assign it to each frame of the image obtained by frame-by-frame decomposition, so as to reduce the error rate of removing fog features in foggy images; the fog frequency layer obtained by fitting the atmospheric scattering model is used, and the fog features in the original foggy image are removed through the fog inversion model, so as to improve the efficiency of removing fog features in foggy images.
[0008] The present invention utilizes the following technical solutions:
[0009] A multi-scale image defogging method applicable to a camera device comprises the following steps in sequence:
[0010] S1: The monitor of the camera device collects the video of the monitoring area and transmits the video of the monitoring area to the internal processor;
[0011] S2: The internal processor decomposes the monitored area video frame by frame to extract the original foggy image;
[0012] S3: preprocessing the extracted original foggy image to enhance the fog features in the original foggy image, thereby obtaining a fog-enhanced image;
[0013] S4: Decompose the fog-enhanced image into several fog frequency layers of different scales;
[0014] S5: The fog frequency layer is obtained by fitting the atmospheric scattering model, and then the fog inversion model is used to remove the fog features in the original foggy image to obtain a preliminary defogged image.
[0015] Preferably, step S1 includes the following specific steps:
[0016] S11: The camera device is powered on and started, and the monitor and internal processor are respectively initialized and verified; the camera device includes a monitor, an internal processor and a control center;
[0017] S12: If the verification passes, that is, the monitor and the internal processor are operating normally, the control center connects to the internal processor and modifies the configuration parameters of the monitor to ensure the quality of the video in the monitored area; if the verification fails, that is, the monitor and the internal processor are operating abnormally, the control center sends a reminder to the administrator and marks the abnormal camera device; the configuration parameters include resolution, frame rate and night vision mode;
[0018] S13: After the configuration parameters of the monitor are modified, the monitor starts collecting video of the monitoring area;
[0019] S14: converting the analog video signal of the monitored area into a digital video signal using a built-in or external encoder;
[0020] S15: compresses and encrypts the digital video signal and transmits it to the internal processor for further processing.
[0021] Preferably, step S2 includes the following specific steps:
[0022] S21: The internal processor receives the compressed and encrypted surveillance area video transmitted by the monitor in real time and records the video receiving time;
[0023] S22: Decrypt and decode the received surveillance area video to convert it into a digital video stream; and generate a unique identifier by calculating the original device serial number of the monitor and the original video reception time according to built-in rules;
[0024] S23: assigning the generated unique identifier to the digital video stream;
[0025] S24: temporarily storing the digital video stream in the cache of the internal processor and transmitting it to the storage space of the control center;
[0026] S25: The internal processor performs signal quality detection and image enhancement on the digital video stream in the cache;
[0027] S26: Decomposing the digital video stream into frame-by-frame images using a video decoder, and synchronizing each frame of the image with the time of the actual scene;
[0028] S27: performing fog detection on each frame of the image that has completed time synchronization to identify and mark the original foggy image;
[0029] S28: Segment and extract the marked original foggy image.
[0030] Preferably, the built-in rules are:
[0031] S2201: converting the original device serial number into binary and converting the original video receiving time into hexadecimal;
[0032] S2202: Concatenate the binary device serial number, the original device serial number, and the hexadecimal video reception time to form a first character string;
[0033] S2203: Concatenate the original video receiving time, the original device serial number, and the hexadecimal video receiving time to form a second character string;
[0034] S2204: Perform hash operations on the first character string and the second character string respectively to convert them into a first hash value and a second hash value;
[0035] S2205: intercepting the first X1 bits of the first hash value and concatenating the last Y1 bits of the second hash value, and performing a hash operation again to obtain a third hash value;
[0036] S2206: intercept the last X2 bits of the first hash value and concatenate them with the first Y2 bits of the second hash value, and perform a hash operation again to obtain a fourth hash value;
[0037] S2207: intercepting the middle Z bits of the first hash value and the second hash value, concatenating them, and performing a hash operation again to obtain a verification hash value;
[0038] S2208: Concatenate the third hash value and the fourth hash value and perform a double interpolation operation to obtain the original identifier;
[0039] S2209: Verify the uniqueness of the original identifier; if the verification passes, the original identifier is used as the unique identifier; if the verification fails, the original identifier is triple-interpolated using the verification hash value, and a character string with the same number of digits as the original identifier is randomly intercepted, and after the uniqueness is verified again, it is used as the unique identifier.
[0040] Preferably, step S3 includes the following specific steps:
[0041] S31: dividing the original foggy image into a foggy area and a fog-free area according to the mark of the original foggy image;
[0042] S32: adjusting the exposure of the foggy area and the fog-free area within a range of -3 to +3, respectively, so as to re-divide the foggy area or the fog-free area that has been incorrectly divided;
[0043] S33: adjusting the foggy area and the fog-free area to different exposures: increasing the exposure of the foggy area and decreasing the exposure of the fog-free area, thereby enhancing the fog characteristics of the foggy area;
[0044] S34: performing a denoising operation on the original foggy image with adjusted exposure using a Gaussian filtering algorithm or a mean filtering algorithm;
[0045] S35: converting the denoised original foggy image from the original RGB color space to the HSV color space;
[0046] S36: In the HSV color space, the saturation and brightness of the original foggy image are increased to further enhance the fog characteristics;
[0047] S37: Using an edge detection algorithm, the edges of fog features in the original foggy image that has undergone color space processing are detected and marked, thereby obtaining a fog-enhanced image.
[0048] Preferably, in step S4, the fog-enhanced image is first mapped to a multi-scale domain constructed by a wavelet function; in the multi-scale domain, the wavelet function performs convolution calculation on the fog-enhanced image through scaling and translation operations to complete the multi-scale refinement of the fog-enhanced image, and then obtain the multi-scale coefficients of the original fog-enhanced image; at the same time, the fog-enhanced image is scaled and resampled using sliding windows with window sizes of k×k, k=2n-1, n=1,2,3,…,12, and a Hadamard operation is performed on the same-scale part in the multi-scale coefficient to obtain a first data layer; then the fog-enhanced image is Fourier transformed and then mapped to a rotationally symmetric coordinate system. , thereby obtaining different frequency components of the fog-enhanced image. Subsequently, the fog-enhanced image is scaled and resampled using a sliding window with a window size of k×k, k=2n, n=1,2,3,…,12, and a Hadamard operation is performed on the same-frequency parts in different frequency components to obtain a second data layer. The residual scale part in the multi-scaling coefficient is then mapped to the residual frequency part in the different frequency components to obtain a mapping data layer. Finally, the fog-enhanced image is converted to the logarithmic domain and fused and spliced with the first data layer, the second data layer, and the mapping data layer to complete the decomposition of the fog frequency layer in the fog-enhanced image at several different scales.
[0049] Preferably, in step S5, the atmospheric scattering model is used to fit the fog frequency layer to obtain morphological artifacts, fog line sets, and atmospheric light values. First, the fog-enhanced image I is obtained by performing a Hadamard operation on the fog-free image J and the fog frequency layer W: The fog frequency layer W is determined according to the atmospheric light value A, the atmospheric scattering coefficient β and the camera depth d: W = A (1-e -βd )β; then the original foggy image P k Perform double gamma correction to generate several foggy images P with different exposures a : k is the image sequence number, p is the first correction factor, γ is the second correction factor; then the obtained foggy images P a According to the weight w a Perform fusion to obtain the fused foggy image Q: M represents the total number of foggy images, and the weight w is obtained by weighted multiplication of contrast, saturation and light polarization. a ; Then use the fused foggy image Q to construct Gaussian pyramid G and Laplacian pyramid L respectively, multiply the minimum scale layer in Gaussian pyramid G with the second data layer and then add them together to obtain the fuzzy layer X b At the same time, all scale layers of the Laplace pyramid L and the first data layer are Hadamard-operated and upsampled to obtain the detail layer X d ; Finally, the fog frequency layer W, blur layer X b and detail layer X d An atmospheric scattering model is fitted to determine morphological artifacts, fog line aggregation, and atmospheric light values in the fog frequency layer.
[0050] Preferably, in step S5, the process of using the fog inversion model to remove the fog features in the original foggy image to obtain the preliminary defogging image includes: inputting the morphological artifacts, fog line set and atmospheric light value into the first branch of the feature extraction network in the fog inversion model to obtain the first-layer feature matrix; inputting the first data layer, the second data layer and the mapping data layer into the second branch of the feature extraction network in the fog inversion model to obtain the second-layer feature matrix; and inputting the fog frequency layer W, the fuzzy layer X and the fog frequency layer W into the second branch of the feature extraction network in the fog inversion model to obtain the second-layer feature matrix. b and detail layer X d , input the third branch of the feature extraction network in the fog inversion model to obtain the final layer feature matrix.
[0051] Preferably, in step S5, the fog inversion model is used to remove the fog features in the original foggy image, and the process of obtaining a preliminary defogged image also includes: after splicing the first-layer feature matrix and the second-layer feature matrix, the first-layer feature matrix is input into the iterative training network in the fog inversion model for several trainings to evolve the different degrees of fog shadows in each data layer to obtain a preliminary weight matrix; at the same time, the first-layer feature matrix and the final-layer feature matrix channels are shuffled and spliced, and input into the generator in the fog inversion model to determine the fog shadow aggregation area and obtain the bias weight matrix; the preliminary weight matrix and the bias weight matrix are then assigned weight coefficients for gated fusion, and then input into the discriminator in the fog inversion model, and the global loss function is used for global update optimization to enhance the removal effect of the fog features in the original foggy image, thereby obtaining a preliminary defogged image.
[0052] Preferably, the multi-scale image defogging method further includes the following steps:
[0053] S6: Perform detail restoration and color correction on the preliminary dehazed image to obtain the final dehazed image;
[0054] S7: Replace the original foggy image in the surveillance area video with the final defogging image to achieve real-time defogging of the surveillance area video.
[0055] The present invention uses built-in rules to calculate the original device serial number and the original video reception time to generate a unique identifier, and assigns it to each frame of the image obtained by frame-by-frame decomposition. The fog frequency layer is obtained by fitting the atmospheric scattering model, and the fog features in the original foggy image are removed through the fog inversion model; the efficiency of removing fog features in foggy images is improved, and the error rate of removing fog features in foggy images is reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0057] Figure 1 The schematic diagram of the multi-scale image defogging method is applicable to camera devices. DETAILED DESCRIPTION
[0058] The present invention is described in detail below with reference to the accompanying drawings and embodiments:
[0059] like Figure 1 As shown, the multi-scale image defogging method applicable to a camera device according to the present invention comprises the following steps in sequence:
[0060] S1: The monitor of the camera device collects the video of the monitoring area and transmits the video of the monitoring area to the internal processor;
[0061] S2: The internal processor decomposes the monitored area video frame by frame to extract the original foggy image;
[0062] S3: preprocessing the extracted original foggy image to enhance the fog features in the original foggy image, thereby obtaining a fog-enhanced image;
[0063] S4: Decompose the fog-enhanced image into several fog frequency layers of different scales;
[0064] S5: The fog frequency layer is obtained by fitting the atmospheric scattering model, and then the fog inversion model is used to remove the fog features in the original foggy image to obtain a preliminary defogged image.
[0065] S6: Perform detail restoration and color correction on the preliminary dehazed image to obtain the final dehazed image;
[0066] In this embodiment, in order to avoid the uneven proportion of RGB color components, the method of adaptive segmented equalization of RGB color channel distribution is used to correct the color cast of the restored image pixel by pixel;
[0067] S7: Replace the original foggy image in the surveillance area video with the final defogging image to achieve real-time defogging of the surveillance area video.
[0068] In the present invention, step S1 includes the following specific steps:
[0069] S11: The camera device is powered on and started, and the monitor and internal processor are respectively initialized and verified; the camera device includes a monitor, an internal processor and a control center;
[0070] S12: If the verification passes, that is, the monitor and the internal processor are operating normally, the control center connects to the internal processor and modifies the configuration parameters of the monitor to ensure the quality of the video in the monitored area; if the verification fails, that is, the monitor and the internal processor are operating abnormally, the control center sends a reminder to the administrator and marks the abnormal camera device; the configuration parameters include resolution, frame rate and night vision mode;
[0071] S13: After the configuration parameters of the monitor are modified, the monitor starts collecting video of the monitoring area;
[0072] S14: converting the analog video signal of the monitored area into a digital video signal using a built-in or external encoder;
[0073] S15: compress and encrypt the digital video signal and transmit it to the internal processor for further processing;
[0074] In the present invention, step S2 includes the following specific steps:
[0075] S21: The internal processor receives the compressed and encrypted surveillance area video transmitted by the monitor in real time and records the video receiving time;
[0076] S22: Decrypt and decode the received surveillance area video to convert it into a digital video stream; and generate a unique identifier by calculating the original device serial number of the monitor and the original video reception time according to built-in rules;
[0077] In this embodiment, the built-in rules are:
[0078] S2201: converting the original device serial number into binary and converting the original video receiving time into hexadecimal;
[0079] S2202: Concatenate the binary device serial number, the original device serial number, and the hexadecimal video reception time to form a first character string;
[0080] S2203: Concatenate the original video receiving time, the original device serial number, and the hexadecimal video receiving time to form a second character string;
[0081] S2204: Perform hash operations on the first character string and the second character string respectively to convert them into a first hash value and a second hash value;
[0082] S2205: intercepting the first X1 bits of the first hash value and concatenating the last Y1 bits of the second hash value, and performing a hash operation again to obtain a third hash value;
[0083] S2206: intercept the last X2 bits of the first hash value and concatenate them with the first Y2 bits of the second hash value, and perform a hash operation again to obtain a fourth hash value;
[0084] S2207: intercepting the middle Z bits of the first hash value and the second hash value, concatenating them, and performing a hash operation again to obtain a verification hash value;
[0085] S2208: Concatenate the third hash value and the fourth hash value and perform a double interpolation operation to obtain the original identifier;
[0086] S2209: Verify the uniqueness of the original identifier; if the verification passes, the original identifier is used as the unique identifier; if the verification fails, a triple interpolation operation is performed on the original identifier using the verification hash value, and a character string with the same number of digits as the original identifier is randomly intercepted. After the uniqueness is verified again, it is used as the unique identifier;
[0087] S23: assigning the generated unique identifier to the digital video stream;
[0088] S24: temporarily storing the digital video stream in the cache of the internal processor and transmitting it to the storage space of the control center;
[0089] S25: The internal processor performs signal quality detection and image enhancement on the digital video stream in the cache;
[0090] S26: Decomposing the digital video stream into frame-by-frame images using a video decoder, and synchronizing each frame of the image with the time of the actual scene;
[0091] S27: performing fog detection on each frame of the image that has completed time synchronization to identify and mark the original foggy image;
[0092] S28: Segment and extract the marked original foggy image;
[0093] In the present invention, step S3 includes the following specific steps:
[0094] S31: dividing the original foggy image into a foggy area and a fog-free area according to the mark of the original foggy image;
[0095] S32: adjusting the exposure of the foggy area and the fog-free area within a range of -3 to +3, respectively, so as to re-divide the foggy area or the fog-free area that has been incorrectly divided;
[0096] In this embodiment, the judgment rule for incorrect classification of foggy and non-fog areas in an image is as follows:
[0097] 1. If the foggy area is misidentified, try increasing the exposure to increase the brightness and reduce the fog's effect. Specifically, increase the exposure to 1 or 2, and then re-determine the foggy and fog-free areas. If the demarcation results improve, it means that there was an error in the original determination.
[0098] 2. If the fog-free area is incorrectly determined, try reducing the exposure to reduce the brightness and increase the fog effect. Specifically, reduce the exposure to -1 or -2 and re-determine the foggy and fog-free areas. If the division results improve, it means that there was an error in the original determination;
[0099] S33: adjusting the foggy area and the fog-free area to different exposures: increasing the exposure of the foggy area and decreasing the exposure of the fog-free area, thereby enhancing the fog characteristics of the foggy area;
[0100] S34: performing a denoising operation on the original foggy image with adjusted exposure using a Gaussian filtering algorithm or a mean filtering algorithm;
[0101] S35: converting the denoised original foggy image from the original RGB color space to the HSV color space;
[0102] S36: In the HSV color space, the saturation and brightness of the original foggy image are increased to further enhance the fog characteristics;
[0103] S37: Using an edge detection algorithm to detect and mark edges of fog features in the original foggy image that has undergone color space processing, thereby obtaining a fog-enhanced image;
[0104] In the present invention, step S4 includes the following specific steps: first, mapping the fog-enhanced image to the multi-scale domain constructed by the wavelet function; in the multi-scale domain, the wavelet function performs convolution calculation on the fog-enhanced image through scaling and translation operations to complete the multi-scale refinement of the fog-enhanced image, and then obtain the multi-scale coefficients of the original fog-enhanced image; at the same time, the fog-enhanced image is scaled and resampled using sliding windows with window sizes of k×k, k=2n-1, n=1,2,3,…,12, and a Hadamard operation is performed on the same-scale part in the multi-scale coefficient to obtain a first data layer; then the fog-enhanced image is Fourier transformed and then mapped to the rotationally symmetric coordinate system, thereby obtaining different frequency components of the fog-enhanced image; then, the fog-enhanced image is scaled and resampled using a sliding window with a window size of k×k, k=2n, n=1,2,3,…,12, and a Hadamard operation is performed on the same frequency parts in different frequency components to obtain a second data layer; the residual scale part in the multi-scale coefficient and the residual frequency part in the different frequency components are then mapped to obtain a mapping data layer; finally, the fog-enhanced image is converted to the logarithmic domain and fused and spliced with the first data layer, the second data layer, and the mapping data layer to complete the decomposition of the fog frequency layer in the fog-enhanced image at several different scales;
[0105] In the present invention, in step S5, the atmospheric scattering model is used to fit the fog frequency layer to obtain morphological artifacts, fog line sets, and atmospheric light values. The process is as follows: first, the fog-free image J and the fog frequency layer W are subjected to a Hadamard operation to obtain the fog-enhanced image I: The fog frequency layer W is determined according to the atmospheric light value A, the atmospheric scattering coefficient β and the camera depth d: W = A (1-e -βd )β; then the original foggy image P k Perform double gamma correction to generate several foggy images P with different exposures a : k is the image sequence number, p is the first correction factor, γ is the second correction factor; then the obtained foggy images P a According to the weight w a Perform fusion to obtain the fused foggy image Q: M represents the total number of foggy images, and the weight w is obtained by weighted multiplication of contrast, saturation and light polarization.a ; Then use the fused foggy image Q to construct Gaussian pyramid G and Laplacian pyramid L respectively, multiply the minimum scale layer in Gaussian pyramid G with the second data layer and then add them together to obtain the fuzzy layer X b At the same time, all scale layers of the Laplace pyramid L and the first data layer are Hadamard-operated and upsampled to obtain the detail layer X d ; Finally, the fog frequency layer W, blur layer X b and detail layer X d Input atmospheric scattering model for fitting to determine morphological artifacts, fog line collection and atmospheric light value in fog frequency layer;
[0106] In this embodiment, morphological artifacts refer to various forms of images that appear in the image even though the originally photographed object does not exist. Morphological artifacts include motion artifacts, aliasing artifacts, wrapping artifacts, and magnetic susceptibility artifacts. The fog line set is grouped according to the azimuth and polar angle of the pixels and is determined using a uniform distribution of a two-dimensional histogram within a certain range.
[0107] The atmospheric light value is first calculated by taking the maximum value of the RGB channels of each pixel in the image as the initial atmospheric light. Next, a morphological closing operation is performed on the initial atmospheric light to eliminate small holes and tiny highlights caused by fog in the image. Finally, a cross-bilateral smoothing filter is used to smooth the image while preserving edge information to obtain the final atmospheric light value.
[0108] In the present invention, in step S5, the fog inversion model is used to remove the fog features in the original foggy image to obtain a preliminary defogging image. The process includes: first, the morphological artifacts, fog line set and atmospheric light value are input into the first branch of the feature extraction network in the fog inversion model to obtain the first-layer feature matrix; and the first data layer, the second data layer and the mapping data layer are input into the second branch of the feature extraction network in the fog inversion model to obtain the second-layer feature matrix; at the same time, the fog frequency layer W, the fuzzy layer X b and detail layer X d , input the third branch of the feature extraction network in the fog inversion model to obtain the final layer feature matrix; then, after splicing the first layer feature matrix and the second layer feature matrix, input the iterative training network in the fog inversion model for several trainings to evolve the different degrees of fog shadows of each data layer and obtain the preliminary weight matrix; at the same time, the first layer feature matrix and the final layer feature matrix channels are shuffled and spliced, and input into the generator in the fog inversion model to determine the fog shadow gathering area and obtain the bias weight matrix; then, after assigning weight coefficients to the preliminary weight matrix and the bias weight matrix for gated fusion, input the discriminator in the fog inversion model, and use the global loss function for global update optimization to enhance the removal effect of fog features in the original foggy image, thereby obtaining a preliminary defogging image;
[0109] In this embodiment, the first branch Y of the feature extraction network 1 It includes one 3×3 convolutional layer, two 5×5 convolutional layers, two convolutional blocks consisting of a 3×1 convolutional layer and a 1×3 convolutional layer, and two 3×3 average pooling layers:
[0110]
[0111] in, represents two 1×3 convolutional layers with a stride of 1 in series, represents two 3×1 convolutional layers with a stride of 1 connected in series, represents two series-connected 3×3 average pooling layers, represents a 3×3 convolutional layer with a stride of 2, represents a 5×5 convolutional layer with a stride of 2, Mp represents the morphological artifact, Al represents the fog line set, and Fg represents the atmospheric light value;
[0112] The second branch Y of the feature extraction network 2 It consists of three 3×3 convolutional layers, two 5×5 convolutional layers, and two 5×5 maximum pooling layers:
[0113]
[0114] The third branch Y of the feature extraction network 3 It consists of two 3×3 convolutional layers, one 5×5 convolutional layer, two 1×1 convolutional layers, and two 5×5 maximum pooling layers:
[0115]
[0116] Among them, Model 0 Represents the feature extraction network, Relu represents the ReLU activation function, Fl represents the first data layer, Sl represents the second data layer, Dl represents the mapping data layer, represents two 5×5 maximum pooling layers with a stride of 2 connected in series, represents a 3×3 convolutional layer with a stride of 2, represents two 1×1 convolutional layers with a stride of 2 connected in series, Represents a 5×5 convolutional layer with a stride of 2;
[0117] The iteratively trained network consists of two 3×3 convolutional layers, two InceptionV3 blocks, one mish activation function, two 3×3 max pooling layers, two 5×5 adaptive pooling layers, and one residual block:
[0118]
[0119] Among them, Model1 Represents iterative training network, AutoPool 5×5 () represents a 5×5 adaptive pooling layer, ResBlock() represents a residual block, represents the Hadamard operation;
[0120] The global loss function is:
[0121] Among them, L represents the global loss function, Gen represents the generator, G represents the quantization function, σ represents the number of iterative training network layers, Dis represents the discriminator, ρ represents the adaptive quantization parameter, E σ Indicates the optimization efficiency of iterative training for each layer, E0 indicates the maximum value of the preset iterative training optimization efficiency, τ indicates the total number of iterative training layers, μ1 indicates the calculation accuracy coefficient, and B σ Represents the network loss function for iterative training of each layer, and μ2 represents the preset optimization weight coefficient;
[0122] Example:
[0123] The camera device is powered on and started, and the monitor and internal processor are initialized and verified respectively. If the verification passes, that is, the monitor and internal processor are operating normally, the control center connects to the internal processor and corrects the monitor's configuration parameters (resolution, frame rate and night vision mode) to ensure the shooting quality of the video in the monitored area. If the verification fails, that is, the monitor and internal processor are operating abnormally, the control center sends a reminder to the manager and marks the abnormal camera device. After the monitor's configuration parameters are corrected, the monitor starts to collect video in the monitored area. The analog signal of the video in the monitored area is converted into a digital video signal using a built-in or external encoder. The digital video signal is compressed and encrypted, and transmitted to the internal processor for further processing.
[0124] The internal processor receives the compressed and encrypted surveillance area video transmitted by the monitor in real time and records the video reception time; decrypts and decodes the received surveillance area video to convert it into a digital video stream; and generates a unique identifier by calculating the monitor's device serial number and the video reception time according to the built-in rules; assigns the generated unique identifier to the digital video stream; temporarily stores the digital video stream in the cache of the internal processor and transmits it to the storage space of the control center; the internal processor performs signal quality detection and image enhancement on the digital video stream in the cache; uses the video decoder to decompose the digital video stream into frame-by-frame images, and synchronizes each frame image with the time of the actual scene; performs fog detection on each frame image that has completed time synchronization to identify and mark the original foggy image; and performs segmentation and extraction on the marked original foggy image;
[0125] The original foggy image is divided into foggy areas and fog-free areas according to the markings of the original foggy image; the exposure of the foggy area and the fog-free area is adjusted in the range of -3 to +3 respectively to re-divide the foggy area or fog-free area that has been misdivided; the foggy area and the fog-free area are adjusted to different exposures respectively: the exposure of the foggy area is increased, and the exposure of the fog-free area is reduced, thereby enhancing the fog characteristics of the foggy area; the original foggy image with adjusted exposure is denoised using a Gaussian filtering algorithm or a mean filtering algorithm; the denoised original foggy image is converted from the original RGB color space to the HSV color space; in the HSV color space, the saturation and brightness of the original foggy image are increased to further enhance the fog characteristics; the edge detection algorithm is used to detect and mark the edges of the fog characteristics in the original foggy image that has completed color space processing, thereby obtaining a fog-enhanced image;
[0126] First, the fog-enhanced image is mapped to the multi-scale domain constructed by the wavelet function; in the multi-scale domain, the wavelet function performs convolution calculation on the fog-enhanced image through scaling and translation operations to complete the multi-scale refinement of the fog-enhanced image, and then obtain the multi-scale coefficients of the original fog-enhanced image; at the same time, the fog-enhanced image is scaled and resampled using sliding windows with window sizes of k×k, k=2n-1, n=1,2,3,…,12, and the Hadamard operation is performed on the same-scale part in the multi-scale coefficients to obtain the first data layer; then the fog-enhanced image is Fourier transformed and then mapped to a rotationally symmetric coordinate system to obtain Different frequency components of the fog-enhanced image are then scaled and resampled using a sliding window with a window size of k×k, k=2n, n=1, 2, 3, …, 12, and a Hadamard operation is performed on the same frequency components in different frequency components to obtain a second data layer; the residual scale components in the multi-scaling coefficients are then mapped to the residual frequency components in the different frequency components to obtain a mapping data layer; finally, the fog-enhanced image is converted to the logarithmic domain and fused with the first data layer, the second data layer, and the mapping data layer to complete the decomposition of the fog frequency layer in the fog-enhanced image at several different scales;
[0127] Next, the atmospheric scattering model is used to fit the fog frequency layer to obtain morphological artifacts, fog line sets, and atmospheric light values. The process is as follows: First, the fog-enhanced image I is obtained by performing a Hadamard operation on the fog-free image J and the fog frequency layer W: The fog frequency layer W is determined according to the atmospheric light value A, the atmospheric scattering coefficient β and the camera depth d: W = A (1-e -βd )β; then the original foggy image P k Perform double gamma correction to generate several foggy images P with different exposures a : k is the image sequence number, p is the first correction factor, γ is the second correction factor; then the obtained foggy images P a According to the weight w a Perform fusion to obtain the fused foggy image Q: M represents the total number of foggy images, and the weight w is obtained by weighted multiplication of contrast, saturation and light polarization. a ; Then use the fused foggy image Q to construct Gaussian pyramid G and Laplacian pyramid L respectively, multiply the minimum scale layer in Gaussian pyramid G with the second data layer and then add them together to obtain the fuzzy layer X b At the same time, all scale layers of the Laplace pyramid L and the first data layer are Hadamard-operated and upsampled to obtain the detail layer X d ; Finally, the fog frequency layer W, blur layer X b and detail layer X d Input atmospheric scattering model for fitting to determine morphological artifacts, fog line collection and atmospheric light value in fog frequency layer;
[0128] Next, the fog inversion model is used to remove the fog features in the original foggy image. The process of obtaining the preliminary defogging image is as follows: first, the morphological artifacts, fog line set and atmospheric light value are input into the first branch of the feature extraction network in the fog inversion model to obtain the first-layer feature matrix; the first data layer, the second data layer and the mapping data layer are input into the second branch of the feature extraction network in the fog inversion model to obtain the second-layer feature matrix; at the same time, the fog frequency layer W, the fuzzy layer X b and detail layer X d , input the third branch of the feature extraction network in the fog inversion model to obtain the final layer feature matrix; then, after splicing the first layer feature matrix and the second layer feature matrix, input the iterative training network in the fog inversion model for several trainings to evolve the different degrees of fog shadows of each data layer and obtain the preliminary weight matrix; at the same time, the first layer feature matrix and the final layer feature matrix channels are shuffled and spliced, and input into the generator in the fog inversion model to determine the fog shadow gathering area and obtain the bias weight matrix; then, after assigning weight coefficients to the preliminary weight matrix and the bias weight matrix for gated fusion, input the discriminator in the fog inversion model, and use the global loss function for global update optimization to enhance the removal effect of fog features in the original foggy image, thereby obtaining a preliminary defogging image;
[0129] Finally, the restored image is restored and colored pixel by pixel using the adaptive piecewise equalization method for RGB color channel distribution to obtain the final defogging image. The original foggy image in the surveillance area video is replaced with the final defogging image to achieve real-time defogging of the surveillance area video.
Claims
1. A multi-scale image defogging method suitable for a camera device, characterized by: The method comprises the following steps in sequence: S1: The monitor of the camera device collects the video of the monitoring area and transmits the video of the monitoring area to the internal processor; S2: The internal processor decomposes the monitored area video frame by frame to extract the original foggy image; S3: preprocessing the extracted original foggy image to enhance the fog features in the original foggy image, thereby obtaining a fog-enhanced image; S4: Decompose the fog-enhanced image into several fog frequency layers of different scales; S5: The fog frequency layer is obtained by fitting the atmospheric scattering model, and then the fog inversion model is used to remove the fog features in the original foggy image to obtain a preliminary defogged image. In step S5, the fog inversion model is used to remove the fog features in the original foggy image to obtain a preliminary defogging image. The process includes: inputting the morphological artifacts, fog line set and atmospheric light value into the first branch of the feature extraction network in the fog inversion model: 1 Convolutional layer, 2 Convolutional layer, 2 by Convolutional layer and The convolutional layer consists of a convolutional block and two Average pooling layer to obtain the first-layer feature matrix; The first data layer, the second data layer, and the mapped data layer are input into the second branch of the feature extraction network in the fog inversion model: 3 Convolutional layer, 2 Convolutional layer and 2 Maximum pooling layer, obtains the sub-layer feature matrix; Among them, the fog-enhanced image is first mapped to the multi-scale domain constructed by the wavelet function. In the multi-scale domain, the wavelet function performs convolution calculation on the fog-enhanced image through scaling and translation operations to complete the multi-scale refinement of the fog-enhanced image, and then obtain the multi-scale coefficients of the original fog-enhanced image; at the same time, the window size is used to The fog-enhanced image is scaled and resampled by a sliding window of , and the Hadamard operation is performed on the same-scale part in the multi-scale coefficient to obtain the first data layer; the fog-enhanced image is then Fourier transformed and then mapped to a rotationally symmetric coordinate system to obtain different frequency components of the fog-enhanced image; then, the window size is used to The fog-enhanced image is scaled and resampled using a sliding window, and a Hadamard operation is performed on the same frequency part in different frequency components to obtain a second data layer; the remaining scale part in the multi-scale coefficient and the remaining frequency part in the different frequency components are mapped to obtain a mapping data layer; At the same time, the fog frequency layer , fuzzy layer and detail layers , input the third branch of the feature extraction network in the fog inversion model: 2 Convolutional layer, 1 Convolutional layer, 2 Convolutional layer and 2 Maximum pooling layer to obtain the final layer feature matrix; Among them, several foggy images are obtained According to weight Perform fusion to obtain a fused foggy image , by weighted multiplication of contrast, saturation and light polarization to obtain the weight ; Reuse the fused foggy image Construct Gaussian pyramids separately and Laplace pyramid , the Gaussian pyramid The minimum scale layer is multiplied by the second data layer and then added to obtain the fuzzy layer , and at the same time the Laplace pyramid All scale layers in the first data layer are subjected to Hadamard operation and upsampling to obtain the detail layer. ; After concatenating the first-layer feature matrix and the second-layer feature matrix, they are input into the iterative training network in the fog inversion model: 2 Convolutional layer, 2 block, 1 mish activation function, 2 Max pooling layer, 2 The adaptive pooling layer and one residual block are trained several times to evolve the different degrees of fog in each data layer and obtain the preliminary weight matrix; At the same time, the first-layer feature matrix and the final-layer feature matrix channels are shuffled and concatenated, and then input into the generator in the fog inversion model to determine the fog shadow gathering area and obtain the bias weight matrix; After assigning weight coefficients to the preliminary weight matrix and the bias weight matrix for gated fusion, they are input into the discriminator in the fog inversion model and globally updated and optimized using the global loss function to enhance the removal effect of fog features in the original foggy image, thereby obtaining a preliminary defogged image.
2. The multi-scale image defogging method applicable to a camera device according to claim 1, characterized in that: The step S1 includes the following specific steps: S11: The camera device is powered on and started, and the monitor and internal processor are respectively initialized and verified; the camera device includes a monitor, an internal processor and a control center; S12: If the verification passes, that is, the monitor and the internal processor are operating normally, the control center connects to the internal processor and corrects the configuration parameters of the monitor to ensure the shooting quality of the video in the monitored area; If the verification fails, meaning the monitor and internal processor are operating abnormally, the control center will send a reminder notification to the administrator and mark the abnormal camera device; configuration parameters include resolution, frame rate, and night vision mode; S13: After the configuration parameters of the monitor are modified, the monitor starts collecting video of the monitoring area; S14: converting the analog video signal of the monitored area into a digital video signal using a built-in or external encoder; S15: compresses and encrypts the digital video signal and transmits it to the internal processor for further processing.
3. The multi-scale image defogging method for a camera device according to claim 1, wherein: The step S2 includes the following specific steps: S21: The internal processor receives the compressed and encrypted surveillance area video transmitted by the monitor in real time and records the video receiving time; S22: Decrypt and decode the received surveillance area video to convert it into a digital video stream; and generate a unique identifier by calculating the original device serial number of the monitor and the original video reception time according to built-in rules; S23: assigning the generated unique identifier to the digital video stream; S24: temporarily storing the digital video stream in the cache of the internal processor and transmitting it to the storage space of the control center; S25: The internal processor performs signal quality detection and image enhancement on the digital video stream in the cache; S26: Decomposing the digital video stream into frame-by-frame images using a video decoder, and synchronizing each frame of the image with the time of the actual scene; S27: performing fog detection on each frame of the image that has completed time synchronization to identify and mark the original foggy image; S28: Segment and extract the marked original foggy image.
4. The multi-scale image defogging method applicable to a camera device according to claim 3, characterized in that: The built-in rules are: S2201: converting the original device serial number into binary and converting the original video receiving time into hexadecimal; S2202: Concatenate the binary device serial number, the original device serial number, and the hexadecimal video reception time to form a first character string; S2203: Concatenate the original video receiving time, the original device serial number, and the hexadecimal video receiving time to form a second character string; S2204: Perform hash operations on the first character string and the second character string respectively to convert them into a first hash value and a second hash value; S2205: intercepting the first X1 bits of the first hash value and concatenating the last Y1 bits of the second hash value, and performing a hash operation again to obtain a third hash value; S2206: intercept the last X2 bits of the first hash value and concatenate them with the first Y2 bits of the second hash value, and perform a hash operation again to obtain a fourth hash value; S2207: intercepting the middle Z bits of the first hash value and the second hash value, concatenating them, and performing a hash operation again to obtain a verification hash value; S2208: Concatenate the third hash value and the fourth hash value and perform a double interpolation operation to obtain the original identifier; S2209: Verify the uniqueness of the original identifier; if the verification passes, the original identifier is used as the unique identifier; if the verification fails, the original identifier is triple-interpolated using the verification hash value, and a character string with the same number of digits as the original identifier is randomly intercepted, and after the uniqueness is verified again, it is used as the unique identifier.
5. The multi-scale image defogging method applicable to a camera device according to claim 1, characterized in that: The step S3 includes the following specific steps: S31: dividing the original foggy image into a foggy area and a fog-free area according to the mark of the original foggy image; S32: adjusting the exposure of the foggy area and the fog-free area within a range of -3 to +3, respectively, so as to re-divide the foggy area or the fog-free area that has been incorrectly divided; S33: adjusting the foggy area and the fog-free area to different exposures: increasing the exposure of the foggy area and decreasing the exposure of the fog-free area, thereby enhancing the fog characteristics of the foggy area; S34: performing a denoising operation on the original foggy image with adjusted exposure using a Gaussian filtering algorithm or a mean filtering algorithm; S35: converting the denoised original foggy image from the original RGB color space to the HSV color space; S36: In the HSV color space, the saturation and brightness of the original foggy image are increased to further enhance the fog characteristics; S37: Using an edge detection algorithm, the edges of fog features in the original foggy image that has undergone color space processing are detected and marked, thereby obtaining a fog-enhanced image.
6. The multi-scale image defogging method applicable to a camera device according to claim 1, characterized in that: In step S4, the fog-enhanced image is converted into a logarithmic domain and then fused and spliced with the first data layer, the second data layer, and the mapping data layer to complete the decomposition of the fog frequency layer in the fog-enhanced image at several different scales.
7. The multi-scale image defogging method applicable to a camera device according to claim 1 or 6, characterized in that: In step S5, the fog frequency layer is fitted using the atmospheric scattering model to obtain the morphological artifacts, fog line set and atmospheric light value. The process is as follows: first, the fog enhanced image is By Fog Free Image and fog frequency layer After Hadamard operation, we get: ;Fog frequency layer According to atmospheric light value , atmospheric scattering coefficient and camera depth Sure: ; Then the original foggy image Perform double gamma correction to generate several foggy images with different exposures : , is the image sequence number, is the first correction factor, is the second correction factor; then obtain the fused foggy image : , Indicates the total number of foggy images; finally, the fog frequency layer , fuzzy layer and detail layers An atmospheric scattering model is fitted to determine morphological artifacts, fog line aggregation, and atmospheric light values in the fog frequency layer.
8. The multi-scale image defogging method applicable to a camera device according to claim 1, characterized in that: The multi-scale image defogging method further comprises the following steps: S6: Perform detail restoration and color correction on the preliminary dehazed image to obtain the final dehazed image; S7: Replace the original foggy image in the surveillance area video with the final defogging image to achieve real-time defogging of the surveillance area video.