A Low-Light Image Enhancement Method Based on Depthwise Separable Convolution

By using depth-separable convolution and self-calibration models in low-illumination image enhancement, the problem of low-illumination image quality in tunnel construction environments is solved, and efficient image enhancement and quality improvement is achieved.

CN115861101BActive Publication Date: 2025-06-27FUZHOU UNIV

Patent Information

Application Number
CN202211512888.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-29
Publication Date
2025-06-27
Estimated Expiration
2042-11-29

AI Technical Summary

Technical Problem

In field tunnel construction environments, low-illumination images are affected by dim light and environmental factors, resulting in low image quality and hidden information, making it difficult to achieve effective image enhancement and quality improvement.

Method used

The low-illumination image enhancement method based on depth separable convolution is adopted. By building a self-calibrated low-light image enhancement model, the depth separable convolution is used to replace the ordinary convolution layer, reducing network parameters, and improving model speed and image quality.

Benefits of technology

It realizes efficiently enhancing image quality under low illumination conditions, reduces calculation amount and inference time, and improves image details preservation and overall color correction effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115861101B_ABST
    Figure CN115861101B_ABST
Patent Text Reader

Abstract

The present invention relates to a low-light image enhancement method based on depthwise separable convolution. It includes: Step S1, collecting the monitoring video of day and night tunnel construction and segmenting it into frame images for enhancing low-light images; Step S2, building a self-calibrating low-light image enhancement model based on depthwise separable convolution, determining the model parameters and the loss function of the neural network, and optimizing the model performance to the best; Step S3, using the trained low-light image enhancement network for enhancing the brightness of dim day and night frame images to obtain the image with enhanced illumination. The illumination estimated by the method of the present invention maintains good smoothness, and very superior performance is obtained in terms of image quality and inference speed, and it can quickly improve the image quality in the complex tunnel environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence algorithms for unsupervised machine learning, and specifically to a low-light image enhancement method based on depthwise separable convolution. Background Art

[0002] In the wild scenes with narrow space, dim light, and single scene, various environmental factors such as temperature, humidity, smoke, water level, and gas composition in the tunnel will affect the imaging effect of each camera in the tunnel (for example, high humidity and corrosive gas in the tunnel). As a result, the low-light images hide information in the darkness, and weak-light image enhancement and improving image quality are extremely challenging. Problems such as controlling partial overexposed information of the image, correcting the overall color, and preserving image details are all urgent problems to be solved.

[0003] In model-based methods, they are limited to defined regularization, and most of them produce unsatisfactory results, requiring manual adjustment of a large number of parameters to adapt to real scenarios. In network-based methods, some carefully designed deep networks are not stable, especially in real scenarios where exposure is ubiquitous due to unknown details, and it is difficult to achieve continuous excellent performance. Summary of the Invention

[0004] The purpose of the present invention is to provide a low-light image enhancement method based on depthwise separable convolution. The illumination estimated by this method maintains good smoothness, which ensures that the reflectivity generated by us is more visually friendly.

[0005] To achieve the above purpose, the technical solution of the present invention is: a low-light image enhancement method based on depthwise separable convolution, including the following steps:

[0006] Step S1, collect the monitoring videos of day and night tunnel construction and segment them into frame images;

[0007] Step S2, build a self-calibrating low-light image enhancement model based on depthwise separable convolution, determine the model parameters and the loss function of the neural network used in the model, and train and optimize the model performance to the best;

[0008] Step S3, use the trained low-light image enhancement model for the dim day and night frame images to enhance the image brightness and obtain the image with enhanced illumination.

[0009] Compared with the prior art, the present invention has the following beneficial effects:

[0010] The first innovation of the present invention lies in providing an efficient and effective learning framework, which realizes accelerating the low-light image enhancement algorithm by using the learning process. By replacing ordinary convolutional layers with depthwise separable convolutions, the number of network model parameters is reduced, and very superior performance is obtained in terms of image quality and inference speed.

[0011] The second innovation of the present invention lies in opening up a new perspective (i.e., introducing an auxiliary process to improve the ability of the basic unit model during the training phase) to improve the practicality of real-world scenarios for other low-level vision problems. The constructed weight-sharing self-calibrated illumination learning process generates a gain that only uses a single basic block for inference, enabling rapid improvement of image quality in complex tunnel environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 It is a schematic diagram of the low-light image enhancement structure used in the process of the embodiment of the present invention.

[0013] Figure 2 It is a schematic diagram of the cascaded self-calibrated illumination learning structure with weight sharing used in the process of the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0014] The technical solutions of the present invention will be specifically described below with reference to the accompanying drawings.

[0015] As Figure 1-2 shown, a low-light image enhancement method based on depthwise separable convolution according to the present invention includes the following steps:

[0016] Step S1: Collect the monitoring video of day and night tunnel construction and segment it into frame images;

[0017] Step S2: Build a self-calibrated low-light image enhancement model based on depthwise separable convolution, determine the model parameters and the loss function of the neural network used in the model, and optimize the model performance to the best;

[0018] Step S3: Use the trained low-light image enhancement model for dim day and night frame images for image brightness enhancement to obtain the image with enhanced illumination.

[0019] In an embodiment of the present invention, the specific implementation of step S1 is as follows:

[0020] Collect the on-site data of the narrow construction environment inside the day and night tunnel construction site through an electronic camera, and randomly sample the video frames at a frame interval of 30 frames per second according to the time sequence to generate 10,000 tunnel construction images as a data set.

[0021] The low - illumination image is affected by high humidity and corrosive gases in the tunnel, which will affect the imaging effect of the camera and thus produce an image with poor visibility, and image enhancement is required.

[0022] In an embodiment of the present invention, step S2 is specifically implemented as follows:

[0023] Step S21: Establish a cascaded self - calibration illumination learning process with weight sharing:

[0024] The weight - sharing mechanism means using the same mapping structure H with the same parameter weights to learn illumination at each stage; among them, the mapping with parameter θ learns the residual representation u between the illumination and the low - light observation t ;

[0025] According to the theoretical consensus, that is, the illumination and the low - light observation are similar or linearly related in most regions. Past experience shows that it is quite difficult to calculate the direct mapping between the low - light observation and the illumination. Therefore, under the premise of ensuring both performance and stability, learning the residual representation can greatly reduce the computational difficulty, especially outstanding in terms of exposure control.

[0026] It is implemented through a neural network. Among them, the neural network is composed of different numbers of basic blocks stacked. Different numbers of basic blocks and different channel numbers in the basic blocks can obtain good PSNR values; in the original neural network, the basic block contains a 3×3 convolutional layer and a ReLU activation function layer, and the H structure is set with 3 channels.

[0027] The innovation of this invention patent lies in redesigning the ordinary 3×3 convolution in the original network by using depth - separable convolution to replace the convolutional layer to reduce network parameters. Each depth - separable convolutional layer consists of a depth convolution and a point convolution. The kernel size of the depth convolution is 3×3 and the stride is 1, and the kernel size of the point convolution is 1×1 and the stride is 1. The present invention reduces network parameters, improves the model speed, and further reduces the computational amount.

[0028] According to the Retinex theory, there is the following relationship between the expected recovered image z and the captured original low - illumination image y: where x represents the illumination component, denotes element - wise multiplication; the illumination component is the core component that needs to be optimized in low - light image enhancement; according to the Retinex theory, by removing the estimated illumination, the enhanced output is further obtained.

[0029] The present invention models the process of learning illumination from a progressive perspective. The basic unit of the process of learning illumination is written as:

[0030]

[0031] where the mapping with parameter θ is the network structure for learning illumination, i.e., the mapping structure H, which learns a simple residual representation between the illumination and the low-light observation; u t represents the residual term at the t-th stage, and x t represents the illumination component at the t-th stage. The range of t is between 0 and T - 1, and the same

[0032] Compared with directly mapping between the low-light observation and the illumination, the large number of learned residual representations in this paper reduce the computational difficulty, ensure both performance and stability, and are particularly reflected in the control of exposure.

[0033] The self-calibration module mentioned above refers to that it will gradually correct the input of each level by integrating physical principles, indirectly affecting the output of each level. Each shared block is expected to output a result as close as possible to the expected target. Ideally, the first block can output the required result to meet the task requirements. At the same time, the output of the subsequent block is similar to or even exactly the same as that of the first block. In this way, in the test stage, only a single block is needed to accelerate the inference speed.

[0034] During the process of re-developing the intermediate output of illumination learning, considering the computational burden of the cascade mode, on the premise of the weight sharing mechanism, a self-calibrated module is constructed and its mapping is added to the original low-light input as the input for the next-stage illumination estimation, endowing a single basic block with stronger representativeness and convergence between the results of each stage, generating a gain in using only a single basic block for inference (not exploited in previous work), greatly reducing the computational cost, and for the first time realizing the use of the learning process to accelerate the low-light image enhancement algorithm.

[0035] Directly learning illumination will cause the image to be overexposed. The process of learning the residual between the illumination and the input suppresses overexposure, but the overall image quality is still not high, especially in the grasp of details. The present invention not only suppresses overexposure but also enriches the image structure.

[0036] A self-calibration module is constructed. Every four basic blocks form a self-calibration module. Specifically, the self-calibration module is expressed as:

[0037]

[0038] where v tis the conversion input for each stage; represents element-wise multiplication; is the introduced parametric operator with learnable variable parameters ; z t represents the image restored at the t-th stage; s t represents the intermediate variable at the t-th stage;

[0039] The rewrite of the basic unit at the t-th stage (t≥1) is:

[0040]

[0041] The t-SNE distributions between the results of each stage converge to the same value;

[0042] In fact, the self-calibration module indirectly affects the output of each stage by integrating physical principles and gradually correcting the input of each stage.

[0043] Step S22, unsupervised training loss:

[0044] Considering the inaccuracy of existing paired data, the present invention uses unsupervised learning to expand the capabilities of the tunnel low-light image enhancement network.

[0045] First, to ensure pixel-level consistency between the estimated illumination and the input of each stage, a fidelity loss is defined:

[0046]

[0047] where T is the total number of stages (set to 3 in this example), and the redefined input (y + s t-1 ) is used to flexibly constrain the output illumination x t , instead of y;

[0048] Meanwhile, a light smoothness loss function is defined. According to the prior smoothness, the illumination in natural images is usually locally smooth. There are two advantages to adopting this prior in the network. First, it helps reduce overfitting and improve the generalization ability of the network. Second, it enhances the image contrast. The present invention adopts a smoothing term with an L1 norm having spatial variables, aiming to minimize the sum of the absolute differences between the target value and the estimated value. The light smoothness loss function is expressed as:

[0049]

[0050] where N represents the total number of pixels, i represents the i-th pixel, represents the neighboring pixels of i in its 5×5 window; respectively represent the illumination components of the i-th and j-th pixels at the t-th stage; Represents the weight, and its formula form is

[0051]

[0052] where c represents the image channel in the YUV color space, σ = 0.1 represents the standard deviation of the Gaussian kernel, and y i,c 、y j,c respectively represent the inputs of the c-th channel of the i-th and j-th pixels, respectively represent the intermediate variables of the c-th channel of the i-th and j-th pixels in the (t - 1)-th stage;

[0053] Thus, the total loss of unsupervised training is obtained:

[0054]

[0055] where α and β are two positive balance parameters. Thus, the ability of the image enhancement network is expanded using unsupervised learning.

[0056] In an embodiment of the present invention, the step S3 is specifically implemented as follows:

[0057] Step S31: Build a video monitoring terminal for the tunnel online system, including a camera, an IP phone, and a hard disk video recorder. The camera and the IP phone are directly connected to the Ethernet port of the 4G router; each 100 meters of the tunnel is a section area control center. Two network cameras are set at the entrance and exit of each fire prevention partition section to monitor the situation of any personnel entering the fire prevention partition; all video monitoring images are controlled and displayed through a remote monitoring platform to achieve full-range monitoring, and the monitoring images of each fire prevention partition can be switched and displayed on the monitor; the environmental scenes within the monitoring range in the tunnel are photographed all day long to complete information data collection;

[0058] Step S32: Apply the optimized low-light image enhancement model in step S2 to each pixel of the tunnel image to be enhanced to obtain an image with enhanced illumination.

[0059] The above are the preferred embodiments of the present invention. All changes made according to the technical solutions of the present invention and whose functional effects do not exceed the scope of the technical solutions of the present invention belong to the protection scope of the present invention.

Claims

1. A low-light image enhancement method based on depthwise separable convolution, characterized in that, It includes the following steps: S1. Collect the monitoring videos of day and night tunnel construction and segment them into frame images; S2. Build a self-calibrating low-light image enhancement model based on depthwise separable convolution, determine the model parameters and the loss function of the neural network used in the model, and train and optimize the model performance to the best; the specific implementation is as follows: S21. Establish a cascaded self-calibrating illumination learning process with weight sharing: The weight sharing mechanism, that is, using the same mapping structure H with the same parameter weights to learn illumination at each stage; among them, the mapping with parameter θ Learn the residual representation u between the illumination and the low-light observation t ; It is implemented through a neural network, where the neural network is composed of different numbers of basic blocks stacked on top of each other. Different numbers of channels are set for the basic blocks, and each basic block contains a depthwise separable convolutional layer and a ReLU activation function layer; each depthwise separable convolutional layer consists of a depthwise convolution and a pointwise convolution. The kernel size of the depthwise convolution is 3×3 and the stride is 1, and the kernel size of the pointwise convolution is 1×1 and the stride is 1; According to the Retinex theory, there is the following association between the expected restored image z and the captured original low-light image y: where x represents the illumination component, denotes element-wise multiplication; the illumination component is the core component that needs to be optimized in low-light image enhancement; Model the learning illumination process from a progressive perspective, and the basic unit of the learning illumination process is written as: Among them, the mapping with parameter θ is the network structure for learning illumination, i.e., the mapping structure H, which learns a simple residual representation between the illumination and the low-light observation; u t represents the residual term at the t-th stage, and x t represents the illumination component at the t-th stage. The range of t is between 0 and T - 1, and the same Construct a self-calibrating module, and every four basic blocks form a self-calibrating module, which is expressed as: Among them, v t is the conversion input for each stage; represents element-wise multiplication; is the introduced parametric operator with learnable variable parameters ; z t represents the restored image at the t-th stage; s t represents the intermediate variable at the t-th stage; The basic unit at the t (t≥1) stage is rewritten as: The t-SNE distributions between the results of each stage converge to the same value; S22. Unsupervised training loss: First, to ensure the pixel-level consistency between the estimated illumination and the input of each stage, define the fidelity loss: where T is the total number of stages, and the re - defined input (y + s t-1 ) is used to flexibly constrain the output illumination x t , replacing y; At the same time, define the illumination smoothness loss function, that is, adopt a smoothing term with the L1 norm of spatial variables, aiming to minimize the sum of the absolute differences between the target value and the estimated value. The illumination smoothness loss function is expressed as: Where N represents the total number of pixels, and i represents the i-th pixel, represents the neighboring pixels of i within its 5×5 window; respectively represent the illumination components of the i-th and j-th pixels at the t-th stage; represents the weight, and its formula form is Among them, c represents the image channel in the YUV color space, σ = 0.1 represents the standard deviation of the Gaussian kernel, y i,c and y j,c respectively represent the inputs of the c-th channel of the i-th and j-th pixels, respectively represent the intermediate variables of the c-th channel of the i-th and j-th pixels in the (t - 1)-th stage; Thus, obtain the total loss of unsupervised training: Among them, α and β are two positive balance parameters; S3. Apply the trained low-light image enhancement model to the dim day and night frame images for image brightness enhancement to obtain the image with enhanced illumination.

2. The low-light image enhancement method based on depthwise separable convolution according to claim 1, wherein, The specific implementation of S1 is as follows: Collect the on-site data of the narrow construction environment inside the day and night tunnel construction site through an electronic camera, and randomly sample the video frames at a frame interval of 30 frames per second according to the time sequence to generate 10,000 tunnel construction images as the data set.

3. A low-light image enhancement method based on depthwise separable convolution according to claim 1, characterized in that, The specific implementation of S3 is as follows: S31. Build a video monitoring terminal for the tunnel online system, including cameras, IP phones, and hard disk recorders. The cameras and IP phones are directly connected to the Ethernet ports of the 4G router; every 100 meters of the tunnel is a regional control center. Two network cameras are set at the entrances and exits of each fire compartment section to monitor the situation of any personnel entering the fire compartment; all video monitoring pictures are controlled and displayed through the remote monitoring platform to achieve full-range monitoring, and the monitoring pictures of each fire compartment can be switched and displayed on the monitor; continuously shoot the environmental scenes within the monitoring range in the tunnel to complete the collection of information data; S32. Apply the optimized low-light image enhancement model in S2 to each pixel of the tunnel image to be enhanced to obtain the image with enhanced illumination.

Citation Information

Patent Citations

  • Retinex-based progressive image enhancement method

    AU2020100175A4

  • Low-illumination image enhancement method based on improved depth separable generative adversarial network

    CN111915525A

Cited By

  • Ceramic package substrate image enhancement method based on scale state space model

    CN121788367A

  • An image enhancement method for ceramic packaging substrates based on a scale-state-space model

    CN121788367B