Low-illumination image enhancement method based on learnable semantic prior

By adopting a learning semantic prior method in low-light image enhancement, predicting and integrating the semantic information of the image, the problem of unnatural low-light image enhancement in the prior art is solved, and higher quality image generation is achieved.

CN120147593APending Publication Date: 2025-06-13XIAN UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510229017.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing low-light image enhancement methods fail to make full use of the complete semantic priors of visual elements in low-light environments, resulting in unnatural or distorted images generated.

Method used

Low-light image enhancement is guided through learnable predictive semantic prior tasks using a learningable predictive semantic prior task. The specific steps include: the input is divided into two branches, one for restoring and reconstruction of image content, and the other for predicting semantic priors of low-light images; using a semantic learner to learn and predict semantic prior information of the image; integrating semantic priors through a semantic perception module, gradually recovering the details of the image, and finally generating an enhanced visible light image.

Benefits of technology

By learning and integrating semantic prior information, the generation quality of low-light images is significantly improved, making the enhanced images more natural and realistic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147593A_ABST
    Figure CN120147593A_ABST
Patent Text Reader

Abstract

The invention discloses a low-illumination image enhancement method based on learnable semantic prior, and the method comprises the steps: firstly dividing the input into two branches, taking a low-light visible light image as the input of the two branches, enabling one branch to be used for the restoration and reconstruction of image contents, and enabling the other branch to be used for predicting the semantic prior of the low-light image; in the first branch, performing enhancement processing on the input low-light visible light image; in the second branch, learning and predicting semantic prior information of the image; and gradually recovering details of the image by using semantic priori acquired from a semantic learning device, and finally generating an enhanced visible light image. And finally, designing a loss function to guide a training process of a low-light image algorithm. Concealed details in the low-light image are disclosed through the learnable semantic priori prediction task, so that the generation quality of the image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer digital image processing, and particularly relates to a low-light image enhancement method based on a learnable semantic prior. Background Art

[0002] Due to environmental factors or technical limitations, low-light imaging is very common. Low-light images not only exhibit poor visibility in visual perception but also face many challenges in processing high-quality images in multimedia computing. Low-light image enhancement (LLIE) technology reveals the hidden details in low-light images, thereby improving the generated quality of the images. However, current methods have not fully utilized the complete semantic prior of visual elements in low-light environments. Therefore, the images generated by these low-light image enhancement methods often look unnatural or distorted. To address this issue, a method for guiding low-light image enhancement through a learnable predictive semantic prior task is proposed. Summary of the Invention

[0003] The object of the present invention is to provide a low-light image enhancement method based on a learnable semantic prior, which reveals the hidden details in low-light images through a learnable predictive semantic prior task, thereby improving the generated quality of the images.

[0004] The technical solution adopted by the present invention is that the low-light image enhancement method based on a learnable semantic prior is specifically implemented according to the following steps:

[0005] Step 1: The input is divided into two branches, and both branches take a low-light visible light image as the input. One branch is used for the restoration and reconstruction of the image content, and the other branch is used for predicting the semantic prior of the low-light image;

[0006] Step 2: In the first branch, the input low-light visible light image is enhanced.

[0007] Step 3: In the second branch, the semantic prior information of the image is learned and predicted.

[0008] Step 4: Using the semantic prior obtained from the semantic learner, gradually restore the details of the image, and finally generate an enhanced visible light image;

[0009] Step 5: Design a loss function to guide the training process of the low-light image algorithm.

[0010] The characteristics of the present invention also lie in that

[0011] Step 1 is specifically implemented according to the following steps:

[0012] The branch for predicting the semantic prior of the low-light image, abbreviated as the semantic learner, is composed of three downsampling layers and five residual blocks. The semantic learner takes the low-light image Ilow As the input, the output is a feature map of size for learning the prior knowledge of low-light images. During the training phase, the high-quality standard image GT is input into the pre-trained BiSeNet, and after 1x1 convolution to adjust the number of channels, a feature map of size is obtained as the supervision signal for the semantic learner. The feature map S output by the semantic learner and the semantic prior output by the pre-trained model jointly constitute the semantic reconstruction loss. During the test phase, the semantic prior of the low-light image is output through this branch.

[0013] Step 2 is specifically implemented according to the following steps:

[0014] Step 2.1: Receive the low-light image as the input. The height of the image is H and the width is W. Denote the low-light image as I low ;

[0015] Step 2.2: Use three downsampling convolutional modules for feature encoding. Each convolutional module extracts features from the image through convolution operations and gradually reduces the image size through the stride. This process is expressed as:

[0016] F I = E I (I low ) (1)

[0017] where H and W respectively represent the height and width of the image, I low represents the encoder of the low-light image, E I represents the encoder of the low-light image, and F I represents the feature map in the encoding stage;

[0018] Step 2.3: Through the encoding process, the obtained feature map F I is used as the input for the subsequent image enhancement or semantic task branch, providing basic feature support for subsequent processing.

[0019] Step 3 is specifically implemented according to the following steps:

[0020] Step 3.1: To ensure that the semantic learner can effectively learn the semantic information in the low-light image, use the semantic segmentation model BiSeNet as the semantic knowledge base;

[0021] Step 3.2: The structure of the semantic learner takes a low-light image as input and consists of three downsampling layers and five residual blocks. Specifically, first, the input low-light image passes through three downsampling layers in sequence. Each downsampling layer uses a 4×4 convolution operation with a stride of 2 and a padding of 1. The number of channels gradually expands from 3 to 128. Meanwhile, the ReLU activation function is used to extract preliminary features. Subsequently, the feature map is input into 5 residual blocks. The number of input and output channels of each residual block is 128. Finally, a 1×1 convolution layer is used to compress the number of channels to the specified output channels, which is 64 by default, to generate semantic prior features. The purpose of this network is to predict semantic prior information from low-light images and provide semantic support for subsequent low-image enhancement;

[0022] Step 3.3: Input the high-quality standard image GT in the LOLv2 dataset into the pre-trained BiSeNet model, and adjust the number of channels through a 1x1 convolution operation to obtain a feature map of size as the supervision signal of the semantic learner;

[0023] Step 3.4: The feature map S output by the semantic learner and the semantic prior output by the pre-trained model jointly constitute the semantic reconstruction loss. This loss function ensures that the semantic learner can effectively learn the semantic information in the low-light image and perform reasonable image enhancement by comparing the output features with the pre-trained semantic prior;

[0024] Step 3.5: Through the above steps, the semantic learner is trained under the guidance of the semantic reconstruction loss to optimize the network weights, and finally achieve the goal of low-light image enhancement and semantic information retention.

[0025] Step 4 is specifically implemented according to the following steps:

[0026] Step 4.1: Through the 1×1 convolution layer of the semantic learner in Step 3.2, map the semantic prior knowledge and the features of the low-light image to the same dimension, and then use cross-modal similarity to calculate the semantic-aware attention map, specifically as follows:

[0027]

[0028] Among them, Conv m (.) and Conv n (.) are convolution layers, C is the feature channel, where represents the semantic-aware attention map, represents the feature map of the low-light image in the k-th interaction, represents the semantic prior feature predicted by the low-light image in the K-th interaction, I represents the first letter of the low-light image, k represents the number of interactions, taking values of 1, 2, 3, m, n represent the number of convolution layers, and s represents the semantic prior feature;

[0029] Step 4.2: Perform pixel-by-pixel weighted summation of the semantically-aware attention map and the feature map of the low-light image to generate the feature map of the next stage:

[0030]

[0031] in, The feature map representing the low-light image of the k+1th interaction;

[0032] Step 4.3: Through interactive learning with the semantic perception module, a semantically aware feature map can be generated. Used for low-light image restoration.

[0033] Step 5 is implemented according to the following steps:

[0034] The training includes two parts: semantic learner and image restoration, so it includes two parts: image restoration and reconstruction loss and semantic reconstruction loss. The image restoration and reconstruction loss includes three parts: Charbonnier loss, perceptual loss, and gradient loss.

[0035] The semantic reconstruction loss in step 5 is as follows:

[0036] The feature map S output by the semantic learner and the semantic prior output by the pre-trained model BiSeNet together constitute the semantic reconstruction loss:

[0037] L prior =|SS gt | (4)

[0038] L prior represents the semantic reconstruction loss, S is the semantic prior predicted by the semantic learner, and S gt is the semantic supervision signal obtained from the BiSeNet module;

[0039] The image restoration and reconstruction loss in step 5 is as follows:

[0040] Step 5.1, Charbonnier loss is expressed as:

[0041]

[0042] L r represents the reconstruction loss, I gt represents the high-quality standard image, i.e., the ground truth value, I en represents the enhanced image, |||| represents the L2 norm, which is used to calculate the distance or difference between the two, and κ is a smoothing term, which is a constant;

[0043] Step 5.2: Perceptual loss compares the ground truth value I through L1 loss gt and the restored image Ien VGG feature distance between:

[0044] L vgg = ||Φ(I gt ) - Φ(I en )|| (6)

[0045] where L vgg represents the perceptual loss, and Φ() represents the operation of extracting features from the VGG network;

[0046] Step 5.3: The gradient loss adopts the L1 norm loss function and is constrained by the eigenvalue of the ground truth on the three RGB channels to ensure that the enhanced low-light image is consistent with the ground truth in details:

[0047]

[0048] where L grad represents the gradient loss, represents the gradient of the high-quality standard image I gt on the c-th channel. represents the gradient of the enhanced image I en on the c-th channel, H and W are the height and width of the image respectively, and c ∈ {1, 2, 3} corresponds to the RGB channels;

[0049] The total loss function in Step 5 is:

[0050] L = L prior + αL r + βL vgg + γL grad (8)

[0051] where α, β, γ are hyperparameters.

[0052] The beneficial effect of the present invention is that, for the low-light image enhancement method based on the learnable semantic prior, a semantic learner is trained, and by performing knowledge distillation on the high-quality standard image, semantic prior features are extracted and learned. Subsequently, the present invention uses the semantic perception module to enable the model to adaptively integrate these learned semantic priors, thereby ensuring the semantic consistency of the enhanced image. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is a schematic structural diagram of the method for guiding low-light image enhancement by the learnable predicted semantic prior task of the present invention;

[0054] Figure 2 is the effect of the enhanced image generated by the present invention under different low-light scenarios in the LOL-v2 dataset. DETAILED DESCRIPTION OF THE INVENTION

[0055] The present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0056] The low-light image enhancement method of the present invention based on learnable semantic prior combines Figure 1 , and is specifically implemented according to the following steps:

[0057] Step 1: The input is divided into two branches, both of which take the low-light visible light image as the input. One branch is used for the restoration and reconstruction of the image content, and the other branch is used for predicting the semantic prior of the low-light image;

[0058] Step 1 is specifically implemented according to the following steps:

[0059] The branch for predicting the semantic prior of the low-light image, abbreviated as the semantic learner, is composed of three downsampling layers and five residual blocks. The semantic learner takes the low-light image I low as the input and outputs a feature map with a size of to learn the prior knowledge of the low-light image. In the training stage, the high-quality standard image GT is input into the pre-trained BiSeNet, and after adjusting the number of channels through 1x1 convolution, a feature map with a size of is obtained as the supervision signal of the semantic learner. The feature map S output by the semantic learner and the semantic prior output by the pre-trained model together constitute the semantic reconstruction loss. In the testing stage, the semantic prior of the low-light image is output through this branch.

[0060] Step 2: In the first branch, three downsampling convolutional modules are used for feature encoding to enhance the input low-light visible light image to restore the brightness and details of the image;

[0061] Step 2 is specifically implemented according to the following steps:

[0062] Step 2.1: Receive the low-light image as the input. The height of the image is H and the width is W. Denote the low-light image as I low ;

[0063] Step 2.2: In order to extract the basic features of the image, three downsampling convolutional modules are used for feature encoding. Each convolutional module extracts the features of the image through convolutional operations and gradually reduces the image size through the stride. This process is expressed as:

[0064] F I = E I (I low ) (1)

[0065] Among them, H and W respectively represent the height and width of the image, and I lowEncoder E for representing low-light images I Encoder F for representing low-light images I Feature map representing the encoding stage;

[0066] Step 2.3: Through the encoding process, the obtained feature map F I As the input for subsequent image enhancement or semantic task branches, it provides basic feature support for subsequent processing.

[0067] Step 3: In the second branch, input the low-light visible light image into the semantic learner network, and supervise and train the semantic learner with high-quality visible light standard images to learn and predict the semantic prior information of the image;

[0068] Step 3 is specifically implemented according to the following steps:

[0069] Step 3.1: To ensure that the semantic learner can effectively learn the semantic information in low-light images, the semantic segmentation model BiSeNet is used as the semantic knowledge base. BiSeNet is a lightweight and efficient real-time semantic segmentation network with strong global context understanding ability, which can identify the semantic information of multiple different objects in the image.

[0070] Step 3.2: The structure of the semantic learner takes low-light images as input and consists of three downsampling layers and five residual blocks. Specifically, first, input the low-light image through three downsampling layers in sequence. Each downsampling layer uses a 4×4 convolution operation with a stride of 2 and a padding of 1, and the number of channels gradually expands from 3 to 128. At the same time, the ReLU activation function is used to extract preliminary features. Subsequently, the feature map is input into 5 residual blocks, and the number of input and output channels of each residual block is 128 to enhance the feature expression ability and retain key information. Finally, the number of channels is compressed to the specified output channels, defaulting to 64, through a 1×1 convolution layer to generate semantic prior features. The purpose of this network is to predict semantic prior information from low-light images and provide semantic support for subsequent low-image enhancement;

[0071] Step 3.3: Input the high-quality standard image GT in the LOLv2 dataset into the pre-trained BiSeNet model, and adjust the number of channels through a 1x1 convolution operation to obtain a feature map of size as the supervision signal of the semantic learner. This process effectively provides semantic-level supervision information for the semantic learner to help it learn semantic priors;

[0072] Step 3.4: The feature map S output by the semantic learner and the semantic prior output by the pre-trained model together constitute the semantic reconstruction loss. This loss function ensures that the semantic learner can effectively learn the semantic information in low-light images and perform reasonable image enhancement by comparing the output features with the pre-trained semantic prior;

[0073] Step 3.5: Through the above steps, the semantic learner is trained under the guidance of the semantic reconstruction loss to optimize the network weights, and finally achieve the goal of low-light image enhancement and semantic information retention.

[0074] Step 4: In the decoder stage, using the semantic prior obtained from the semantic learner, three upsampling convolutional modules are adopted and tightly coupled through the semantic perception module. Through multi-stage cascading, combined with gradual upsampling and semantic perception guidance, the network can gradually restore the details of the image and finally generate the enhanced visible light image.

[0075] Step 4 is specifically implemented according to the following steps:

[0076] Step 4.1: Through the 1×1 convolutional layer of the semantic learner in Step 3.2, map the semantic prior knowledge and the features of the low-light image to the same dimension, and then use cross-modal similarity to calculate the attention map of semantic perception, specifically as follows:

[0077]

[0078] Among them, Conv m (.), Conv n (.) are convolutional layers, C is the feature channel, where represents the attention map of semantic perception, represents the feature map of the low-light image in the k-th interaction, represents the predicted semantic prior feature of the low-light image in the K-th interaction, I represents the first letter of the low-light image, k represents the number of interactions, taking values of 1, 2, 3, m, n represent the number of convolutional layers, and s represents the semantic prior feature;

[0079] Step 4.2: Perform pixel-wise weighted summation of the attention map of semantic perception and the feature map of the low-light image to generate the feature map of the next stage:

[0080]

[0081] Among them, represents the feature map of the low-light image in the (k + 1)-th interaction;

[0082] Step 4.3: Through interactive learning with the semantic perception module, finally, a feature map with semantic perception can be generated for the restoration of the low-light image.

[0083] Step 5: Design a loss function to guide the training process of the low-light image algorithm. In the low-light image enhancement method of the present invention, the training process is optimized by the dual constraints of content reconstruction loss and semantic reconstruction loss.

[0084] Step 5 is specifically implemented according to the following steps:

[0085] The training includes two parts: a semantic learner and image restoration. Therefore, it includes two parts: an image restoration reconstruction loss and a semantic reconstruction loss. The image restoration reconstruction loss includes three parts: a Charbonnier loss, a perceptual loss, and a gradient loss.

[0086] The semantic reconstruction loss in Step 5 is specifically as follows:

[0087] The feature map S output by the semantic learner and the semantic prior output by the pre-trained model BiSeNet together constitute the semantic reconstruction loss:

[0088] L prior = |S - S gt | (4)

[0089] L prior represents the semantic reconstruction loss, S is the semantic prior predicted by the semantic learner, and S gt is the semantic supervision signal obtained from the BiSeNet module;

[0090] The image restoration reconstruction loss in Step 5 is specifically as follows:

[0091] Step 5.1, The Charbonnier loss is expressed as:

[0092]

[0093] L r represents the reconstruction loss, I gt represents the high-quality standard image, i.e., the ground truth, and I en represents the enhanced image. |||| represents the L2 norm, which is used to calculate the distance or difference between the two, and κ is a smoothing term, which is a constant;

[0094] Step 5.2, The perceptual loss compares the VGG feature distance between the ground truth I gt and the restored image I en through the L1 loss:

[0095] L vgg = ||Φ(I gt ) - Φ(I en )|| (6)

[0096] Among them, L vgg represents the perceptual loss, and Φ() represents the operation of extracting features from the VGG network;

[0097] Step 5.3. The gradient loss represents information such as the details and textures of an image through edges and gradients. Therefore, in the present invention, the L1 norm loss function is adopted and constrained by the eigenvalue on the RGB three channels with the ground truth to ensure that the enhanced low-light image is consistent with the ground truth in details:

[0098]

[0099] Among them, L grad represents the gradient loss, represents the gradient of the high-quality standard image I gt on the c-th channel. represents the gradient of the enhanced image I en on the c-th channel, where H and W are the height and width of the image respectively, and c ∈ {1, 2, 3} corresponds to the RGB channels;

[0100] The total loss function in Step 5 is:

[0101] L = L prior + αL r + βL vgg + γL grad (8)

[0102] Among them, α, β, γ are hyperparameters.

[0103] Example 1

[0104] The low-light image enhancement method based on learnable semantic prior of the present invention combines Figure 1 and is specifically implemented according to the following steps:

[0105] Step 1. The input is divided into two branches, both of which take the low-light visible light image as the input. One branch is used for the restoration and reconstruction of the image content, and the other branch is used for predicting the semantic prior of the low-light image;

[0106] Step 2. In the first branch, three downsampling convolutional modules are adopted for feature encoding to enhance the input low-light visible light image to restore the brightness and details of the image;

[0107] Step 3. In the second branch, the low-light visible light image is input into the semantic learner network, and the semantic learner is supervised and trained by the high-quality visible light standard image to learn and predict the semantic prior information of the image;

[0108] Step 4: At the decoder stage, using the semantic prior obtained from the semantic learner, three upsampling convolutional modules are adopted and tightly coupled through the semantic perception module. Through multi-stage cascading, combining progressive upsampling and semantic perception guidance, the network can gradually restore the details of the image and finally generate the enhanced visible light image;

[0109] Step 5: Design a loss function to guide the training process of the low-light image algorithm. In the low-light image enhancement method of the present invention, the training process is optimized by the dual constraints of content reconstruction loss and semantic reconstruction loss.

[0110] Embodiment 2

[0111] The low-light image enhancement method based on learnable semantic prior of the present invention combines Figure 1 , and is specifically implemented according to the following steps:

[0112] Step 1: The input is divided into two branches, both branches take the low-light visible light image as the input. One branch is used for the restoration and reconstruction of the image content, and the other branch is used for predicting the semantic prior of the low-light image;

[0113] Step 1 is specifically implemented according to the following steps:

[0114] The branch for predicting the semantic prior of the low-light image, abbreviated as the semantic learner, is composed of three downsampling layers and five residual blocks. The semantic learner takes the low-light image I low as the input and outputs a feature map of size to learn the prior knowledge of the low-light image. In the training stage, the high-quality standard image GT is input into the pre-trained BiSeNet, and after adjusting the number of channels through 1x1 convolution, a feature map of size is obtained as the supervision signal of the semantic learner. The feature map S output by the semantic learner and the semantic prior output by the pre-trained model jointly constitute the semantic reconstruction loss. In the test stage, the semantic prior of the low-light image is output through this branch.

[0115] Step 2: In the first branch, three downsampling convolutional modules are adopted for feature encoding to enhance the input low-light visible light image to restore the brightness and details of the image;

[0116] Step 3: In the second branch, the low-light visible light image is input into the semantic learner network, and the semantic learner is supervised and trained by the high-quality visible light standard image to learn and predict the semantic prior information of the image;

[0117] Step 4: In the decoder stage, using the semantic prior obtained from the semantic learner, three upsampling convolutional modules are adopted and tightly coupled through the semantic perception module. Through multi-stage cascading, combined with gradual upsampling and semantic perception guidance, the network can gradually recover the details of the image and finally generate the enhanced visible light image;

[0118] Step 5: Design a loss function to guide the training process of the low-light image algorithm. In the low-light image enhancement method of the present invention, the training process is optimized through the dual constraints of content reconstruction loss and semantic reconstruction loss.

[0119] Embodiment 3

[0120] The low-illumination image enhancement method based on learnable semantic prior of the present invention combines Figure 1 , and is specifically implemented according to the following steps:

[0121] Step 1: The input is divided into two branches, both branches take the low-light visible light image as the input. One branch is used for the recovery and reconstruction of the image content, and the other branch is used for predicting the semantic prior of the low-light image;

[0122] Step 1 is specifically implemented according to the following steps:

[0123] The branch for predicting the semantic prior of the low-light image, abbreviated as the semantic learner, is composed of three downsampling layers and five residual blocks. The semantic learner takes the low-light image I low as the input and outputs a feature map with a size of to learn the prior knowledge of the low-light image. In the training stage, the high-quality standard image GT is input into the pre-trained BiSeNet, and after adjusting the number of channels through 1x1 convolution, a feature map with a size of is obtained as the supervision signal of the semantic learner. The feature map S output by the semantic learner and the semantic prior output by the pre-trained model together constitute the semantic reconstruction loss. In the test stage, the semantic prior of the low-light image is output through this branch.

[0124] Step 2: In the first branch, three downsampling convolutional modules are adopted for feature encoding to enhance the input low-light visible light image and restore the brightness and details of the image;

[0125] Step 2 is specifically implemented according to the following steps:

[0126] Step 2.1: Receive the low-light image as the input. The height of the image is H and the width is W. Denote the low-light image as I low ;

[0127] Step 2.2: To extract the basic features of the image, three downsampling convolutional modules are used for feature encoding. Each convolutional module extracts features from the image through convolution operations and gradually reduces the image size through the stride. This process is expressed as:

[0128] F I = E I (I low ) (1)

[0129] where H and W respectively represent the height and width of the image, I low represents the encoder of the low-light image, E I represents the encoder of the low-light image, F I represents the feature map in the encoding stage;

[0130] Step 2.3: Through the encoding process, the obtained feature map F I is used as the input for the subsequent image enhancement or semantic task branch, providing basic feature support for subsequent processing.

[0131] Step 3: In the second branch, the low-light visible light image is input into the semantic learner network, and the semantic learner is supervised and trained with high-quality visible light standard images to learn and predict the semantic prior information of the image;

[0132] Step 4: In the decoder stage, using the semantic prior obtained from the semantic learner, three upsampling convolutional modules are adopted and tightly coupled through the semantic perception module. Through multi-stage cascading, combined with gradual upsampling and semantic perception guidance, the network can gradually restore the details of the image and finally generate the enhanced visible light image;

[0133] Step 5: Design a loss function to guide the training process of the low-light image algorithm. In the low-light image enhancement method of the present invention, the training process is optimized through the dual constraints of content reconstruction loss and semantic reconstruction loss.

[0134] Example 4

[0135] The low-illumination image enhancement method based on learnable semantic prior of the present invention combines Figure 1 , and is specifically implemented according to the following steps:

[0136] Step 1: The input is divided into two branches, both of which take the low-light visible light image as the input. One branch is used for the restoration and reconstruction of the image content, and the other branch is used for predicting the semantic prior of the low-light image;

[0137] Step 1 is specifically implemented according to the following steps:

[0138] The semantic prior branch for predicting low-light images, abbreviated as the semantic learner, consists of three downsampling layers and five residual blocks. The semantic learner takes the low-light image I low as input and outputs a feature map of size to learn the prior knowledge of low-light images. In the training stage, the high-quality standard image GT is input into the pre-trained BiSeNet. After adjusting the number of channels through 1x1 convolution, a feature map of size is obtained as the supervision signal for the semantic learner. The feature map S output by the semantic learner and the semantic prior output by the pre-trained model together constitute the semantic reconstruction loss. In the test stage, the semantic prior of the low-light image is output through this branch.

[0139] Step 2: In the first branch, three downsampling convolutional modules are used for feature encoding to enhance the input low-light visible light image to restore the brightness and details of the image.

[0140] Step 3: In the second branch, the low-light visible light image is input into the semantic learner network, and the semantic learner is supervised and trained with the high-quality visible light standard image to learn and predict the semantic prior information of the image.

[0141] Step 3 is specifically implemented according to the following steps:

[0142] Step 3.1: To ensure that the semantic learner can effectively learn the semantic information in low-light images, the semantic segmentation model BiSeNet is used as the semantic knowledge base. BiSeNet is a lightweight and efficient real-time semantic segmentation network with strong global context understanding ability, which can identify the semantic information of multiple different objects in the image.

[0143] Step 3.2: The structure of the semantic learner takes the low-light image as input and consists of three downsampling layers and five residual blocks. Specifically, first, the low-light image is input and passes through three downsampling layers in sequence. Each downsampling layer uses a 4×4 convolution operation with a stride of 2 and a padding of 1. The number of channels gradually expands from 3 to 128, and the ReLU activation function is used to extract preliminary features. Subsequently, the feature map is input into 5 residual blocks, and the number of input and output channels of each residual block is 128 to enhance the feature expression ability and retain key information. Finally, the number of channels is compressed to the specified output channel number, defaulting to 64, through a 1×1 convolution layer to generate the semantic prior features. The purpose of this network is to predict the semantic prior information from the low-light image and provide semantic support for subsequent low-image enhancement.

[0144] Step 3.3: The high-quality standard image GT in the LOLv2 dataset is input into the pre-trained BiSeNet model, and the number of channels is adjusted through 1x1 convolution operation to obtain a size of The feature map, as the supervision signal of the semantic learner, effectively provides semantic-level supervision information for the semantic learner in this process, helping it learn semantic priors;

[0145] Step 3.4: The feature map S output by the semantic learner and the semantic prior output by the pre-trained model jointly constitute the semantic reconstruction loss. By comparing the output features with the pre-trained semantic prior, this loss function ensures that the semantic learner can effectively learn the semantic information in the low-light image and perform reasonable image enhancement;

[0146] Step 3.5: Through the above steps, the semantic learner is trained under the guidance of the semantic reconstruction loss to optimize the network weights, and finally achieve the goal of low-light image enhancement and semantic information preservation.

[0147] Step 4: In the decoder stage, using the semantic prior obtained from the semantic learner, three upsampling convolutional modules are adopted and tightly coupled through the semantic perception module. Through multi-stage cascading, combining progressive upsampling and semantic perception guidance, the network can gradually restore the details of the image and finally generate the enhanced visible light image;

[0148] Step 5: Design a loss function to guide the training process of the low-light image algorithm. In the low-light image enhancement method of the present invention, the training process is optimized by the dual constraints of the content reconstruction loss and the semantic reconstruction loss.

[0149] Embodiment 5

[0150] The low-light image enhancement method based on learnable semantic priors of the present invention combines Figure 1 , and is specifically implemented according to the following steps:

[0151] Step 1: The input is divided into two branches, both branches take the low-light visible light image as the input. One branch is used for the restoration and reconstruction of the image content, and the other branch is used for predicting the semantic prior of the low-light image;

[0152] Step 1 is specifically implemented according to the following steps:

[0153] The branch for predicting the semantic prior of the low-light image, abbreviated as the semantic learner, is composed of three downsampling layers and five residual blocks. The semantic learner takes the low-light image I low as the input and outputs a feature map with a size of for learning the prior knowledge of the low-light image. In the training stage, the high-quality standard image GT is input into the pre-trained BiSeNet, and the number of channels is adjusted through 1x1 convolution to obtain a size of The feature map, as the supervision signal of the semantic learner, the feature map S output by the semantic learner and the semantic prior output by the pre-trained model jointly constitute the semantic reconstruction loss. In the test phase, the semantic prior of the low-light image is output through this branch.

[0154] Step 2: In the first branch, three downsampling convolutional modules are used for feature encoding to enhance the input low-light visible light image to restore the brightness and details of the image.

[0155] Step 2 is specifically implemented according to the following steps:

[0156] Step 2.1: Receive the low-light image as the input. The height of the image is H and the width is W. Denote the low-light image as I low ;

[0157] Step 2.2: To extract the basic features of the image, three downsampling convolutional modules are used for feature encoding. Each convolutional module extracts features from the image through convolutional operations and gradually reduces the image size through the stride. This process is expressed as:

[0158] F I = E I (I low ) (1)

[0159] Among them, H and W respectively represent the height and width of the image, I low represents the encoder of the low-light image, E I represents the encoder of the low-light image, F I represents the feature map in the encoding stage;

[0160] Step 2.3: Through the encoding process, the obtained feature map F I is used as the input for the subsequent image enhancement or semantic task branch, providing basic feature support for subsequent processing.

[0161] Step 3: In the second branch, the low-light visible light image is input into the semantic learner network, and the semantic learner is supervised and trained with high-quality visible light standard images to learn and predict the semantic prior information of the image;

[0162] Step 4: In the decoder stage, using the semantic prior obtained from the semantic learner, three upsampling convolutional modules are adopted and tightly coupled through the semantic perception module. Through multi-stage cascading, combined with step-by-step upsampling and semantic perception guidance, the network can gradually restore the details of the image and finally generate the enhanced visible light image;

[0163] Step 4 is specifically implemented according to the following steps:

[0164] Step 4.1: Through the 1×1 convolutional layer of the semantic learner in Step 3.2, map the semantic prior knowledge and the features of the low-light image to the same dimension, and then calculate the semantic-aware attention map using cross-modal similarity as follows:

[0165]

[0166] where Conv m (.), Conv n (.) are convolutional layers, C is the feature channel, and where represents the semantic-aware attention map, represents the feature map of the low-light image at the k-th interaction, represents the predicted semantic prior feature of the low-light image in the K-th interaction, I represents the first letter of the low-light image, k represents the number of interactions, taking values of 1, 2, 3, m, n represents the number of convolutional layers, and s represents the semantic prior feature;

[0167] Step 4.2: Perform pixel-wise weighted summation of the semantic-aware attention map and the feature map of the low-light image to generate the feature map for the next stage:

[0168]

[0169] where represents the feature map of the low-light image at the (k + 1)-th interaction;

[0170] Step 4.3: Through interactive learning with the semantic-aware module, finally, a semantic-aware feature map can be generated for the restoration of the low-light image.

[0171] Step 5: Design a loss function to guide the training process of the low-light image algorithm. In the low-light image enhancement method of the present invention, the training process is optimized by the dual constraints of content reconstruction loss and semantic reconstruction loss.

[0172] Step 5 is specifically implemented according to the following steps:

[0173] The training includes two parts: the semantic learner and image restoration, so it includes two parts: the image restoration reconstruction loss and the semantic reconstruction loss. The image restoration reconstruction loss includes three parts: Charbonnier loss, perceptual loss, and gradient loss.

[0174] Example 6

[0175] The low-illumination image enhancement method based on learnable semantic prior of the present invention, combined with Figure 1 、 Figure 2 is specifically implemented according to the following steps:

[0176] Step 1: The input is divided into two branches, both of which take the low-light visible light image as the input. One branch is used for the restoration and reconstruction of the image content, and the other branch is used for predicting the semantic prior of the low-light image.

[0177] Step 1 is specifically implemented according to the following steps:

[0178] The branch for predicting the semantic prior of the low-light image, abbreviated as the semantic learner, is composed of three downsampling layers and five residual blocks. The semantic learner takes the low-light image I low as the input and outputs a feature map with a size of to learn the prior knowledge of the low-light image. In the training stage, the high-quality standard image GT is input into the pre-trained BiSeNet, and after adjusting the number of channels through 1x1 convolution, a feature map with a size of is obtained as the supervision signal of the semantic learner. The feature map S output by the semantic learner and the semantic prior output by the pre-trained model together constitute the semantic reconstruction loss. In the testing stage, the semantic prior of the low-light image is output through this branch.

[0179] Step 2: In the first branch, three downsampling convolutional modules are used for feature encoding to enhance the input low-light visible light image to restore the brightness and details of the image.

[0180] Step 3: In the second branch, the low-light visible light image is input into the semantic learner network, and the semantic learner is supervised and trained by the high-quality visible light standard image to learn and predict the semantic prior information of the image.

[0181] Step 4: In the decoder stage, using the semantic prior obtained from the semantic learner, three upsampling convolutional modules are adopted and tightly coupled through the semantic perception module. Through multi-stage cascading, combined with step-by-step upsampling and semantic perception guidance, the network can gradually restore the details of the image and finally generate the enhanced visible light image.

[0182] Step 5: Design a loss function to guide the training process of the low-light image algorithm. In the low-light image enhancement method of the present invention, the training process is optimized by the dual constraints of the content reconstruction loss and the semantic reconstruction loss.

[0183] Step 5 is specifically implemented according to the following steps:

[0184] The training includes two parts: the semantic learner and image restoration, so it includes two parts: the image restoration and reconstruction loss and the semantic reconstruction loss. The image restoration and reconstruction loss includes three parts: the Charbonnier loss, the perceptual loss, and the gradient loss.

[0185] The semantic reconstruction loss in Step 5 is specifically as follows:

[0186] The feature map S output by the semantic learner and the semantic prior output by the pre-trained model BiSeNet together constitute the semantic reconstruction loss:

[0187] L prior = |S - S gt | (4)

[0188] L prior represents the semantic reconstruction loss, S is the semantic prior predicted by the semantic learner, and S gt is the semantic supervision signal obtained from the BiSeNet module;

[0189] The image restoration and reconstruction loss in step 5 is specifically as follows:

[0190] Step 5.1, The Charbonnier loss is expressed as:

[0191]

[0192] L r represents the reconstruction loss, I gt represents the high-quality standard image, i.e., the ground truth, and I en represents the enhanced image, |||| represents the L2 norm, which is used to calculate the distance or difference between the two, and κ is a smoothing term, which is a constant;

[0193] Step 5.2, The perceptual loss compares the VGG feature distance between the ground truth I gt and the restored image I en :

[0194] L vgg = ||Φ(I gt ) - Φ(I en )|| (6)

[0195] where, L vgg represents the perceptual loss, and Φ() represents the operation of extracting features from the VGG network;

[0196] Step 5.3, The gradient loss represents information such as the details and texture of the image through edges and gradients. Therefore, in the present invention, the L1 norm loss function is adopted and constrained by the eigenvalue of the ground truth on the three RGB channels to ensure that the enhanced low-light image is consistent with the ground truth in details:

[0197]

[0198] where, L grad represents the gradient loss, represents the gradient of the high-quality standard image I gt on the c-th channel. Denote the enhanced image as I en Gradient in the c-th channel, where H and W are the height and width of the image respectively, and c ∈ {1, 2, 3} corresponding to the RGB channels;

[0199] The total loss function in step 5 is:

[0200] L = L prior + αL r + βL vgg + γL grad (8)

[0201] where α, β, γ are hyperparameters.

Claims

1. A low-light image enhancement method based on learnable semantic priors, characterized in that: Follow the steps below to implement it: Step 1: The input is divided into two branches, both of which take low-light visible light images as input. One branch is used to restore and reconstruct the image content, and the other branch is used to predict the semantic prior of the low-light image. Step 2: In the first branch, the input low-light visible light image is enhanced; Step 3: In the second branch, learn and predict the semantic prior information of the image; Step 4: Using the semantic priors obtained from the semantic learner, gradually restore the details of the image and finally generate an enhanced visible light image; Step 5: Design a loss function to guide the training process of the low-light image algorithm.

2. The low-light image enhancement method based on learnable semantic prior according to claim 1, characterized in that: The step 1 is specifically implemented according to the following steps: The semantic prior branch for predicting low-light images is referred to as the semantic learner. The semantic learner consists of three downsampling layers and five residual blocks. The semantic learner is based on the low-light image I low As input, the output size is The feature map is used to learn the prior knowledge of low-light images. In the training stage, the high-quality standard image GT is input into the pre-trained BiSeNet, and the number of channels is adjusted through 1x1 convolution to obtain a size of The feature map S output by the semantic learner and the semantic prior output by the pre-trained model together constitute the semantic reconstruction loss. In the test phase, the semantic prior of the low-light image is output through this branch.

3. The low-light image enhancement method based on learnable semantic prior according to claim 2, characterized in that: The step 2 is specifically implemented according to the following steps: Step 2.1: Receive a low-light image as input. The image height is H and the width is W. The low-light image is recorded as I. low ; Step 2.2: Three down-sampling convolution modules are used for feature encoding. Each convolution module extracts features from the image through convolution operation and gradually reduces the image size through stride. This process is expressed as: F I =E I (I low ) (1) in, H and W represent the height and width of the image respectively, I low The encoder representing the low-light image, E I The encoder representing the low-light image, F I Feature map representing the encoding stage; Step 2.3: Through the encoding process, the feature map F is obtained I As the input of subsequent image enhancement or semantic task branches, it provides basic feature support for subsequent processing.

4. The low-light image enhancement method based on learnable semantic prior according to claim 3, characterized in that: The step 3 is specifically implemented according to the following steps: Step 3.1: To ensure that the semantic learner can effectively learn the semantic information in low-light images, the semantic segmentation model BiSeNet is used as the semantic knowledge base; Step 3.2: The structure of the semantic learner takes the low-light image as input and consists of three downsampling layers and five residual blocks. Specifically, the input low-light image first passes through three downsampling layers in sequence. Each downsampling layer uses a 4×4 convolution operation with a step size of 2 and a padding of 1. The number of channels is gradually expanded from 3 to 128. At the same time, the ReLU activation function is used to extract preliminary features. Subsequently, the feature map is input into 5 residual blocks. The number of input and output channels of each residual block is 128. Finally, a 1×1 convolution layer is used to compress the number of channels to the specified number of output channels, which is 64 by default, to generate semantic prior features. The purpose of this network is to predict semantic prior information from low-light images and provide semantic support for subsequent low-light image enhancement. Step 3.3: Input the high-quality standard image GT in the LOLv2 dataset into the pre-trained BiSeNet model, and adjust the number of channels through a 1x1 convolution operation to obtain a size of The feature map of , as the supervision signal of the semantic learner; Step 3.4: The feature map S output by the semantic learner and the semantic prior output by the pre-trained model together constitute the semantic reconstruction loss. This loss function ensures that the semantic learner can effectively learn the semantic information in the low-light image and perform reasonable image enhancement by comparing the output features with the pre-trained semantic prior. Step 3.5: Through the above steps, the semantic learner is trained under the guidance of semantic reconstruction loss to optimize the network weights, and finally achieve the goal of low-light image enhancement and semantic information preservation.

5. The low-light image enhancement method based on learnable semantic prior according to claim 4, characterized in that: The step 4 is specifically implemented according to the following steps: Step 4.1: Map the semantic prior knowledge and the features of the low-light image to the same dimension through the 1×1 convolutional layer of the semantic learner in step 3.2, and then use the cross-modal similarity to calculate the semantically aware attention map as follows: Among them, Conv m (.), Conv n (.) is the convolution layer, C is the feature channel, where represents the semantically aware attention map, The feature map representing the low-light image of the kth interaction, represents the semantic prior features of low-light image prediction in the Kth interaction, I represents the first letter of the low-light image, k represents the number of interactions, and its value is 1, 2, 3, m, n represents the number of convolutional layers, and s represents the semantic prior features; Step 4.2: Perform pixel-by-pixel weighted summation of the semantically-aware attention map and the feature map of the low-light image to generate the feature map of the next stage: in, The feature map representing the low-light image of the k+1th interaction; Step 4.3: Through interactive learning with the semantic perception module, a semantically aware feature map can be generated. Used for low-light image restoration.

6. The low-light image enhancement method based on learnable semantic prior according to claim 5, characterized in that: The step 5 is specifically implemented according to the following steps: The training includes two parts: semantic learner and image restoration, so it includes two parts: image restoration and reconstruction loss and semantic reconstruction loss. The image restoration and reconstruction loss includes three parts: Charbonnier loss, perceptual loss, and gradient loss.

7. The low-light image enhancement method based on learnable semantic prior according to claim 6, characterized in that: The semantic reconstruction loss in step 5 is as follows: The feature map S output by the semantic learner and the semantic prior output by the pre-trained model BiSeNet together constitute the semantic reconstruction loss: L prior =|S-S gt | (4) L prior represents the semantic reconstruction loss, S is the semantic prior predicted by the semantic learner, and S gt is the semantic supervision signal obtained from the BiSeNet module.

8. The low-light image enhancement method based on learnable semantic prior according to claim 7, characterized in that: The image restoration and reconstruction loss in step 5 is specifically as follows: Step 5.1, Charbonnier loss is expressed as: L r represents the reconstruction loss, I gt represents the high-quality standard image, i.e., the ground truth value, I en represents the enhanced image, |||| represents the L2 norm, which is used to calculate the distance or difference between the two, and κ is a smoothing term, which is a constant; Step 5.2: Perceptual loss compares the ground truth value I through L1 loss gt and the restored image I en The VGG feature distance between: L vgg =||Φ(I gt )-Φ(I en )|| (6) Among them, L vgg represents the perceptual loss, Φ() represents the operation of extracting features from the VGG network; Step 5.3: The gradient loss uses the L1 norm loss function and is constrained by the feature values ​​of the three RGB channels of the ground truth to ensure that the enhanced low-light image is consistent with the ground truth in details: Among them, L grad represents the gradient loss, Represents a high quality standard image I gt The gradient in channel c, Represents the enhanced image I en The gradient at the cth channel, H and W are the height and width of the image, respectively, and c∈{1,2,3} corresponds to the RGB channels.

9. The low-light image enhancement method based on learnable semantic prior according to claim 8, characterized in that: The total loss function in step 5 is: L=L prior +αL r +βL vgg +γL grad (8) Among them, α, β, and γ are hyperparameters.

Citation Information

Cited By

  • Zero-sample low-illumination image enhancement method based on brightness and semantic prior guidance

    CN122367846A