Self-adaptive weak light image enhancement method based on customized prompt learning

By introducing a multi-level prompt mechanism, the details retention and semantic consistency of image enhancement in complex low-light conditions are solved, and high-quality image recovery effect is achieved, suitable for video surveillance, autonomous driving and medical image processing.

CN120259149APending Publication Date: 2025-07-04HANGZHOU DIANZI UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510218373.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art In complex low-light conditions, low-light image enhancement methods have a trade-off between detail retention and semantic consistency, and the adaptability and generalization are insufficient when multiple degradation factors are superimposed.

Method used

Adaptive low-light image enhancement method based on customized prompt learning is adopted, and the image enhancement process is guided by designing a multi-level prompt mechanism, including semantic, non-priori and texture prompt generation modules, combining U-Net structure and multi-head self-attention mechanism.

Benefits of technology

It significantly improves the brightness, detail performance and color restoration capabilities of low-light images, improves clarity and overall visual effects, and is suitable for scenes such as video surveillance, autonomous driving and medical image processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259149A_ABST
    Figure CN120259149A_ABST
Patent Text Reader

Abstract

The invention discloses a self-adaptive weak light image enhancement method based on customized prompt learning. The method comprises the following steps: firstly, constructing a self-adaptive weak light image enhancement model based on customized prompt learning, and designing a dual-path processing mechanism by adopting a U-Net architecture; performing feature extraction on the input weak light image through an encoder of the encoding and decoding path to obtain multi-scale and multi-level image feature information; and further processing the image feature information extracted by the encoder through the decoder, in the processing process of the decoder, sequentially combining a plurality of different prompt features output by the prompt generation path, assisting the recovery of the image information through prompts, and finally outputting an enhanced image. According to the method, the brightness, detail representation and color rendition capability of the low-light image are remarkably improved. Experiments on a plurality of public data sets show that the method is superior to an existing method in the aspects of definition improvement, detail retention and overall visual effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision and deep learning, and particularly relates to a method for adaptively enhancing low-light images based on prompt learning. This method guides the model to perform adaptive processing on low-light images through a multi-level prompt mechanism, aiming to improve the restoration quality of images under complex lighting conditions. Background Art

[0002] Low-light image enhancement is one of the important research topics in the field of computer vision, aiming to improve the visual quality of images in low-light environments and is widely applied in practical scenarios such as video surveillance, autonomous driving, and medical imaging. Under low-light conditions, images are usually affected by problems such as insufficient contrast, noise interference, and color distortion, which not only reduce the human eye's recognition of image content but also pose great challenges to subsequent visual analysis tasks (such as object detection and object recognition).

[0003] With the rapid development of deep learning technology, low-light image enhancement has gradually evolved from traditional methods to data-driven intelligent technologies. Traditional methods include histogram equalization, Retinex theory, etc., which are effective for simple scenes but have limited effects under complex conditions. In recent years, end-to-end learning methods based on deep neural networks have become mainstream, especially convolutional neural networks (CNNs) and generative adversarial networks (GANs) have achieved remarkable results in detail restoration, noise suppression, and color recovery. At the same time, technologies combining generative models such as variational autoencoder layers (VAEs) have further improved the enhancement effect.

[0004] Although existing methods have made progress in many aspects, in the face of complex low-light scenes (such as extremely uneven lighting, locally over-bright or over-dark), there is still a trade-off between detail preservation and semantic consistency in the existing technology. In particular, in the case of the superposition of multiple degradation factors, the adaptability and generalization of existing models need to be improved. Therefore, introducing innovative technologies such as prompt learning and guiding the enhancement process by combining multi-level prompt mechanisms (such as texture detail prompts, semantic prompts, implicit knowledge prompts) provides a new idea for solving the complex degradation problems of low-light images. Summary of the Invention

[0005] Aiming at the deficiencies in the existing technology, the present invention provides an adaptive low-light image enhancement method based on customized prompt learning. By designing a multi-level prompt mechanism, including a semantic-based prompt generation module (SemanticPrompt Block, SPB), a prior-free prompt generation module (Prior-Free Prompt Block, FPB), and a texture-based prompt generation module (Texture Prompt Block, TPB), the adaptive enhancement of low-light images is realized. This method adopts a U-Net structure and combines a multi-head self-attention mechanism with a cross-modal prompt fusion strategy to effectively integrate prompt learning into the image enhancement process.

[0006] Specifically, the present invention introduces a pluggable prompt mechanism in the enhancement model in stages, including the following three modules: a semantic-based prompt generation module, which extracts the semantic features of low-light images through a pre-trained CLIP model and generates semantic prompts to enhance the high-level understanding ability of the model; a prior-free prompt generation module, which dynamically generates implicit semantic prompts to achieve adaptive enhancement for diverse degradation types in low-light images; a texture-based prompt generation module, which introduces high-frequency texture detail prompts to guide detail restoration and suppress noise. The above prompt information is combined layer by layer with the features of the encoding layer through a feature fusion module to gradually improve the image quality and finally output a high-quality normal-light image.

[0007] The adaptive low-light image enhancement method based on customized prompt learning includes the following steps:

[0008] Step (1) Construct an adaptive low-light image enhancement model based on customized prompt learning;

[0009] The adaptive low-light image enhancement model adopts a U-Net architecture and designs a dual-path processing mechanism, including an encoding-decoding path composed of an encoder and a decoder and a prompt generation path composed of a semantic-based prompt generation module (SemanticPromptBlock, SPB), a prior-free prompt generation module (Prior-Free Prompt Block, FPB), and a texture-based prompt generation module (Texture PromptBlock, TPB).

[0010] Step (2) First, the encoder of the encoding-decoding path extracts the features of the input low-light image to obtain multi-scale and multi-level image feature information;

[0011] Step (3) The decoder further processes the image feature information extracted by the encoder, uses the multi-head attention mechanism for information fusion, and gradually restores the size of the image in combination with the upsampling operation.

[0012] Step (4) During the decoder processing, multiple different prompt features are sequentially combined with the output of the prompt generation path, and the recovery of the image information is assisted by this prompt, and finally the enhanced image is output.

[0013] Step (5) Construct the overall loss function of the adaptive low-light image enhancement model and train the model.

[0014] Furthermore, the encoders and decoders of different layers are all composed of multiple Transformer blocks.

[0015] Furthermore, in step (2), first, the input low-light image is subjected to channel number expansion through the input layer, and then the processed image is passed into the encoder for subsequent processing;

[0016] The input low-light image with a size of img ∈ R H×W×3 first passes through a 3×3 convolution (i.e., the input layer) to obtain the underlying feature F0 ∈ R H×W×C . Where C represents the number of channels, H is the height of the image, and W is the width of the image;

[0017] Furthermore, the encoder is divided into three layers, namely encoding layer 1, encoding layer 2, and encoding layer 3, which are successively composed of 1, 2, and 4 Transform blocks. The image processed by the input layer is subjected to feature extraction through three consecutive encoding layers to obtain features F1 ∈ R H×W×C , where between each encoding layer, the features are tiled through downsampling operations to obtain features

[0018] Furthermore, the decoder includes decoding layer 1, decoding layer 2, and decoding layer 3, which are successively composed of 1, 2, and 4 Transform blocks.

[0019] Furthermore, during the decoder processing, the prompt features output by the prompt generation path are combined, and the specific operations are as follows:

[0020] The prompt mechanism is introduced in stages in the decoder of the adaptive low-light image enhancement model: the semantics-based prompt generation module uses the pre-trained CLIP model to generate semantic symbols, which are then fused with the original features to generate semantic prompts to enhance the high-level understanding ability of the model; the no-prior-prompt generation module adaptively enhances the model by dynamically generating implicit semantic prompts for diverse degradation types in low-light images; the texture-based prompt generation module effectively guides detail recovery and noise suppression by introducing high-frequency texture detail prompts. The obtained prompt features are combined layer by layer with the features generated by the encoder through the feature fusion module, gradually improving the image quality, and finally outputting a high-quality normal illumination image.

[0021] Furthermore, the semantics-based prompt generation module uses a pre-trained CLIP model to generate semantic tokens, which are then fused with the original features to generate semantic prompts to enhance the model's high-level understanding ability. The specific operations are as follows:

[0022] In the semantics-based prompt generation module, first, a randomly generated learnable vector is passed as input to the frozen CLIP text encoding layer. The obtained semantic token Emb and the input feature F4 are fused through a dual-gated self-attention mechanism to generate a semantic prompt S. Finally, the semantic prompt S is concatenated with the input feature F4, and then through a 3×3 convolution operation, the semantic prompt feature F4', which is the high-level prompt information, is obtained. The corresponding structured encoding formula is as follows:

[0023] Emb = CLIP text (Token) (Formula 1)

[0024] v = [Reshape(F4); FC(Emb)] (Formula 2)

[0025] S = Reshape(F4) + λ × tanh(θ) × Attn(v) (Formula 3)

[0026] F4' = Conv 1×1 ([F4; Reshape(S)]) (Formula 4)

[0027] Among them, Token represents a randomly generated learnable vector with the same shape as the CLIP text input, Emb represents the semantic token, and CLIP text represents the CLIP text encoding layer. is the output feature of the last encoding layer. Reshape represents the reshaping operation, FC represents the linear transformation of the channel dimension, and Attn represents the multi-head attention mechanism. represents the semantic prompt feature.

[0028] Finally, after the semantic prompt feature F4' is upsampled, it is concatenated with the feature F3 output by the encoding layer 3, and after the channel is adjusted through a convolution operation, it is input to the decoding layer 3 to obtain the output feature

[0029] Furthermore, the prior-free prompt generation module adaptively enhances for diverse degradation types in low-light images by dynamically generating implicit semantic prompts. The specific operations are as follows:

[0030] In the prior-free prompt generation module, first, global average pooling is performed on the output feature F5 of the decoding layer 3, and then through a 1×1 convolution operation and an operation to generate weights, the weight w is obtained i. Immediately afterwards, multiply the weight w i by the learnable vector P c . Then, through a 3×3 convolution operation, obtain the prior-free implicit semantic cue G. Concatenate the generated prior-free implicit semantic cue G with the input feature F5, and then through a 3×3 convolution operation, obtain the prior-free semantic cue feature F5'. This process is described by the following equations:

[0031]

[0032] w i = Softmax(Conv 1×1 (GAP(F5))) (Equation 7)

[0033] F5' = Conv 1×1 ([F5; G↑]) (Equation 8)

[0034] where P c ∈R N×16×16×4C represents the randomly initialized learnable vector cue component, N is the number of learning components, Conv 3×3 represents the 3×3 convolution operation, is the output feature from the decoding layer 3. GAP represents the global average pooling operation, w i represents the weight calculated through the convolution operation, Softmax is used to generate the weight, G represents the generated prior-free implicit semantic cue, is the prior-free semantic cue feature.

[0035] After upsampling F5' and concatenating it with the encoding layer 2 feature F2, and adjusting the channels through the convolution operation, input it to the decoding layer 2 to obtain the output feature

[0036] Furthermore, the texture-based cue generation module effectively guides detail restoration and noise suppression by introducing high-frequency texture detail cues. The specific operations are as follows:

[0037] The texture-based cue generation module processes the input low-light image through the texture encoder to generate the texture map T, that is, the high-frequency texture detail cue. The texture encoder consists of six convolutional layers, and the first five convolutional layers are followed by ReLU activation layers. Fuse the texture map T generated by the texture encoder with the output feature of the decoding layer 2 to generate the texture-based cue feature This process is described by the following formula:

[0038] F6' = Conv 1×1 ([F6; F6×scale(T)+shift(T)]) (Equation 9)

[0039] Among them, T represents the generated texture map, and scale and shift respectively represent 3×3 convolution operations for controlling the effective range and offset of features.

[0040] After upsampling F6' and concatenating it with the feature F1 output by the encoding layer 1, and adjusting the channels through a convolution operation, it is output to the decoding layer 1 to obtain the output feature F7 ∈ R H×W×C ;

[0041] Furthermore, after passing through an output layer composed of an image reconstruction self-attention mechanism and a 3×3 convolution operation, the output feature F7 of the decoding layer 1 obtains the enhanced image P.

[0042] An adaptive low-light image enhancement method based on customized prompt learning, whose overall loss function L All includes three parts: enhancement loss L E , texture encoder training loss L T and regularization loss L R . Its overall formula is as follows:

[0043] L All = λ1 × L E + λ2 × L T + λ3 × L R (Formula 10)

[0044] Among them, λ1, λ2, and λ3 are weights used to balance the contributions of each loss term to the overall loss function.

[0045] The enhancement loss is composed of the mean absolute error (MAE) loss and the structural similarity (SSIM) loss, and its formula is:

[0046] L E = MAE(P, I) + 1 - SSIM(P, I) (Formula 11)

[0047] Among them, P represents the enhanced image, and I represents the target normal illumination image.

[0048] The texture encoder training loss L T includes the mean square error (MSE) loss and the total variation (TV) loss, aiming to promote the training of the texture encoder for generating detailed texture maps. The weights of these losses are controlled by α and β to achieve balance. The formula for L T is as follows:

[0049] L T = α × MSE(T, S) + β × TV(T) (Formula 12)

[0050] Among them, T represents the generated texture map, and S is the high-frequency texture map obtained by applying the Sobel filter to the normal illumination image. To ensure the stable generation of T, the TV loss is introduced to enhance smoothness.

[0051] Regularization loss L R is used to prevent overfitting, and its formula is:

[0052]

[0053] where p i represents the value of the model parameter, m represents the total number of model parameters, and μ is a constant, which is default set to 1×10 -6 .

[0054] The beneficial effects of the present invention are as follows:

[0055] The present invention significantly improves the brightness, detail performance, and color restoration ability of low-light images. Experiments on multiple public datasets show that this method is superior to existing methods in terms of clarity improvement, detail retention, and overall visual effect, and has broad application prospects. It can be applied to scenarios that require image enhancement such as video surveillance, autonomous driving, and medical image processing, providing an innovative solution to the image processing problem under complex lighting conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 is the flow chart of the present invention;

[0057] Figure 2 is the overall framework schematic diagram of the present invention;

[0058] Figure 3 is the detailed process of three generation prompts of the present invention;

[0059] Figure 4 is the effect display of the enhanced image of the present invention;

[0060] Figure 5 is the effect display of the self-built test set;

[0061] Figure 6 is the effect display of each method for testing on the unpaired dataset. DETAILED DESCRIPTION OF THE INVENTION

[0062] The present invention will be further described in detail below with reference to the drawings.

[0063] As Figure 1As shown in the figure, the present invention provides a method for adaptive enhancement of low-light images based on prompt learning, aiming to enhance low-light images and assist downstream tasks. The present invention realizes an efficient conversion from low-light images to enhanced images by designing a dual-path processing mechanism. Specifically, the low-light image first expands the number of channels through the input layer, and then the encoder extracts multi-scale and multi-level feature information. On the one hand, the extracted feature information is directly transmitted to the decoder for processing, and on the other hand, it is used to generate prompt features. The prompt features and the image feature information are jointly input into the decoder to further optimize the effect of image enhancement. During the feature extraction process, the encoder converts the input low-light image into multi-dimensional features, which not only provide basic support for subsequent image enhancement, but also are input into the prompt generation module to generate prompt features suitable for the current image enhancement requirements. The prompt generation module generates targeted prompt features according to the features extracted by the encoder. Subsequently, the feature information and the generated prompt features are sent to the decoder, and the decoder generates an optimized feature representation by synergistically fusing the image features and the prompt features, thereby improving the enhancement effect of low-light images. Finally, the decoder combines the feature information of the encoder and the prompt features generated by the prompt generation path to output a high-quality enhanced image. Through the above process, the present invention realizes a highly collaborative optimization of image feature extraction, enhancement prompt generation, and decoding output, significantly improving the enhancement effect of low-light images and the adaptability of the model.

[0064] The specific implementation steps are as Figure 2 shown:

[0065] Step (1): Input the low-light image into the input layer to expand the number of channels, obtaining the expanded image. This image is then transmitted to Encoding Layer 1 to extract feature information, obtaining feature F1 ∈ R H×W×C . Subsequently, after downsampling F1, the feature

[0066] Step (2): Transmit feature F1 to Decoding Layer 3 for subsequent use; feature F1' continues to be transmitted to the next Encoding Layer 2, obtaining feature Subsequently, after downsampling F2, the feature

[0067] Step (3): Transmit feature F2 to Decoding Layer 2 for subsequent use; feature F2' then continues to be transmitted to Encoding Layer 3, obtaining feature Subsequently, after downsampling feature F3, the feature

[0068] Step (4): Transmit F3 to Decoding Layer 3 for subsequent use; feature F4 is transmitted to the semantic-based prompt generation module;

[0069] Step (5), the structure of the semantic-based prompt generation module is as follows Figure 3 As shown, this module first takes a randomly generated learnable vector as a prompt and inputs it into a large language model with frozen parameters to obtain semantic symbols. Subsequently, the semantic symbols are passed into the attention mechanism and fused with the feature F4. The fused result is then concatenated with F4. Finally, semantic prompt features are obtained through a convolution operation

[0070] Step (6), after the semantic prompt feature F4' is upsampled, it is passed into the decoding layer 3 and fused with the feature F3 to obtain a feature Then, the feature F5 is input into the prior-free prompt generation module;

[0071] Step (7), the structure of the prior-free prompt generation module is as follows Figure 3 As shown, first the feature F5 is processed through a pooling layer, then normalized, and multiplied by a randomly initialized learnable vector. Next, after a convolution operation, it is concatenated with the feature F5, and finally, prior-free semantic prompt features are obtained through a convolution operation

[0072] Step (8), after the prior-free semantic prompt feature F5' is upsampled, it is input into the decoding layer 2 and fused with the feature F2 to obtain a feature Then, the feature F6 is input into the texture-based prompt generation module;

[0073] Step (9), the structure of the texture-based prompt generation module is as follows Figure 3 As shown, after calculating the scale and shift generated by the texture encoder (consisting of 6 convolutional layers) with the feature F6, it is concatenated with F6. Finally, texture-based prompt features are obtained through a convolution operation

[0074] Step (10), the texture-based prompt feature F6' is input into the decoding layer 1 and fused with the feature F1, and then an enhanced image is obtained through the output layer.

[0075] Step (11), the specific process of model training is as follows: The training set consists of a mixture of the LOLv1, LOLv2_Real_captured, LOLv2_Synthetic, and Sony datasets, with a total of 2235 images. Only conventional data preprocessing operations such as random horizontal flipping, random vertical flipping, and random cropping are used for dataset processing to increase the diversity of the training set. The overall loss function L of the model All contains three parts: the enhancement loss L E , the texture encoder training loss L Tand the regularization loss L R . Its overall formula is as follows:

[0076] L All = λ1 × L E + λ2 × L T + λ3 × L R

[0077] where λ1, λ2, and λ3 are weights used to balance the contributions of each loss term to the overall loss function.

[0078] The enhancement loss is composed of the mean absolute error (MAE) loss and the structural similarity (SSIM) loss, and its formula is:

[0079] L E = MAE(P, I) + 1 - SSIM(P, I)

[0080] where P represents the enhanced image and I represents the target normal illumination image.

[0081] The texture encoder training loss L T includes the mean squared error (MSE) loss and the total variation (TV) loss, aiming to promote the training of the texture encoder for generating detailed texture maps. The weights of these losses are controlled by α and β to achieve balance. The formula for L T is as follows:

[0082] L T = α × MSE(T, S) + β × TV(T)

[0083] where T represents the generated texture map, and S is the high-frequency texture map obtained by applying the Sobel filter to the normal illumination image. To ensure the stable generation of T, the TV loss is introduced to enhance smoothness.

[0084] The regularization loss L R is used to prevent overfitting, and its formula is:

[0085]

[0086] where p i represents the value of the model parameter, m represents the total number of model parameters, and μ is a constant, default set to 1×10 -6 .

[0087] Training was carried out using an NVIDIA GeForce 3090 GPU. In addition, the set training image size was 128×128. The batch size was set to 8, and the optimizer used for optimization was the AdamW optimizer with an initial learning rate of 1×10^(-4). The total number of training epochs was set to 1000. During the training process, the texture encoder was also trained together without pre-training. In the relevant loss functions, λ1, λ2, and λ3 were all set to 1, while α and β were set to 0.1 and 0.01 respectively.

[0088] To test the performance of the method of the present invention, qualitative and quantitative experiments were carried out on three publicly available datasets, LOLv1, LOLv2_Real_captured, and LOLv2_Synthetic. The results of the qualitative experiments are as Figure 4 shown. The present invention can generate enhanced images. By comparing with other enhancement methods, it can be observed that the results enhanced by the method of the present invention are closer to normal images. In addition, we also collected different datasets and created a self-built test set. The effects are as Figure 5 shown, and it can be seen that the effect of this method is closer to normal images. In addition, we also conducted tests on unpaired datasets, testing on two publicly available datasets, LIME and MEF. Figure 6 shows the effects of each method in the unpaired dataset test.

[0089] In terms of quantitative analysis, the quantitative comparison experiment results of the method of the present invention with existing image enhancement networks on three publicly available datasets, LOLv1, LOLv2_Real_captured, and LOLv2_Synthetic, and a self-built dataset, Mixed, are shown in Table 1, Table 2, Table 3, and Table 4. In addition, the comparison experiment results of unpaired datasets, LIME and MEF, are shown in Table 5. PSNR in the table is an index commonly used for image compression and restoration quality evaluation, representing the ratio of the maximum power of the signal to the power of the noise. It is usually used to measure the difference between the quality of the restored image and the original image. SSIM is an image quality index that takes into account factors such as image structure, brightness, and contrast, aiming to simulate the perception of image quality by the human visual system. LPIPS is a deep learning-based perceptual image quality evaluation index, aiming to quantify the perceptual similarity between images. And IQE in Table 5 is a reference-free image quality assessment method, mainly used to evaluate the naturalness of images, especially the perceptual quality of images. NIQE does not depend on the original reference image but evaluates its naturalness by modeling the statistical characteristics of the image. ST is a variant of SSIM, targeting applications or variants without a reference image.

[0090] Table 1 Performance comparison of LOLv1 dataset (the optimal result is marked in bold, and the sub-optimal is underlined)

[0091]

[0092] Table 2 Performance comparison of LOLv2_Real_captured dataset (the best results are marked in bold, and the second-best are underlined)

[0093]

[0094]

[0095] Table 3 Performance comparison of LOLv2_Synthetic dataset (the best results are marked in bold, and the second-best are underlined)

[0096]

[0097] Table 4 Performance comparison of Mixed dataset (the best results are marked in bold, and the second-best are underlined)

[0098]

[0099] Table 5 Performance comparison of LIME and MEF datasets (the best results are marked in bold, and the second-best are underlined)

[0100]

[0101]

[0102] By comparing the results in Tables 1 to 4, it can be seen that the present invention has an advantage in consistency in multiple metrics, especially outstanding in SSIM and LPIPS metrics. In terms of PSNR, the present invention performs well on multiple datasets. Especially on the LOLv2_Synthetic (25.3101) and Mixed (23.7893) datasets, it has achieved very high scores, superior to most existing methods. For example, on the LOLv1 and LOLv2_Real_captured datasets, the present invention has also achieved excellent results of 25.9533 and 28.6360. Although it is slightly lower than DCCNet (28.6642) and LLFlow (28.3521) on the LOLv2_Real_captured dataset, the gap is extremely small and the overall effect is still excellent. In contrast, many existing methods such as ZeroDCE and ZeroIG have relatively low PSNR scores, indicating their deficiencies in noise suppression and detail restoration; in terms of SSIM, the present invention performs particularly outstandingly, especially on the LOLv2_Synthetic (0.9400) and Mixed (0.7605) datasets, superior to methods such as RetinexMamba (0.9346) and LLFlow (0.6810) respectively. The SSIM scores of other methods are relatively low, especially on the LOLv1 dataset. Only the SSIM of LLFlow is close to that of the present invention, but the present invention still maintains a high advantage (0.8693). SSIM is a key metric for measuring the structural similarity of images. The high score of the present invention indicates that it can retain the structural information of the image to the greatest extent when enhancing the image, making the restored image more consistent with the original image and the enhancement effect more natural; in terms of LPIPS, the present invention performs particularly excellently. Its LPIPS values on all datasets are the smallest. Especially on the LOLv1 (0.1128) and LOLv2_Real_captured (0.0920) datasets, its LPIPS values are significantly lower than other methods. A low LPIPS value means that the generated image is more natural in perceptual quality and has the least visual distortion. Compared with other methods such as LLFlow and RetinexMamba, although their PSNR and SSIM are relatively high, they perform poorly in LPIPS, indicating that there are certain problems in their visual perceptual quality. In contrast, the present invention can minimize perceptual distortion while enhancing the image and maintain the naturalness of the image.

[0103] As can be seen from Table 5, the present invention shows the optimal image quality on both the LIME and MEF datasets, especially outstanding in terms of naturalness and detail retention.

[0104] Based on the comparison and analysis of the above performance indicators, the present invention performs excellently on all test data sets, demonstrating its outstanding capabilities in tasks such as image enhancement, restoration, and multi-exposure fusion. In particular, its advantages in image structure fidelity, detail restoration, and perceptual quality prove that it is superior to existing mainstream methods.

[0105] The above content is a further detailed description of the present invention in combination with specific / preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, they can also make several substitutions or modifications to these described embodiments, and these substitution or modification methods should all be regarded as belonging to the protection scope of the present invention.

[0106] The parts not detailed in the present invention belong to the well-known technology in the art.

Claims

1. An adaptive low-light image enhancement method based on customized prompt learning, characterized in that, It includes the following steps: Step (1): Construct an adaptive low-light image enhancement model based on customized prompt learning; The adaptive low-light image enhancement model adopts a U-Net architecture and designs a dual-path processing mechanism, including an encoding-decoding path composed of an encoder and a decoder, and a prompt generation path composed of a semantics-based prompt generation module, a prior-free prompt generation module, and a texture-based prompt generation module; Step (2): First, the encoder of the encoding-decoding path extracts features from the input low-light image to obtain multi-scale and multi-level image feature information; Step (3): The decoder further processes the image feature information extracted by the encoder, uses the multi-head attention mechanism for information fusion, and gradually restores the size of the image in combination with the upsampling operation; Step (4): During the decoder processing, multiple different prompt features output by the prompt generation path are sequentially combined, and the recovery of image information is assisted by the prompts, and finally the enhanced image is output; Step (5): Construct the overall loss function of the adaptive low-light image enhancement model and train the model.

2. The adaptive low-light image enhancement method based on customized prompt learning according to claim 1, wherein The encoders and decoders of different layers are all composed of multiple Transformer blocks.

3. The adaptive low-light image enhancement method based on customized prompt learning according to claim 1, characterized in that, In step (2), first, the input layer expands the number of channels of the input low-light image, and then the processed image is passed into the encoder for subsequent processing; The input size is img ∈ R H×W×3 of the low-light image, which first passes through a 3×3 convolution, namely the input layer, to obtain the underlying feature F0 ∈ R H×W×C ; Where C represents the number of channels, H is the height of the image, and W is the width of the image.

4. The adaptive low-light image enhancement method based on customized prompt learning according to claim 3, characterized in that The encoder mentioned above is divided into three layers, namely Encoding Layer 1, Encoding Layer 2, and Encoding Layer 3, which are composed of 1, 2, and 4 Transform blocks in sequence. The image processed by the input layer is subjected to feature extraction through three consecutive encoding layers to obtain features F1 ∈ R H ×W×C , where between each encoding layer, the features are tiled through downsampling operations to obtain features 5. The adaptive low-light image enhancement method based on customized prompt learning according to claim 4, wherein, During the decoder processing, the prompt features output by the prompt generation path are combined. The specific operations are as follows: The decoder includes decoding layer 1, decoding layer 2, and decoding layer 3, which are sequentially composed of 1, 2, and 4 Transform blocks; a prompt mechanism is introduced in stages in the decoder of the adaptive low-light image enhancement model: the semantics-based prompt generation module uses the pre-trained CLIP model to generate semantic symbols, and then fuses them with the original features to generate semantic prompts to enhance the high-level understanding ability of the model; the prior-free prompt generation module adaptively enhances the model by dynamically generating implicit semantic prompts for diverse degradation types in low-light images; the texture-based prompt generation module effectively guides detail recovery and noise suppression by introducing high-frequency texture detail prompts; the obtained prompt features are combined layer by layer with the features generated by the encoder through the feature fusion module to gradually improve the image quality, and finally a high-quality normal illumination image is output.

6. The adaptive low-light image enhancement method based on customized prompt learning according to claim 5, characterized in that The semantics-based prompt generation module uses the pre-trained CLIP model to generate semantic symbols, and then fuses them with the original features to generate semantic prompts to enhance the high-level understanding ability of the model. The specific operations are as follows: For the semantics-based prompt generation module, first, a randomly generated learnable vector is used as the input and passed to the frozen CLIP text encoding layer; the obtained semantic symbol Emb and the input feature F4 are fused through a dual-gated self-attention mechanism to generate a semantic prompt S; finally, the semantic prompt S is concatenated with the input feature F4, and then passed through a 3×3 convolution operation to obtain the semantic prompt feature F4′, that is, the high-level prompt information; the corresponding structured encoding formula is as follows: Emb=CLIP text (Token)(Formula 1) v = [ReshapeF4; FCEmb] (Formula 2) S = ReshapeF4 + λ × tanh(θ) × Attn(v) (Equation 3) F4′ = Conv 1×1 ([F4; Reshape(S)]) (Equation 4) Among them, Token represents a learnable vector randomly generated with the same shape as the CLIP text input, Emb represents semantic symbols, and CLIP text represents the CLIP text encoding layer; is the output feature of the last encoding layer; Reshape represents a reshaping operation, FC represents a linear transformation of the channel dimension, and Attn represents a multi-head attention mechanism; represents the semantic prompt feature; Finally, after the semantic hint feature F4′ undergoes an upsampling operation, it is concatenated with the feature F3 output by the encoding layer 3. After adjusting the channels through a convolution operation, it is input into the decoding layer 3 to obtain the output feature 7. The adaptive low-light image enhancement method based on customized prompt learning according to claim 6, characterized in that, The prior-free hint generation module adaptively enhances various degradation types in low-light images by dynamically generating implicit semantic hints. The specific operations are as follows: In the prior-free prompt generation module, first, global average pooling operation is performed on the output feature F5 of the decoding layer 3, and then through a 1×1 convolution operation and an operation to generate weights, the weight w is obtained i ; immediately afterwards , multiply the weight w i by the learnable vector P c , and then obtain the prior-free implicit semantic cue G through a 3×3 convolution operation; concatenate the generated prior-free implicit semantic cue G with the input feature F5, and then obtain the prior-free semantic cue feature F5′ through a 3×3 convolution operation; this process is described by the following equation: w i = Softmax(Conv 1×1 (GAP(F5)))(Equation 7) F5′ = Conv 1×1 F5; G↑(Equation 8) Among them, P c ∈R N×16×16×4C represents a randomly initialized learnable vector prompt component, N is the number of learning components, Conv 3×3 represents a 3×3 convolution operation, is the output feature from the decoding layer 3; GAP represents the global average pooling operation, w i represents the weight calculated through the convolution operation, Softmax is used to generate the weight, G represents the generated prior-free implicit semantic prompt, is the prior-free semantic prompt feature; After upsampling F5' and concatenating it with the feature F2 of the encoding layer 2, and adjusting the channels through a convolution operation, it is input to the decoding layer 2 to obtain the output feature 8. The adaptive low-light image enhancement method based on customized prompt learning according to claim 7, wherein The texture-based hint generation module effectively guides detail restoration and noise suppression by introducing high-frequency texture detail hints. The specific operations are as follows: The texture-based hint generation module processes the input low-light image through a texture encoder to generate a texture map T, i.e., a high-frequency texture detail hint; the texture encoder consists of six convolutional layers, followed by ReLU activation layers in the first five convolutional layers; the texture map T generated by the texture encoder is fused with the output features of decoding layer 2 to generate texture-based hint features This process is described by the following formula: F6′ = Conv 1×1 ([F6; F6 × scaleT + shift(T)]) (Equation 9) Among them, T represents the generated texture map, and scale and shift respectively represent 3×3 convolution operations for controlling the effective range and offset of features; After upsampling F6′, it is concatenated with the feature F1 output by the encoding layer 1, and after adjusting the channels through a convolutional operation, it is output to the decoding layer 1 to obtain the output feature F7∈R H×W×C .

9. The adaptive low-light image enhancement method based on customized prompt learning according to claim 8, wherein The output feature F7 of the decoding layer 1 passes through an output layer composed of a self-attention mechanism for image reconstruction and a 3×3 convolution operation to obtain the enhanced image P.

10. The adaptive low-light image enhancement method based on customized prompt learning according to claim 1, characterized in that, The overall loss function L of the adaptive low-light image enhancement model All consists of three parts: the enhancement loss L E , the texture encoder training loss L T and the regularization loss L R ; its overall formula is as follows: L All = λ1 × L E + λ2 × L T + λ3 × L R (Equation 10) Among them, λ1, λ2, and λ3 are weights used to balance the contributions of each loss term to the overall loss function; The enhancement loss is composed of the mean absolute error (MAE) loss and the structural similarity (SSIM) loss. Its formula is: L E = MAE(P, I) + 1 - SSIM(P, I) (Equation 11) Among them, P represents the enhanced image, and I represents the target normal illumination image; Texture encoder training loss L T includes mean squared error (MSE) loss and total variation (TV) loss, aiming to facilitate the training of a texture encoder for generating detailed texture maps; the weights of these losses are controlled by α and β to achieve a balance; the formula for L T is as follows: L T = α × MSE(T, S) + β × TV(T) (Equation 12) Among them, T represents the generated texture map, and S is the high-frequency texture map obtained by applying the Sobel filter to the normal illumination image; Regularization loss L R Used to prevent overfitting, and its formula is: where p i represents the value of the model parameter, m represents the total number of model parameters, and μ is a constant.

Citation Information

Cited By

  • Layered hybrid network-based underlying visual color imaging learning method and device

    CN121353105A

  • No-reference AIGC image quality evaluation method for theme generation image group

    CN121599947A