Low-illumination image enhancement method, electronic equipment and storage medium

By introducing a contrast language-image pretrained model (CLIP) to learn positive and negative prompt pairs and alternately trained with the low-illumination enhancement model, the problem of insufficient robustness and generalization ability of low-illumination image enhancement models in the prior art is solved, and a more efficient low-illumination image enhancement effect is achieved.

CN120070256APending Publication Date: 2025-05-30XIAN NOVASTAR TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311568964.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-22
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The robustness and generalization ability of the existing low-illumination image enhancement methods are lacking, making it difficult to effectively process low-illumination images in real scenes.

Method used

The contrast language-image pre-trained model (CLIP) is used to learn the positive and negative prompt pairs of low-illumination images and normal illumination images, and alternately train and update with the low-illumination enhancement model to improve the training effect of the model.

Benefits of technology

By introducing the CLIP model, the robustness and generalization performance of the low-illumination image enhancement model are improved, and low-illumination images can be processed more accurately and visual quality can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070256A_ABST
    Figure CN120070256A_ABST
Patent Text Reader

Abstract

The invention is applicable to the technical field of image processing, and provides a low-illumination image enhancement method, electronic equipment and a storage medium, and the low-illumination image enhancement method comprises the steps: learning positive and negative prompt pairs of a low-illumination image and a normal-illumination image through employing a comparison language-image pre-training model; inputting the low-illumination image into a low-illumination enhancement model to obtain an enhanced image output by the low-illumination enhancement model; inputting the enhanced image into a contrast language-image pre-training model to obtain image features of the enhanced image; updating the positive and negative prompt pairs based on the image features of the enhanced image; updating parameters of the low-illumination enhancement model at least based on the similarity between the image features of the enhanced image and the updated positive and negative prompt pairs; and inputting the low-illumination image into the low-illumination enhancement model and the following steps repeatedly until a stop condition is met, and obtaining a trained low-illumination enhancement model which is used for performing illumination enhancement on the low-illumination image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of image processing, and particularly relates to a low-light image enhancement method, an electronic device, and a storage medium. Background Art

[0002] Low-light enhancement is to process problems such as low brightness, low contrast, noise, and artifacts existing in low-illumination images, that is, images taken under insufficient lighting conditions, so as to improve their visual quality.

[0003] Low-illumination enhancement methods can be divided into two categories: traditional methods and deep learning methods. Traditional low-illumination enhancement algorithms are prone to color difference, artifact halos, and loss of a large amount of detail information when enhancing images in practice, and their speed, accuracy, and robustness are not as good as deep learning methods.

[0004] In the deep learning methods of low-illumination enhancement, supervised learning methods occupy the mainstream position. They mainly use different network models to learn the mapping relationship and feature representation between paired images, effectively improving the visual quality of enhanced images. However, due to the limitations of the size of the training dataset and the diversity of low-illumination scenarios, the robustness and generalization ability of supervised learning methods are lacking, and it is difficult to process low-illumination images in real scenarios. For unsupervised learning low-illumination enhancement methods, they often rely on ideal assumptions such as average brightness and gray world models, and the robustness and generalization ability of these methods are also lacking. Summary of the Invention

[0005] Embodiments of this application provide a low-illumination image enhancement method, an electronic device, and a storage medium, which can solve the problem of the lack of robustness and generalization ability in low-illumination enhancement methods in related technologies.

[0006] In a first aspect, embodiments of this application provide a low-illumination image enhancement method, which includes: using a contrastive language-image pre-training model to learn positive and negative prompt pairs of low-illumination images and normal-illumination images; inputting the low-illumination image into a low-illumination enhancement model to obtain an enhanced image output by the low-illumination enhancement model; inputting the enhanced image into the contrastive language-image pre-training model to obtain the image features of the enhanced image; updating the positive and negative prompt pairs based on the image features of the enhanced image; updating the parameters of the low-illumination enhancement model at least based on the similarity between the image features of the enhanced image and the updated positive and negative prompt pairs; repeating the steps of inputting the low-illumination image into the low-illumination enhancement model and subsequent steps until a stop condition is met to obtain a trained low-illumination enhancement model, and the low-illumination enhancement model is used to perform light enhancement on low-illumination images.

[0007] Second aspect, an embodiment of the present application provides a low-light image enhancement device, which includes: a learning module, configured to learn positive and negative prompt pairs of low-light images and normal-light images by using a contrastive language-image pre-training model; an enhancement module, configured to input a low-light image into a low-light enhancement model to obtain an enhanced image output by the low-light enhancement model; a feature module, configured to input the enhanced image into the contrastive language-image pre-training model to obtain image features of the enhanced image; a first update module, configured to update the positive and negative prompt pairs based on the image features of the enhanced image; a second update module, configured to update the parameters of the low-light enhancement model at least based on the similarity between the image features of the enhanced image and the updated positive and negative prompt pairs; an acquisition module, configured to obtain a trained low-light enhancement model, and the low-light enhancement model is used to perform light enhancement on low-light images.

[0008] Third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable by the processor. When the processor executes the computer program, the low-light image enhancement method described in any item of the first aspect above is implemented.

[0009] Fourth aspect, an embodiment of the present application provides a computer-readable storage medium, storing a computer program, and when the computer program is executed by a processor, the low-light image enhancement method described in any item of the first aspect above is implemented.

[0010] Fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on an electronic device, the electronic device is enabled to execute the low-light image enhancement method described in any item of the first aspect above.

[0011] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: By introducing a contrastive language-image pre-training (CLIP) model, which has undergone a large number of classification trainings of natural language-pictures and contains rich prior information, and its classification accuracy and generalization performance are significantly higher than those of other image classification models trained on a single dataset. Using the CLIP model to perform binary classification on the illuminance of images can achieve good robustness and generalization performance. Using CLIP for classification requires providing corresponding text prompts for the CLIP model, and generally manual design of prompts is used in the related art. However, in the present application, manual design is abandoned, and positive and negative prompt pairs regarding image illuminance are selected and alternately trained and updated with the low-light enhancement model to obtain more accurate positive and negative prompt pairs, improving the training effect of the low-light enhancement model. Description of the Drawings

[0012] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0013] Figure 1 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application;

[0014] Figure 2 It is a schematic flowchart of a low-light image enhancement method provided by an embodiment of the present application;

[0015] Figure 3 is Figure 2 The specific flowchart of S1 in;

[0016] Figure 4 It is combined with the contrastive language-image pre-training model Figure 2 The schematic diagram of S1 in;

[0017] Figure 5 It is a schematic structural diagram of a low-light enhancement model provided by an embodiment of the present application;

[0018] Figure 6 It is a schematic structural diagram of a multi-scale refinement hybrid spatial attention module provided by an embodiment of the present application;

[0019] Figure 7 It is a schematic structural diagram of a light-selective normalization module provided by an embodiment of the present application;

[0020] Figure 8 It is a schematic structural diagram of a low-light image enhancement device provided by an embodiment of the present application. Specific Embodiments

[0021] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures and technologies are presented to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from hindering the description of the present application.

[0022] It should be understood that when used in the specification and claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0023] It should also be understood that the term "and / or" as used in the specification and appended claims of this application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0024] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" depending on the context.

[0025] In addition, in the description of the specification and appended claims of this application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0026] Reference to "one embodiment" or "some embodiments" etc. described in the specification of this application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0027] The low-light image enhancement method provided by the embodiments of this application can be applied to an electronic device, and the electronic device includes, but is not limited to, electronic devices with computing functions such as servers, server clusters, mobile phones, tablet computers, laptop computers, desktop computers, personal digital assistants, and wearable devices, etc. The embodiments of this application do not impose any restrictions on the specific type of the electronic device.

[0028] Figure 1 Shown is a block diagram of a part of the structure of the electronic device provided by the embodiments of this application. Refer to Figure 1 , the electronic device includes: a processor 10, a memory 20, a bus 30, an input device 40, an output device 50, and a communication device 60. The processor 10 and the memory 20 are connected to each other through the bus 30, and the input device 40, the output device 50, and the communication device 60 are also connected to the bus 30. Those skilled in the art can understand that Figure 1The structure of the electronic device shown does not limit the electronic device, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0029] The following will specifically introduce each component of the electronic device in conjunction with Figure 1 the electronic device:

[0030] The processor 10 is the control center of the electronic device and can execute various functions and process data by running programs stored in the memory 20. The processor 10 can be a Central Processing Unit (CPU), and the processor 10 can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In some embodiments, the processor 10 may include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0031] The memory 20 is used to store the operating system, application programs, BootLoader, data, and other programs, such as the program code of computer programs. The memory 20 can also be used to temporarily store data required for and generated during the execution of programs. The memory 20 can include high-speed random access memory and can also include non-volatile memory, such as flash memory, hard disks, multimedia cards, card-type memories, etc. The memory 20 can include storage units provided inside the electronic device, such as the hard disk of the electronic device, and / or removable external storage units, such as external hard disks, USB flash drives, Smart Media Cards (SMCs), Secure Digital (SD) cards, etc.

[0032] The input device 40 may include at least one of a keyboard, a mouse, a touch panel, a joystick, etc., and is used to collect the input operations of the user to generate corresponding input signals.

[0033] The output device 50 is used to output information to be provided to the user. The output device 50 generally includes a display. Optionally, a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. can be adopted. In addition, the output device may further include a speaker.

[0034] The communication device 60 may include a modem, a network card, etc., and is used to establish a network connection with other electronic devices and communicate with each other.

[0035] The low-light image enhancement method provided by the embodiments of the present application can be implemented as a computer software program. For example, the embodiments of the present application provide a computer program product, which includes a computer program carried on a computer-readable medium. The computer program contains program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 60, and / or installed from a removable external storage unit. When the computer program is executed by the processor 10, various functions defined in the low-light image enhancement method provided by the embodiments of the present application are implemented.

[0036] Figure 2 A schematic flowchart of the low-light image enhancement method provided by the present application is shown. By way of example and not limitation, this method can be applied to the above-mentioned electronic device.

[0037] S1: Use the contrastive language-image pre-training model to learn positive and negative cue pairs of low-light images and normal-light images.

[0038] The contrastive language-image pre-training (CLIP) model is a pre-trained multi-modal neural network model, originally used to match images and texts. The CLIP model includes an image encoder Φ image and a text encoder Φ text , the image encoder is used to extract image features of the input image, and the text encoder is used to extract text features of the input natural language text. According to the similarity, or rather, the similarity degree, of the image features and the text features, it can be determined whether there is a match between the input image and the input natural language text.

[0039] The CLIP model can be used for image classification. To achieve this task, a set of prompts needs to be prepared for the target categories, and the prompts correspond to the target categories one by one. Generally, common prompts are designed manually. For example, "A photo of a {object}", where {object} is filled with the target category respectively. Input the prompts of each target category into the text encoder of the CLIP model to obtain the text features of each target category; input the image to be recognized into the image encoder of the CLIP model to obtain the image features of the image to be recognized; calculate the similarity between the image features and the text features of each target category. Since both the image features and the text features are in the format of vectors, calculating the similarity generally means calculating the distance between vectors; the target category corresponding to the text feature with the highest similarity to the image features, that is, the closest distance, is the classification result of the image to be recognized.

[0040] The design of the prompts will affect the performance of CLIP in image classification. For example, in order to achieve better classification performance, for fine-grained classification tasks, more detailed prompts need to be designed. In this application, the CLIP model is introduced to perform binary classification on low-light images and normally illuminated images. Since the description of light is relatively complex, the optimal pair of prompts varies depending on the image situation, and it is difficult to design a pair of prompts manually and the generalization performance is limited. Therefore, in this application, manual design of prompts is abandoned, and instead, positive and negative prompts for image illumination are learned based on CLIP.

[0041] Please refer to Figure 3 and Figure 4 , in some embodiments provided in this application, learning the positive and negative prompts specifically includes:

[0042] S11: Initialize a pair of natural language prompts.

[0043] Initialize a pair of natural language prompts T p ∈R N×512 and T n ∈R N×512 , where N represents the number of words included in the prompt, T p corresponds to the positive prompt for normal illumination, and T n corresponds to the negative prompt for low light. The content of the natural language prompt can be completely random or directly set, and the content of the two natural language prompts can be the same or different.

[0044] S12: Input a low-light image, a normally illuminated image, and a pair of natural language prompts into the contrastive language-image pre-training model to obtain the image features of the low-light image, the image features of the normally illuminated image, and the initial positive and negative prompts.

[0045] The content of the normal illumination image and the low illumination image can be the same or different. The low illumination image is input into the image encoder of the CLIP model to obtain the image features of the low illumination image, and the normal illumination image used as the reference image is input into the image encoder of the CLIP model to obtain the image features of the normal illumination image. The positive and negative prompt pairs include a positive prompt representing normal illumination and a negative prompt representing low illumination. The initial positive prompt is the text feature of T p and the initial negative prompt is the text feature of T n .

[0046] S13: Update the positive and negative prompt pairs according to the similarity between the image features of the low illumination image, the image features of the normal illumination image and the positive and negative prompt pairs until the iteration condition is met.

[0047] In the case of correct classification, the similarity between the image features of the low illumination image and the negative prompt is higher than the similarity between the image features of the low illumination image and the positive prompt, and the similarity between the image features of the low illumination image and the negative prompt is lower than the similarity between the image features of the low illumination image and the positive prompt. However, the initial prompt pairs often cannot meet such conditions. Therefore, the loss can be calculated according to the similarity, and then the positive and negative prompt pairs can be updated according to the loss. In this application, learning the positive and negative prompt pairs is to update the text features output by the text encoder of the CLIP model, rather than directly updating the natural language. That is, the positive and negative prompt pairs refer to Φ text (T p ) and Φ text (T n ), rather than T p and T n . Specifically, the binary cross-entropy loss can be calculated according to the similarity between the image features of the low illumination image, the image features of the normal illumination image and the positive and negative prompt pairs. The calculation formula is as follows:

[0048]

[0049]

[0050] where I ∈ {I low , I normal} represents the input low illumination image or normal illumination image, y represents the classification label of the current image, and for the negative sample I low the value is 0, and for the positive sample I normal the value is 1.

[0051] Then, the positive and negative prompt pairs can be updated according to the binary cross-entropy loss until the iteration condition is met. The iteration condition can be that the number of iterations reaches the threshold, or the binary cross-entropy loss is less than or equal to the set value.

[0052] S2: Input the low illumination image into the low illumination enhancement model to obtain the enhanced image output by the low illumination enhancement model.

[0053] Please refer to Figure 5 , the low-light enhancement model includes 2n multi-scale refinement hybrid spatial attention (MRSAT) modules, where n is an integer greater than 1, and the parameters of each MRSAT are independent and not shared. Here, n = 3 is taken as an example for illustration.

[0054] The input image Input ∈ R H×W×3 , first passes through a convolutional layer with a 3×3 convolutional kernel to increase the feature channels to 32, and then is input into the first MRSAT module. The output features of each of the first three MRSAT modules are downsampled and then input into the next MRSAT module. The specific formula is as follows:

[0055] Out 1 = ReLU(conv_1(Input) + b 1 ) (3)

[0056] Out 2 = Down 1 (MRSAT 1 (Out 1 )) (4)

[0057] Out 3 = Down 2 (MRSAT 2 (Out 2 )) (5)

[0058] Out 4 = Down 3 (MRSAT 3 (Out 3 )) (6)

[0059] Among them, conv_1 represents the convolutional operation on the input image, b 1 represents the bias parameter of the network, ReLU(·) represents the activation function after the convolutional layer, and Down(·) represents the downsampling operation.

[0060] After the output features of each of the last three MRSAT modules are upsampled, they are subjected to residual connection with one of the input features of the first three MRSAT modules with the same size, and then input into the next MRSAT module. The specific formula is as follows:

[0061] Out 5 = Up 1 (MRSAT 4 (Out 4 )) (7)

[0062] Out 6 = Up2 (MRSAT 5 (Out 3 +Out 5 )) (8)

[0063] Out 7 = Up 3 (MRSAT 6 (Out 2 +Out 6 )) (9)

[0064] Out 8 = ReLU(conv_2(Out 1 +Out 7 ) + b 2 ) (10)

[0065] Among them, conv_2 represents the convolution operation, and the parameters of this convolution operation are not shared with those of conv_1. b 2 represents the bias parameter of the network, ReLU(·) represents the activation function after the convolutional layer, and Up(·) represents the upsampling operation.

[0066] Specifically, please refer to Figure 6 , the MRSAT module includes a light selectivity normalization module, an attention module, and a perceptron module. The first n MRSAT modules do not enable the hybrid mechanism and only adopt the self-attention mechanism. The specific formula is as follows.

[0067] Out 1 = MLP(MSA(CSNorm(input))) + input (11)

[0068] Out 2 = MLP(CSNorm(Out 1 )) + Out 1 (12)

[0069] Among them, input represents the input image features, Out 1 represents the output features in the middle of the module processing, and Out 2 represents the output image features. CSNorm(·) represents the light selectivity normalization module, MSA(·) represents the multi-head self-attention module, and MLP(·) represents the multi-layer perceptron module.

[0070] The last n MRSAT modules perform residual connection on the features obtained by the previous module processing and the input features, and enable the hybrid attention mechanism. The specific formula is as follows.

[0071] Out 1 = MSA(CSNorm(Xcur )) + MCA(CSNorm(X pre ) + CSNorm(X cur )) (13)

[0072] Out 2 = MLP(Out 1 ) + X cur (14)

[0073] Out 3 = MLP(CSNorm(Out 1 )) + Out 2 (15)

[0074] Among them, X cur represents the currently input image feature, X pre represents the image feature obtained by processing the previous module, Out 1 and Out 2 represent the output features in the middle of module processing, Out 3 represents the output image feature, CSNorm(·) represents the illumination selective normalization module, MSA(·) represents the multi-head self-attention module, MCA(·) represents the multi-head cross-attention module, and MLP(·) represents the multi-layer perceptron module.

[0075] Please refer to Figure 7 , the illumination selective normalization module is composed of a brightness selective normalization module and an adaptive light-related channel selection gating module. Given an image feature x ∈ R H×W×C , IN normalizes x by subtracting the mean μ(x) and dividing by the standard deviation σ(x), and the specific formula is as follows.

[0076]

[0077] Among them, μ(x) and σ(x) are calculated independently in each channel and spatial dimension, and γ, β ∈ R C are scalable parameters learned from the data. Since IN(·) can reduce the brightness difference between instances, the normalized feature σ(x) has a robust representation independent of the brightness condition, enabling the network to adapt to various brightness scenarios and improving its generalization ability.

[0078] The gating module is responsible for the adaptive light-related channel selection operation, which can effectively enhance the generalization ability of the model and reduce the information loss caused by normalization. The gating module outputs a series of binary values to combine the normalized channels and the original channels, and the specific formula is as follows.

[0079] x n+1 = (1 - g)·x n + g·x′n (17)

[0080] Among them, g represents the binary value in the channel dimension, and · represents channel multiplication. The gating operation activates or deactivates the channel through the binary value to selectively normalize the channel. Thus, the generated feature x n+1 eliminates the influence of brightness and retains the basic information with the channel unchanged.

[0081] In the MRSAT module, the combination of self-attention and cross-attention expands the receptive field, thereby obtaining rich spatial correlations and semantic features. The multi-scale refinement allows the model to extract feature information at different levels, which helps to enhance the perception ability of the brightness and darkness of different regions of the image.

[0082] S3: Input the enhanced image into the contrastive language-image pre-training model to obtain the image features of the enhanced image.

[0083] S4: Update the positive and negative prompt pairs based on the image features of the enhanced image.

[0084] The marginal ranking loss can be calculated based on the image features of the enhanced image, and the specific formula is as follows:

[0085]

[0086]

[0087]

[0088] Among them, m 0 ∈[0,1] represents the target distance between the normal illumination image and the low illumination image obtained by the CLIP model. Ideally, m 0 is 1. In practice, the value of m 0 can be set to be less than 1 and close to 1. For example, a value is selected from the range of 0.85 to 0.95 as the value of m 0 to increase the distance between the low illumination image and the normal illumination image as much as possible. m 1 represents the target distance between the enhanced image obtained by the CLIP model and the normal illumination image. Ideally, m 1 is 0. In practice, the value of m 1 can be set to be greater than 0 and close to 0. For example, a value is selected from the range of 0.1 to 0.3 as the value of m 1 to ensure that the enhanced result is similar to the normal illumination image. Iresult represents the enhanced image, and Φ image (I) obtained by substituting Iresult into Φ image (I result ) is the image feature of the enhanced image. S n(I) represents the similarity score between the image and the negative prompt, S p (I) represents the similarity score between the image and the positive prompt, S p (I result ) The higher it is, the more similar the enhanced image is to the positive prompt.

[0089] Then, the positive and negative prompt pairs can be updated using the margin ranking loss.

[0090] S5: Update the parameters of the low-light enhancement model based at least on the similarity between the image features of the enhanced image and the updated positive and negative prompt pairs.

[0091] Specifically, the contrastive language-image perception loss can be calculated using the similarity between the image features of the enhanced image and the updated positive and negative prompt pairs. The specific formula is as follows.

[0092]

[0093] Among them, I result represents the enhanced image, Φ image and Φ text represent the image encoder and the text encoder in the CLIP model, respectively.

[0094] Then, the parameters of the low-light enhancement model can be updated based at least on the contrastive language-image perception loss. Specifically, the identity loss can be calculated based on the structural similarity between the low-light image and the enhanced image, and the parameters of the low-light enhancement model can be updated by combining the contrastive language-image perception loss and the identity loss. The specific calculation formula of the identity loss is as follows.

[0095]

[0096] Among them, α l is the weight of the l-th layer of the image encoder in the contrastive language-image pre-training model, represents the second norm.

[0097] S6: Determine whether the stop condition is met.

[0098] The stop condition can be that the number of iterations reaches a threshold, or the loss is less than or equal to a preset value.

[0099] If so, jump to S7; otherwise, jump to S2.

[0100] S7: Obtain the trained low-light enhancement model.

[0101] The low-light enhancement model is used to perform light enhancement on low-light images.

[0102] Through the implementation of this embodiment, the CLIP model is introduced. This model has undergone a large amount of natural language-image classification training, which contains rich prior information. Its classification accuracy and generalization performance are significantly higher than those of other image classification models trained on a single dataset. Using the CLIP model to perform binary classification on the illuminance of images can achieve good robustness and generalization performance.

[0103] When using CLIP for classification, corresponding text prompts need to be provided for the CLIP model. In related technologies, these are generally manually designed. In this embodiment, however, manual design is abandoned, and positive and negative prompt pairs regarding image illuminance are chosen to be learned, and are alternately trained and updated with the low-light enhancement model to obtain more accurate positive and negative prompt pairs, improving the training effect of the low-light enhancement model.

[0104] In addition, this embodiment uses a multi-scale refined hybrid attention module to capture the long-range correlations of features in different scale spatial domains of the image, effectively enhancing the representation ability of the network. The introduced light-selective normalization module can adaptively normalize the feature channels related to brightness, improving the generalization of the network.

[0105] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.

[0106] Figure 8 FIG. shows a schematic structural diagram of a low-light image enhancement device provided by an embodiment of the present application. The low-light image enhancement device includes a learning module 11, an enhancement module 12, a feature module 13, a first update module 14, a second update module 15, and an acquisition module 16.

[0107] The learning module 11 is used to learn positive and negative prompt pairs of low-light images and normal-illuminance images using a contrastive language-image pre-training model.

[0108] The enhancement module 12 is used to input a low-light image into a low-light enhancement model to obtain an enhanced image output by the low-light enhancement model.

[0109] The feature module 13 is used to input the enhanced image into a contrastive language-image pre-training model to obtain the image features of the enhanced image.

[0110] The first update module 14 is used to update the positive and negative prompt pairs based on the image features of the enhanced image.

[0111] The second update module 15 is used to update the parameters of the low-light enhancement model at least based on the similarity between the image features of the enhanced image and the updated positive and negative prompt pairs.

[0112] An acquisition module 16 is configured to obtain a trained low-light enhancement model, which is used to perform light enhancement on low-light images.

[0113] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.

[0114] The embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in the foregoing method embodiments can be implemented.

[0115] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of the present application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps in the foregoing method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / electronic device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0116] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not described or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0117] Those of ordinary skill in the art will realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0118] In the embodiments provided in this application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0119] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0120] The above embodiments are only used to illustrate the technical solutions of this application, not to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included in the protection scope of this application.

Claims

1. A low - illumination image enhancement method, characterized in that, it includes: Using a contrastive language - image pre - training model to learn positive and negative cue pairs of low - illumination images and normal - illumination images; Inputting a low - illumination image into a low - illumination enhancement model to obtain an enhanced image output by the low - illumination enhancement model; Inputting the enhanced image into the contrastive language - image pre - training model to obtain the image features of the enhanced image; Updating the positive and negative cue pairs based on the image features of the enhanced image; Updating the parameters of the low - illumination enhancement model at least based on the similarity between the image features of the enhanced image and the updated positive and negative cue pairs; Repeating the steps of inputting the low - illumination image into the low - illumination enhancement model and subsequent steps until a stop condition is met, to obtain a trained low - illumination enhancement model, and the low - illumination enhancement model is used to perform illumination enhancement on low - illumination images.

2. The method according to claim 1, characterized in that, The step of using a contrastive language - image pre - training model to learn positive and negative cue pairs of low - illumination images and normal - illumination images includes: Initializing a pair of natural language cues; Inputting a low - illumination image, a normal - illumination image, and the pair of natural language cues into the contrastive language - image pre - training model to obtain the image features of the low - illumination image, the image features of the normal - illumination image, and the initial positive and negative cue pairs, where the positive and negative cue pairs include a positive cue representing normal illumination and a negative cue representing low illumination; Updating the positive and negative cue pairs according to the similarity between the image features of the low - illumination image, the image features of the normal - illumination image, and the positive and negative cue pairs until an iteration condition is met.

3. The method according to claim 2, characterized in that, The step of updating the positive and negative cue pairs according to the similarity between the image features of the low - illumination image, the image features of the normal - illumination image, and the positive and negative cue pairs includes: Calculating a binary cross - entropy loss according to the similarity between the image features of the low - illumination image, the image features of the normal - illumination image, and the positive and negative cue pairs; Updating the positive and negative cue pairs according to the binary cross - entropy loss.

4. The method according to claim 1, characterized in that, The step of updating the positive and negative cue pairs based on the image features of the enhanced image includes: Calculating a margin ranking loss based on the image features of the enhanced image; Updating the positive and negative cue pairs using the margin ranking loss.

5. The method according to claim 1, characterized in that, The step of updating the parameters of the low - illumination enhancement model at least based on the similarity between the image features of the enhanced image and the updated positive and negative cue pairs includes: Calculating a contrastive language - image perception loss using the similarity between the image features of the enhanced image and the updated positive and negative cue pairs; Updating the parameters of the low - illumination enhancement model at least based on the contrastive language - image perception loss.

6. The method according to claim 5, characterized in that, The step of updating the parameters of the low - illumination enhancement model at least based on the contrastive language - image perception loss includes: Calculating an identity loss based on the structural similarity between the low - illumination image and the enhanced image; Update the parameters of the low-light enhancement model by combining the contrastive language-image perception loss and the identity loss.

7. The method according to claim 1, wherein, the low-light enhancement model includes 2n multi-scale refined hybrid spatial attention modules. The output features of each of the first n multi-scale refined hybrid spatial attention modules are downsampled and then input into the next multi-scale refined hybrid spatial attention module. The output features of each of the last n multi-scale refined hybrid spatial attention modules are upsampled and then subjected to a residual connection with one of the input features of the first n multi-scale refined hybrid spatial attention modules having the same size, and then input into the next multi-scale refined hybrid spatial attention module.

8. The method according to claim 7, wherein, the multi-scale refined hybrid spatial attention module includes a light selective normalization module, a self-attention module, a cross-attention module, and a perceptron module.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein, when the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium storing a computer program, wherein, when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.