Image enhancement model training method and image enhancement method

WO2026166295A1PCT designated stage Publication Date: 2026-08-13TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-01-12
Publication Date
2026-08-13

Smart Images

  • Figure CN2026071905_13082026_PF_FP_ABST
    Figure CN2026071905_13082026_PF_FP_ABST
Patent Text Reader

Abstract

An image enhancement model training method and an image enhancement method. The image enhancement model training method is executed by an electronic device, and comprises: acquiring a plurality of sample images and first image enhancement prompt information of each sample image (202); for each sample image, inputting the sample image and the first image enhancement prompt information into an initial image enhancement model, and by means of the initial image enhancement model, using the first image enhancement prompt information as condition information for guiding image enhancement, to perform image enhancement on the sample image to obtain an enhanced image (204); on the basis of the first image enhancement prompt information, generating image restoration prompt information for restoring the enhanced image to the sample image, inputting the enhanced image and the image restoration prompt information into an initial image restoration model, and by means of the initial image restoration model, using the image restoration prompt information as condition information for guiding image restoration, to perform image restoration on the enhanced image to obtain a restored image (206); comparing a difference between the sample image and the restored image to determine a sample training loss value of the sample image (208); and on the basis of the respective sample training loss values of the plurality of sample images, adjusting model parameters of the initial image enhancement model to obtain a trained image enhancement model (210).
Need to check novelty before this filing date? Find Prior Art

Description

Image augmentation model training methods and image augmentation methods

[0001] Related applications

[0002] This application claims priority to Chinese patent application filed on February 10, 2025, with application number 202510149160.0 and entitled "Image Enhancement Model Training Method and Image Enhancement Method", the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of image processing technology, and in particular to an image enhancement model training method, apparatus, computer device, computer-readable storage medium, and computer program product, as well as an image enhancement method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Technology

[0004] With the development of computer technology, images are increasingly used in various situations. Images captured in low-light environments often suffer from problems such as unclear details, insufficient contrast, and color distortion. These problems severely affect the visual quality of the image and subsequent visual processing tasks, such as object detection, image recognition, and image understanding. Therefore, enhancing low-light images is an urgent problem to be solved.

[0005] In traditional techniques, the method for enhancing low-light images is to train an image enhancement model using a fully supervised learning-based low-light image enhancement method, and then use the image enhancement model to enhance the low-light image.

[0006] However, traditional methods typically require paired low-light and normal-light images for supervised learning, which can lead to overfitting of the image enhancement model and result in poor image enhancement performance. Summary of the Invention

[0007] Therefore, it is necessary to provide an image enhancement model training method, apparatus, computer device, computer-readable storage medium, and computer program product, as well as an image enhancement method, apparatus, computer device, computer-readable storage medium, and computer program product, to address the aforementioned technical problems.

[0008] In a first aspect, this application provides a method for training an image enhancement model, executed by an electronic device, comprising:

[0009] Acquire multiple sample images and a first image enhancement prompt message for each of the sample images;

[0010] For each sample image, the sample image and the first image enhancement prompt information are input into an initial image enhancement model. The initial image enhancement model uses the first image enhancement prompt information as conditional information to guide image enhancement, and the sample image is enhanced to obtain an enhanced image.

[0011] Based on the first image enhancement prompt information, image restoration prompt information is generated to restore the enhanced image to the sample image. The enhanced image and the image restoration prompt information are input into an initial image restoration model. Through the initial image restoration model, using the image restoration prompt information as conditional information to guide image restoration, the enhanced image is restored to obtain the restored image.

[0012] By comparing the differences between the sample image and the restored image, the sample training loss value of the sample image is determined; and

[0013] Based on the sample training loss values ​​of each of the multiple sample images, the model parameters of the initial image enhancement model are adjusted to obtain a trained image enhancement model.

[0014] Secondly, this application also provides an image enhancement model training device, comprising:

[0015] The sample data acquisition module is used to acquire multiple sample images and first image enhancement prompt information for each of the sample images;

[0016] An image enhancement module is used to input the sample image and the first image enhancement prompt information into an initial image enhancement model for each sample image, and to enhance the sample image using the first image enhancement prompt information as condition information to guide image enhancement, thereby obtaining an enhanced image.

[0017] The image restoration module is used to generate image restoration prompts to restore the enhanced image to the sample image based on the first image enhancement prompts, input the enhanced image and the image restoration prompts into an initial image restoration model, and restore the enhanced image to the sample image through the initial image restoration model, using the image restoration prompts as conditional information to guide image restoration, to obtain the restored image.

[0018] The loss calculation module is used to compare the differences between the sample image and the restored image to determine the sample training loss value of the sample image;

[0019] The parameter adjustment module is used to adjust the model parameters of the initial image enhancement model based on the sample training loss values ​​of the multiple sample images, so as to obtain a trained image enhancement model.

[0020] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described image enhancement model training method.

[0021] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described image enhancement model training method.

[0022] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described image enhancement model training method.

[0023] Sixthly, this application provides an image enhancement method performed by an electronic device, comprising:

[0024] Obtain the image to be enhanced and the second image enhancement prompt information for the image to be enhanced; and

[0025] The image to be enhanced and the second image enhancement prompt information are input into the trained image enhancement model. The trained image enhancement model uses the second image enhancement prompt information as the condition information to guide image enhancement, and performs image enhancement on the image to be enhanced to obtain the enhanced image. The trained image enhancement model is trained by the above-described image enhancement model training method.

[0026] In a seventh aspect, this application also provides an image enhancement apparatus, comprising:

[0027] The data acquisition module is used to acquire the image to be enhanced and the second image enhancement prompt information of the image to be enhanced;

[0028] The processing module is used to input the image to be enhanced and the second image enhancement prompt information into a trained image enhancement model, and to perform image enhancement on the image to be enhanced by the trained image enhancement model, using the second image enhancement prompt information as the condition information for guiding image enhancement, to obtain an enhanced image; the trained image enhancement model is trained by the above-described image enhancement model training method.

[0029] Eighthly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described image enhancement method.

[0030] Ninthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described image enhancement method.

[0031] In a tenth aspect, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described image enhancement method.

[0032] Details of one or more embodiments of this application are set forth in the following drawings and description. Other features, objects, and advantages of this application will become apparent from the specification, drawings, and claims. Attached Figure Description

[0033] To more clearly illustrate the technical solutions in the embodiments of this application or the conventional technology, the drawings used in the description of the embodiments or the conventional technology will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the disclosed drawings without creative effort.

[0034] Figure 1 shows the application environment of an image enhancement model training method in one embodiment;

[0035] Figure 2 is a flowchart illustrating an image enhancement model training method in one embodiment;

[0036] Figure 3 is a schematic diagram of obtaining an enhanced image through an initial image enhancement model in one embodiment;

[0037] Figure 4 is a schematic diagram of obtaining the first enhanced image features in one embodiment;

[0038] Figure 5 is a schematic diagram of cross-attention interaction in one embodiment;

[0039] Figure 6 is a schematic diagram of cross-attention interaction in another embodiment;

[0040] Figure 7 is a schematic diagram of obtaining the restored image through the initial image restoration model in one embodiment;

[0041] Figure 8 is a schematic diagram of the content consistency loss value obtained in one embodiment;

[0042] Figure 9 is a schematic diagram of the supervised loss value obtained in one embodiment;

[0043] Figure 10 is a schematic diagram of the image enhancement effect in one embodiment;

[0044] Figure 11 is a flowchart illustrating an image enhancement method in one embodiment;

[0045] Figure 12 is a model framework diagram for training a dark light enhancement model in one embodiment;

[0046] Figure 13 is a schematic diagram of multi-scale feature extraction in one embodiment;

[0047] Figure 14 is a schematic diagram of cross-attention interaction in a loop attention adapter in one embodiment;

[0048] Figure 15 is a schematic diagram of visual quality comparison of a dark light enhancement method in one embodiment;

[0049] Figure 16 is a schematic diagram of the comparison of the dark light enhancement method in a dark light face detection task in one embodiment;

[0050] Figure 17 is a schematic diagram of the low-light enhancement method in nighttime image segmentation and comparison in one embodiment;

[0051] Figure 18 is a schematic diagram of the low-light enhancement method for nighttime image segmentation and comparison in another embodiment;

[0052] Figure 19 is a structural block diagram of an image enhancement model training device in one embodiment;

[0053] Figure 20 is a structural block diagram of an image enhancement device in one embodiment;

[0054] Figure 21 is an internal structure diagram of a computer device in one embodiment. Detailed Implementation

[0055] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0056] In order to clearly describe the technical solution of this application and facilitate understanding of the technical solution of this application, the key concepts involved in this application will be explained below.

[0057] 1. Sample image.

[0058] Sample images refer to image data used to train image enhancement models. For example, sample images can be images of poor quality that require enhancement. Specifically, sample images can be images that appear dark and have low contrast due to insufficient ambient light, i.e., low-light images. As another example, sample images can be images with color deviations due to incorrect white balance settings or uneven lighting conditions, i.e., color-unbalanced images.

[0059] 2. First image enhancement prompt information.

[0060] The first image enhancement prompt is a prompt used to guide the enhancement of sample images. It can help to better understand the needs and goals of image enhancement, thereby improving the quality of sample images more effectively.

[0061] 3. Image enhancement.

[0062] Image enhancement is an image processing technique designed to adjust the visual appearance of an image to better suit specific applications or improve human viewing experience. The goal of image enhancement is to improve image quality, highlighting useful information while suppressing noise or unimportant details. For example, image enhancement can be used to increase the brightness and contrast of low-light images to obtain images with normal lighting. It can also be used to correct color deviations in color-unbalanced images to obtain color-balanced images.

[0063] 4. Image restoration prompts.

[0064] Image restoration prompts are information used to guide the restoration of an enhanced image back to a sample image. They help to better understand the needs and goals of image restoration, thereby achieving image restoration more effectively.

[0065] 5. Image restoration.

[0066] Image restoration refers to the process of processing an enhanced image to eliminate its enhancement effects and restore it to the original image. The goal of image restoration is to recover the original state of the sample image as closely as possible. For example, if the sample image is a low-light image, image restoration can be used to reduce the brightness and contrast of a normal-light image to restore its original state. Similarly, if the sample image is a color-imbalanced image, image restoration can be used to correct color deviations and restore the original state of the color-imbalanced image.

[0067] In an exemplary embodiment, the image enhancement model training method provided in this application, as shown in Figure 1, can be applied to a terminal. The terminal acquires multiple sample images and first image enhancement prompts for each sample image. For each sample image, the sample image and the first image enhancement prompts are input into an initial image enhancement model. Using the initial image enhancement model and the first image enhancement prompts as guiding conditions for image enhancement, the sample image is enhanced to obtain an enhanced image. Based on the first image enhancement prompts, image restoration prompts are generated to restore the enhanced image to the sample image. The enhanced image and the image restoration prompts are input into an initial image restoration model. Using the initial image restoration model and the image restoration prompts as guiding conditions for image restoration, the enhanced image is restored to obtain a restored image. The differences between the sample image and the restored image are compared to determine the sample training loss value of the sample image. Based on the sample training loss values ​​of the multiple sample images, the model parameters of the initial image enhancement model are adjusted to obtain a trained image enhancement model.

[0068] The terminals can be, but are not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle systems, and projection devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted displays. Head-mounted displays can include virtual reality (VR) devices, augmented reality (AR) devices, and smart glasses.

[0069] In an exemplary embodiment, as shown in Figure 2, an image enhancement model training method is provided. This embodiment illustrates the method by applying it to a terminal. It is understood that this method can also be applied to a server, or to a system including both a terminal and a server, and is implemented through interaction between the terminal and the server. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. In this embodiment, the method includes steps 202 to 210. Wherein:

[0070] Step 202: Obtain multiple sample images and the first image enhancement prompt information for each sample image.

[0071] Here, "sample image" refers to the image data used to train the image enhancement model. For example, a sample image can be a low-quality image that requires enhancement. Specifically, a sample image can be a dark, low-contrast image due to insufficient ambient light, i.e., a low-light image. Another example is a color-unbalanced image, which may exhibit color deviation due to incorrect white balance settings or uneven lighting conditions.

[0072] The first image enhancement prompt refers to the prompt information used to guide image enhancement of the sample image. It helps to better understand the needs and goals of image enhancement, thereby more effectively improving the quality of the sample image. For example, the first image enhancement prompt can be a first text prompt used to guide image enhancement of the sample image, indicating a textual description of the needs and goals of image enhancement. Alternatively, the first image enhancement prompt can also be a first image prompt used to guide image enhancement of the sample image, indicating an image description of the needs and goals of image enhancement.

[0073] For example, the terminal acquires multiple sample images and first image enhancement prompts for each sample image. In a specific application, the first image enhancement prompts include first text prompts and first image prompts guiding image enhancement of the sample images. The first text prompts can be text descriptions generated based on the image enhancement requirements and objectives, and the first image prompts can be obtained based on image feature information extracted from the sample images.

[0074] In a specific application, the first text prompt can be input by the user based on their image enhancement needs and goals, or it can be automatically generated by the terminal based on those needs and goals. For example, by analyzing a sample image, the terminal can determine its image enhancement needs and goals, and then generate the first text prompt accordingly. For instance, if the sample image is a low-light image, the terminal can determine its image enhancement needs and goals by analyzing the low-light image. To adjust the low-light image to a normal-light image through image enhancement, the normal-light image can be used as the first text prompt.

[0075] In a specific application, the first image prompt can be obtained by the terminal based on image feature information extracted from the sample image. For example, the first image prompt can specifically be image feature information characterizing illumination perception, which reflects the illumination conditions of the enhanced image. This helps to better understand the illumination environment that needs enhancement, thereby enabling targeted optimization. For instance, the first image prompt can specifically be the inverse brightness channel map, illumination distribution map, or illumination feature map generated by a pre-trained illumination perception network from the sample image.

[0076] Step 204: For each sample image, input the sample image and the first image enhancement prompt information into the initial image enhancement model. Through the initial image enhancement model, using the first image enhancement prompt information as the condition information to guide image enhancement, perform image enhancement on the sample image to obtain the enhanced image.

[0077] The initial image augmentation model refers to the image augmentation model before parameter tuning; it is the basic version of the image augmentation model and is typically used in subsequent training, tuning, or optimization processes. For example, the parameters of the initial image augmentation model are pre-set, either based on a general image augmentation task or randomly initialized, and have not yet been trained or optimized for a specific dataset or task. In this embodiment, the initial image augmentation model can be optimized by training on multiple sample images to obtain a trained image augmentation model.

[0078] Image enhancement, a type of image processing technique, aims to adjust the visual effects of an image to make it more suitable for specific applications or easier for humans to observe. The goal of image enhancement is to improve image quality, highlight useful information, and suppress noise or unimportant details. For example, image enhancement can be used to increase the contrast of a low-light image to obtain a normal-light image. It can also be used to correct color deviations in a color-unbalanced image to obtain a color-balanced image.

[0079] For example, for each sample image, the terminal inputs the sample image and the first image enhancement prompt information into the initial image enhancement model. Through the initial image enhancement model, using the first image enhancement prompt information as the condition information to guide image enhancement, the sample image is enhanced to obtain an enhanced image.

[0080] In specific applications, the initial image enhancement model includes a first prompt information encoder and a first generator for image enhancement. By inputting the first image enhancement prompt information into the first prompt information encoder, the first image enhancement prompt information can be encoded to obtain a first encoded feature. By inputting the first encoded feature and a sample image into the first generator, the first generator can perform image enhancement processing on the sample image using the first encoded feature as a control condition to obtain an enhanced image.

[0081] In a specific application, the first generator includes an image enhancement encoder, a multi-scale enhancement feature extraction network, and an image enhancement decoder. The image enhancement encoder is used to encode the sample image to obtain a second encoded feature. The multi-scale enhancement feature extraction network is used to perform multi-scale feature extraction on the second encoded feature based on the first encoded feature to obtain a first enhanced image feature. The image enhancement decoder is used to decode the first enhanced image feature to obtain an enhanced image.

[0082] Step 206: Based on the first image enhancement prompt information, generate image restoration prompt information to restore the enhanced image to the sample image. Input the enhanced image and the image restoration prompt information into the initial image restoration model. Through the initial image restoration model, using the image restoration prompt information as the condition information to guide image restoration, perform image restoration on the enhanced image to obtain the restored image.

[0083] Image restoration prompts refer to information used to guide the restoration of an enhanced image back to the original sample image. They help to better understand the needs and goals of image restoration, thus achieving more effective restoration. For example, image restoration prompts can specifically be second textual prompts that indicate the needs and goals of image restoration. Alternatively, image restoration prompts can be second image prompts that provide the semantic features of the sample image, helping to better understand the sample image and maintain semantic consistency during the restoration process.

[0084] The initial image restoration model refers to the image restoration model before parameter adjustments; it is the basic version of the image restoration model and is typically used in subsequent training, adjustment, or optimization processes. The parameters of the initial image restoration model are pre-set, possibly based on general image restoration task settings or randomly initialized, and have not yet been trained or optimized for a specific dataset or task. In this embodiment, the initial image restoration model is mainly used to restore the enhanced image output by the initial image enhancement model to a sample image, so as to use the restored image for unsupervised image enhancement model training.

[0085] Image restoration refers to the process of processing an enhanced image to eliminate its enhancement effects and restore it to the original image. The goal of image restoration is to recover the original state of the sample image as closely as possible. For example, if the sample image is a low-light image, image restoration can be used to reduce the brightness and contrast of a normal-light image to restore its original state. Similarly, if the sample image is a color-imbalanced image, image restoration can be used to correct color deviations and restore the original state of the color-imbalanced image.

[0086] For example, the terminal generates image restoration prompts based on the first image enhancement prompts, which restore the enhanced image to the sample image. The enhanced image and the image restoration prompts are input into the initial image restoration model. The initial image restoration model uses the image restoration prompts as conditional information to guide the image restoration, and performs image restoration on the enhanced image to obtain the restored image.

[0087] In specific applications, the terminal can determine the image enhancement direction based on the first image enhancement prompt information, and then use the opposite direction of the image enhancement direction as the image restoration direction. The image restoration prompt information is generated by using the image restoration direction to restore the enhanced image to the sample image. That is, the image restoration prompt information can be obtained by inverting the first image enhancement prompt information.

[0088] In a specific application, taking the first image enhancement prompt information, which includes a first text prompt information and a first image prompt information for guiding the image enhancement of the sample image, as an example, the terminal can generate a second text prompt information with opposite semantics by analyzing the text semantics of the first text prompt information, or the user can manually input the second text prompt information. Simultaneously, the terminal can obtain the second image prompt information by inverting the first image prompt information.

[0089] In specific applications, the initial image restoration model includes a second prompt information encoder and a second generator for image restoration. By inputting the image restoration prompt information into the second prompt information encoder, the image restoration prompt information can be encoded to obtain a third encoded feature. By inputting the third encoded feature and the enhanced image into the second generator, the enhanced image can be processed by the second generator with the third encoded feature as the control condition to obtain the restored image.

[0090] In a specific application, the second generator includes an image restoration encoder, a multi-scale restoration feature extraction network, and an image restoration decoder. The image restoration encoder is used to encode the restored image to obtain a fourth encoded feature. The multi-scale restoration feature extraction network is used to extract multi-scale features from the fourth encoded feature based on the third encoded feature to obtain the restored image features. The image restoration decoder is used to decode the restored image features to obtain the restored image.

[0091] Step 208: Compare the differences between the sample image and the restored image to determine the sample training loss value of the sample image.

[0092] For example, by comparing the sample image and the restored image, the terminal can determine the differences between them, and then use these differences to determine the sample training loss value of the sample image. In specific applications, the differences between the sample image and the restored image can be obtained by comparing the pixel values ​​at each pixel location in the sample image and the restored image. The terminal can then determine the sample training loss value of the sample image by calculating the difference between the pixel values ​​at each pixel location in the sample image and the restored image.

[0093] In a specific application, the sample training loss value can be L1 loss. The terminal will calculate the absolute value of the difference between the pixel values ​​at each pixel position in the sample image and the restored image, and then sum the absolute differences of all pixel positions and divide by the total number of pixels in the sample image to obtain the mean absolute error as the sample training loss value.

[0094] Step 210: Based on the sample training loss values ​​of each of the multiple sample images, adjust the model parameters of the initial image enhancement model to obtain the trained image enhancement model.

[0095] For example, after obtaining the sample training loss values ​​of each of the multiple sample images, the terminal will adjust the model parameters of the initial image enhancement model based on the sample training loss values ​​of each of the multiple sample images to obtain the trained image enhancement model.

[0096] In practical applications, after obtaining the individual training loss values ​​of multiple sample images, the terminal can calculate the total loss value of the model during training based on these individual loss values. Based on this total loss value, the gradient of each parameter in the initial image enhancement model is calculated using the backpropagation algorithm. Then, an optimization algorithm is used to update the parameters based on the gradients, thereby adjusting the model parameters of the initial image enhancement model. It can be understood that the total loss value of the model during training can be obtained by superimposing the individual training sample loss values ​​of multiple sample images, or by further calculating other loss values ​​and combining them with the individual training sample loss values ​​of multiple sample images and other loss values.

[0097] Specifically, the gradient represents the contribution of a parameter to the model's loss. A positive gradient indicates that increasing the parameter will increase the loss, and the parameter should be decreased. Conversely, a negative gradient indicates that increasing the parameter will decrease the loss, and the parameter should be increased. The magnitude of the gradient reflects the urgency of parameter adjustment. The optimization algorithm can be a stochastic gradient descent algorithm, which controls parameter updates using a predefined learning rate. Alternatively, an adaptive optimizer can control parameter updates by calculating the first-order moment estimate (momentum) and the second-order moment estimate (adaptive learning rate).

[0098] The aforementioned image enhancement model training method, upon acquiring multiple sample images and first image enhancement prompts for each sample image, inputs the sample image and the first image enhancement prompts into an initial image enhancement model for each sample image. This allows the initial image enhancement model to enhance the sample image using the first image enhancement prompts as guiding conditions, resulting in an enhanced image. Based on the first image enhancement prompts, image restoration prompts are generated to restore the enhanced image back to the sample image. Inputting the enhanced image and the image restoration prompts into an initial image restoration model allows the initial image restoration model to restore the enhanced image using the image restoration prompts as guiding conditions, resulting in a restored image. By comparing the differences between the sample image and the restored image, the sample training loss value of the sample image can be determined. Furthermore, based on the sample training loss values ​​of multiple sample images, the model parameters of the initial image enhancement model are adjusted to obtain a trained image enhancement model. The entire process defines a loop generation process that includes image enhancement and image restoration, enabling unsupervised training of the image enhancement model. This avoids the problem of overfitting in the image enhancement model, thus obtaining an image enhancement model that can support improved image enhancement effects.

[0099] In an exemplary embodiment, the initial image enhancement model includes a first cue information encoder, an image enhancement encoder, a multi-scale enhancement feature extraction network, and an image enhancement decoder. Using the initial image enhancement model, and with the first image enhancement cue information as conditional information to guide image enhancement, the sample image is enhanced to obtain an enhanced image, including:

[0100] The first image enhancement prompt information is encoded by the first prompt information encoder to obtain the first encoded feature, and the sample image is encoded by the image enhancement encoder to obtain the second encoded feature;

[0101] By using a multi-scale enhanced feature extraction network, multi-scale feature extraction is performed on the second encoded feature based on the first encoded feature to obtain the first enhanced image feature;

[0102] The enhanced image is obtained by decoding the features of the first enhanced image using an image enhancement decoder.

[0103] The system comprises the following components: a first prompt information encoder, which compresses the input first image enhancement prompt information into a low-dimensional feature representation, i.e., the first encoded feature; an image enhancement encoder, which compresses the input sample image into a low-dimensional feature representation, i.e., the second encoded feature; a multi-scale enhancement feature extraction network, which is a deep learning network combining multi-scale feature extraction techniques, aiming to enhance the model's understanding of the details and global picture of the sample image through feature fusion at different scales; and an image enhancement decoder, which recovers the enhanced image from the low-dimensional first enhanced image features.

[0104] For example, the initial image enhancement model includes a first prompt information encoder, an image enhancement encoder, a multi-scale enhancement feature extraction network, and an image enhancement decoder, as shown in Figure 3. When enhancing a sample image using the initial image enhancement model, the terminal first encodes the first image enhancement prompt information using the first prompt information encoder to obtain the first encoded feature, and then encodes the sample image using the image enhancement encoder to obtain the second encoded feature. Next, the multi-scale enhancement feature extraction network extracts multi-scale features from the second encoded feature based on the first encoded feature to obtain the first enhanced image feature. Finally, the image enhancement decoder decodes the first enhanced image feature to obtain the enhanced image.

[0105] In specific applications, the type of the first prompt information encoder is determined based on the form of the first image enhancement prompt information. For example, if the first image enhancement prompt information is a first text prompt information, then the first prompt information encoder is an encoder used to extract text features; if the first image enhancement prompt information is a first image prompt information, then the first prompt information encoder can be an encoder used to extract image features.

[0106] In specific applications, an image augmentation encoder, specifically an encoder in a variational autoencoder, is used to map a sample image to a low-dimensional latent space and learn a compressed representation of the sample image, while an image augmentation decoder, specifically a decoder in a variational autoencoder, is used to generate an augmented image from the latent space.

[0107] In this embodiment, by encoding the first image enhancement prompt information, the amount of data can be compressed, and the first image enhancement prompt information can be represented with a first encoding feature that has a smaller amount of data. By encoding the sample image, the amount of data can be compressed, and the sample image can be represented with a second encoding feature that has a smaller amount of data. Then, by performing multi-scale feature extraction based on the first encoding feature and the second encoding feature, image enhancement can be performed with the first encoding feature as the control condition to obtain the first enhanced image feature. Thus, the enhanced image can be obtained by decoding the first enhanced image feature.

[0108] In an exemplary embodiment, the first image enhancement prompt information includes first text prompt information and first image prompt information; the first image enhancement prompt information is encoded by a first prompt information encoder to obtain a first encoded feature, including:

[0109] The first text prompt information is encoded by the first text encoder in the first prompt information encoder to obtain the first text encoding feature, and the first image prompt information is encoded by the first image encoder in the first prompt information encoder to obtain the first image encoding feature.

[0110] The first text encoding feature and the first image encoding feature are used as the first encoding feature;

[0111] By using a multi-scale enhanced feature extraction network, multi-scale feature extraction is performed on the second encoded feature based on the first encoded feature to obtain the first enhanced image feature, including:

[0112] By using a multi-scale enhanced feature extraction network, based on the first text encoding features and the first image encoding features, multi-scale feature extraction is performed on the second encoding features to obtain the first enhanced image features.

[0113] For example, when the first image enhancement prompt information includes both first text prompt information and first image prompt information, the terminal encodes the first text prompt information using the first text encoder in the first prompt encoder to obtain first text encoding features, and encodes the first image prompt information using the first image encoder in the first prompt information encoder to obtain first image encoding features. The first text encoding features and the first image encoding features are then used as the first encoding features. When the first encoding features include both first text encoding features and first image encoding features, the method of multi-scale feature extraction based on the first encoding features using a multi-scale enhancement feature extraction network is as follows: Multi-scale feature extraction is performed on the second encoding features using the first text encoding features and the first image encoding features to obtain the first enhanced image features.

[0114] In specific applications, the first text encoder and the first image encoder features can be a pre-trained CLIP (Contrastive Language-Image Pre-training) model, which is a multimodal pre-trained model that includes a text encoder for text feature extraction and an image encoder for image feature extraction.

[0115] In this embodiment, when the first image enhancement prompt information includes first text prompt information and first image prompt information, the first text prompt information and the first image prompt information can be represented by encoding them. Then, using the first text encoding features and the first image encoding features as control conditions, multi-scale feature extraction can be performed on the second encoding features to obtain the first enhanced image features. It is understood that by simultaneously providing text features and image features as control conditions, sufficient semantic information can be provided during the multi-scale feature extraction process, which is beneficial for maintaining the semantic consistency of image enhancement.

[0116] In one exemplary embodiment, the image enhancement model training method further includes:

[0117] Extract image information representing illumination perception from the sample image, and invert the pixel value of each pixel in the image information representing illumination perception to obtain the first image prompt information.

[0118] Image information characterizing illumination perception refers to image information that can directly reflect the intensity and distribution of light in a sample image. For example, image information characterizing illumination perception can specifically represent single-channel images of illumination perception (such as brightness channel images) and illumination estimation maps.

[0119] For example, the terminal extracts image information representing illumination perception from the sample image, inverts the pixel value of each pixel in the image information representing illumination perception, and obtains the first image prompt information. In specific applications, the image information representing illumination perception can be a single-channel image representing illumination perception, such as a luminance channel image.

[0120] In a specific application, when extracting the luminance channel image, the terminal first determines whether the color space type of the sample image is a color space that includes the luminance channel image. If the color space type of the sample image is a color space that includes the luminance channel image, the terminal can directly extract the luminance channel image from the sample image. If the color space type of the sample image is not a color space that includes the luminance channel image, the terminal needs to first convert the sample image to obtain a converted image in a color space that includes the luminance channel image, and then extract the luminance channel image from the converted image.

[0121] For example, taking a luminance channel image as an example, which is a V (Value, luminance) channel image, the color space containing the luminance channel image is the HSV (Hue, Saturation, Value) color space. If the sample image is an HSV image, the terminal can directly extract the V channel image from the sample image. If the sample image is not an HSV image (specifically, it can be an RGB (Red, Green, Blue) image, etc.), then it is necessary to first convert the sample image to an HSV image and then extract the V channel image as the first image prompt information.

[0122] In this embodiment, by inverting the pixel value of each pixel in the image information representing illumination perception, the first image prompt information is obtained. This can provide illumination perception semantic information as guidance during image enhancement, helping the initial image enhancement model to better understand and process the illumination conditions in the sample image, thereby achieving a more natural and effective enhancement effect and improving the image enhancement effect.

[0123] In an exemplary embodiment, the multi-scale enhanced feature extraction network includes N-level first encoding layers, N-level first decoding layers, and a first bridging layer; N is a positive integer greater than 1; through the multi-scale enhanced feature extraction network, based on the first text encoding features and the first image encoding features, multi-scale feature extraction is performed on the second encoding features to obtain the first enhanced image features, including:

[0124] Through N levels of first coding layers, based on the first text coding features and the first image coding features, multi-scale feature extraction is performed on the second coding features to obtain the enhanced coding features output by the Nth level of first coding layers;

[0125] Through the first bridging layer, based on the first text coding features and the first image coding features, feature extraction is performed on the enhanced coding features output by the first coding layer at the Nth level to obtain the enhanced coding features output by the first bridging layer;

[0126] Through N levels of first decoding layers, based on the first text encoding features and the first image encoding features, multi-scale feature extraction is performed on the enhanced encoding features output by the first bridging layer to obtain the enhanced decoding features output by the Nth level of first decoding layer;

[0127] The enhanced decoding features output by the first decoding layer of the Nth level are used as the first enhanced image features.

[0128] Among them, the N-level first coding layer is used to extract features of sample images and reduce spatial resolution by gradually utilizing the second coding features based on the first text coding features and the first image coding features. The second coding features can be compressed into a low-dimensional feature representation, namely the enhanced coding features output by the N-level first coding layer.

[0129] In this design, the first encoding layer can extract features through downsampling and cross-attention interaction. Downsampling refers to reducing the spatial dimensions (width and height) of the original data while increasing the number of feature channels. Downsampling can be achieved through a combination of convolution and pooling. Cross-attention interaction is a mechanism used to enhance feature fusion and information transfer. In this embodiment, in the first encoding layer, cross-attention interaction can fuse the first text-encoded features and the first image-encoded features with the downsampled features.

[0130] The first bridging layer connects the N-level first coding layer and the N-level first decoding layer, playing a crucial role in bridging the gap between them. It is typically located after the N-level first coding layer and before the first level first decoder, and is responsible for passing the enhanced coding features output by the N-level first coding layer to the first level first decoder.

[0131] Among them, the N-level first decoding layer is used to gradually restore the low-dimensional features extracted by the encoder to their original size based on the first text encoding features and the first image encoding features. It can gradually restore the enhanced encoding features output by the first bridging layer to features of the same size as the second encoding features, that is, the enhanced decoding features output by the N-level first decoding layer.

[0132] The first decoding layer extracts features through upsampling and cross-attention interaction. Upsampling refers to increasing the spatial resolution of the original data from low to high resolution, often used to recover detailed information in an image. In the first decoding layer, cross-attention interaction allows the first text-encoded features and the first image-encoded features to be fused with the upsampled features.

[0133] For example, as shown in Figure 4, the terminal performs multi-scale feature extraction on the second encoding features based on the first text encoding features and the first image encoding features through N-level first encoding layers to obtain the enhanced encoding features output by the Nth-level first encoding layer. Through the first bridging layer, it performs feature extraction on the enhanced encoding features output by the Nth-level first encoding layer based on the first text encoding features and the first image encoding features to obtain the enhanced encoding features output by the first bridging layer. Through the N-level first decoding layer, it performs multi-scale feature extraction on the enhanced encoding features output by the first bridging layer based on the first text encoding features and the first image encoding features to obtain the enhanced decoding features output by the Nth-level first decoding layer. The enhanced decoding features output by the Nth-level first decoding layer are used as the first enhanced image features.

[0134] In practical applications, the multi-scale enhanced feature extraction network in this embodiment can be implemented using the U-Net network. U-Net is a deep learning model with a symmetrical encoder and decoder structure, connected by a bottleneck layer. The U-Net encoder layer includes multiple encoding levels, which can be used to construct the first encoding layer of N levels. The U-Net decoder includes multiple decoding levels, which can be used to construct the first decoding layer of N levels. The bottleneck layer can be used to construct the first bridging layer.

[0135] In this embodiment, based on the first text encoding features and the first image encoding features, multi-scale feature extraction of the second encoding features can be achieved by first encoding and then decoding layer by layer. This can obtain the first enhanced image features that fully integrate text semantics and image semantics. Then, by decoding the first enhanced image features, an enhanced image with good image enhancement effect can be obtained.

[0136] In an exemplary embodiment, through N-level first coding layers, based on first text coding features and first image coding features, multi-scale feature extraction is performed on second coding features to obtain enhanced coding features output by the Nth-level first coding layer, including:

[0137] For the nth level of the first coding layer in N levels, when n equals 1, the second coding feature is used as part of the input data of the nth level of the first coding layer. When n is greater than 1 and less than or equal to N, the enhanced coding feature output by the (n-1)th first coding layer is used as part of the input data of the nth level of the first coding layer.

[0138] In the first encoding layer of the nth level, feature extraction is performed on a portion of the input data of the first encoding layer of the nth level to obtain the first scale feature. Cross-attention interaction is performed on the first scale feature, the first text encoding feature and the first image encoding feature to obtain the enhanced encoding feature output by the first encoding layer of the nth level.

[0139] For example, for the nth level of the first coding layer in N levels, when n equals 1, the terminal uses the second coding feature as part of the input data for the nth level of the first coding layer. When n is greater than 1 and less than or equal to N, the terminal uses the enhanced coding feature output by the (n-1)th first coding layer as part of the input data for the nth level of the first coding layer. In the nth level of the first coding layer, the terminal performs feature extraction on a portion of the input data of the nth level of the first coding layer based on downsampling processing to obtain the first scale feature. Cross-attention interaction is then performed on the first scale feature, the first text coding feature, and the first image coding feature to obtain the enhanced coding feature output by the nth level of the first coding layer.

[0140] In this embodiment, by means of this method, at each encoding level, feature extraction processing can be performed on a portion of the input data based on the first text encoding features and the first image encoding features. Thus, through layer-by-layer feature extraction processing, richer feature representations can be learned, which is beneficial for improving the image enhancement effect.

[0141] In an exemplary embodiment, cross-attention interaction is performed on the first scale feature, the first text encoding feature, and the first image encoding feature to obtain the enhanced encoding feature output by the first encoding layer at the nth level, including:

[0142] Cross-attention interaction is performed on the first scale feature and the first text encoding feature to obtain the first interactive feature, and cross-attention interaction is performed on the first scale feature and the first image encoding feature to obtain the second interactive feature;

[0143] By concatenating the first and second interactive features, the enhanced coding features output by the first coding layer of the nth level are obtained.

[0144] For example, in order to fully utilize the illumination-aware semantic features of the first image encoding features and promote the semantic representation learning of latent features, the terminal will perform cyclic cross-attention interaction on the first scale features, the first text encoding features, and the first image encoding features to obtain the enhanced encoding features output by the nth level first encoding layer. Here, cyclic cross-attention interaction refers to querying the first image encoding features through one cross-attention layer and providing the feature response to the first image encoding features through another cross-attention layer.

[0145] For example, as shown in Figure 5, in the first attention stage, the terminal performs cross-attention interaction on the first scale feature and the first text encoding feature to obtain the first interaction feature, and performs cross-attention interaction on the first scale feature and the first image encoding feature to obtain the third interaction feature. In the second attention stage, the terminal feeds back the first interaction feature and the third interaction feature to the first image prompt feature. After performing cross-attention interaction on the first scale feature and the third interaction feature to obtain the second interaction feature, the terminal obtains the enhanced encoding feature output by the first encoding layer of the nth level by concatenating the first interaction feature and the second interaction feature.

[0146] In practical applications, when performing cross-attention interaction on the first scale feature and the first text encoding feature, the terminal can use the first scale feature as the query vector and the first text encoding feature as the key vector and value vector. The first interaction feature is obtained by performing cross-attention calculation on the query vector, key vector and value vector through the cross-attention layer.

[0147] In this embodiment, by means of the interaction of features of the first scale feature, the first text encoding feature and the first image encoding feature, the first scale feature and the first text encoding feature can guide the first image encoding feature, so as to learn a richer feature representation, maintain semantic consistency, and help to improve the image enhancement effect.

[0148] In an exemplary embodiment, a second interactive feature is obtained by performing cross-attention interaction on the first scale feature and the first image coding feature, including:

[0149] A linear mapping is performed on the first image encoding features to obtain a first query vector, and the first scale features are used as the first key vector and the first value vector;

[0150] A third interactive feature is obtained by performing cross-attention interaction on the first query vector, the first key vector, and the first value vector.

[0151] The third interactive feature is used as the second query vector, and the first image encoding feature is linearly mapped to obtain the second key vector and the second value vector.

[0152] The second query vector, the second key vector, and the second value vector are subjected to cross-attention interaction to obtain the second interaction feature.

[0153] For example, as shown in Figure 6, when performing cross-attention interaction on the first scale feature and the first image coding feature, the terminal performs linear mapping on the first image coding feature to obtain the first query vector, and uses the first scale feature as the first key vector and the first value vector. Cross-attention interaction is then performed on the first query vector, the first key vector, and the first value vector to obtain the third interaction feature. The third interaction feature is then used as the second query vector, and linear mapping is performed on the first image coding feature to obtain the second key vector and the second value vector. Cross-attention interaction is then performed on the second query vector, the second key vector, and the second value vector to obtain the second interaction feature.

[0154] In specific applications, the third interaction feature Z i It can be defined by the following equation:

[0155] Q i =c i W′ q , is the first query vector, c i It is the first image coding feature, W′ q It is a learnable linear projection layer used to linearly map the encoded features of the first image, K. u =Z u Z is the first key vector. u It is the first-scale feature, V u =Z u , is the first value vector, d is the dimension of the first key vector, and the softmax function converts the attention scores into a probability distribution, representing the attention weight of the query vector to each key vector.

[0156] In practical applications, the second interaction feature can be defined by the following equation:

[0157] Z i K is the second query vector. i =c i W′ k , is the second key vector, c i It is the first image coding feature, W′ k A learnable linear projection layer, used to linearly map the encoded features of the first image, V i =c i W′ v , is the second value vector, where W′ v The learnable linear projection layer is used to linearly map the encoded features of the first image. d is the dimension of the second key vector. The softmax function converts the attention scores into a probability distribution, representing the attention weight of the query vector on each key vector.

[0158] In this embodiment, this approach can fully utilize the illumination-aware semantic features of the first image encoding features and promote the semantic representation learning of latent features, thereby learning richer feature representations, maintaining semantic consistency, and improving image enhancement effects.

[0159] In an exemplary embodiment, through N-level first decoding layers, based on first text encoding features and first image encoding features, multi-scale feature extraction is performed on the enhanced encoding features output by the first bridging layer to obtain the enhanced decoding features output by the Nth-level first decoding layer, including:

[0160] For the nth level of the first decoding layer in N levels, when n equals 1, the enhanced coding features output by the first bridging layer are used as part of the input data of the nth first decoder. When n is greater than 1 and less than or equal to N, the enhanced decoding features output by the (n-1)th first decoding layer are used as part of the input data of the nth level of the first decoding layer.

[0161] In the first decoding layer of the nth level, feature extraction is performed on a portion of the input data of the first decoding layer of the nth level to obtain the second scale feature. The second scale feature and the enhanced coding feature output by the first coding layer corresponding to the first decoding layer of the nth level are fused to obtain the fused feature. Cross-attention interaction is performed on the fused feature, the first text coding feature and the first image coding feature to obtain the enhanced decoding feature output by the first decoding layer of the nth level.

[0162] For example, for the nth level of the first decoding layer in N levels, when n equals 1, the terminal uses the enhanced coding features output by the first bridging layer as part of the input data of the nth first decoder. When n is greater than 1 and less than or equal to N, the terminal uses the enhanced decoding features output by the (n-1)th first decoding layer as part of the input data of the nth level of the first decoding layer. In the nth level of the first decoding layer, the terminal performs feature extraction on a portion of the input data of the nth level of the first decoding layer based on upsampling processing to obtain second-scale features. It then fuses the second-scale features with the enhanced coding features output by the first coding layer corresponding to the nth level of the first decoding layer to obtain fused features. Finally, it performs cross-attention interaction on the fused features, the first text coding features, and the first image coding features to obtain the enhanced decoding features output by the nth level of the first decoding layer.

[0163] In practical applications, the first coding layer corresponding to the first decoding layer of the nth level refers to the coding layer whose output enhanced coding features and second-scale features have the same feature size. By fusing the second-scale features with enhanced coding features of the same size, a fused feature can be obtained.

[0164] In specific applications, the terminal first performs cross-attention interaction on the fused feature and the first text encoding feature to obtain the fourth interactive feature, and then performs cross-attention interaction on the fused feature and the first image encoding feature to obtain the fifth interactive feature. Finally, the fourth interactive feature and the fifth interactive feature are concatenated to obtain the enhanced decoding feature output by the first decoding layer of the nth level.

[0165] In practical applications, when performing cross-attention interaction on the fused feature and the first text encoding feature, the terminal can use the fused feature as the query vector and the first text encoding feature as the key vector and value vector. The cross-attention layer is used to perform cross-attention calculation on the query vector, key vector and value vector to obtain the fourth interaction feature.

[0166] In specific applications, the method of cross-attention interaction between the fused feature and the first image coding feature can be the same as the method of cross-attention interaction between the first scale feature and the first image coding feature. The specific interaction process can be as follows: perform linear mapping on the first image coding feature to obtain the query vector, and use the fused feature as the key vector and value vector. Perform cross-attention interaction on the query vector, key vector, and value vector to obtain the interaction feature. Use the interaction feature as the new query vector, and perform linear mapping on the first image coding feature to obtain the new key vector and new value vector. Perform cross-attention interaction on the new query vector, new key vector, and new value vector to obtain the fifth interaction feature.

[0167] In this embodiment, by means of this method, at each decoding level, feature extraction processing can be performed on a portion of the input data based on the first text encoding features and the first image encoding features. Thus, through layer-by-layer feature extraction processing, richer feature representations can be learned, which is beneficial to improving the image enhancement effect.

[0168] In an exemplary embodiment, the first image enhancement prompt information includes first text prompt information and first image prompt information. Based on the first image enhancement prompt information, image restoration prompt information is generated to restore the enhanced image to the sample image, including:

[0169] Based on the first text prompt information, a second text prompt information is generated to restore the enhanced image to the sample image, and the pixel value of each pixel in the first image prompt information is reversed to obtain the second image prompt information;

[0170] Use the second text prompt and the second image prompt as image restoration prompts.

[0171] For example, when the first image enhancement prompt includes a first text prompt and a second text prompt, the terminal can analyze the semantics of the first text prompt and generate a second text prompt with the opposite semantics that restores the enhanced image to the sample image. Alternatively, the user can manually input the second text prompt. Simultaneously, the terminal inverts the pixel values ​​of each pixel in the first image prompt to obtain the second image prompt. After obtaining the second text prompt and the second image prompt, the terminal uses both as image restoration prompts.

[0172] In specific applications, if the second text prompt is automatically generated by the terminal, the terminal can use a pre-trained text question-answering model to obtain the second text prompt. The specific data input to the text question-answering model can be: Please output information that is the opposite of the semantics of the first text prompt.

[0173] In practical applications, inverting pixel values ​​refers to changing the pixel value from its original value to its maximum value minus the current value; this can also be called reverse processing or color inversion. In a specific application, if the first image prompt is obtained by inverting the pixel value of each pixel in the image information representing illumination perception extracted from the sample image, the terminal can directly use the image information representing illumination perception extracted from the sample image as the second image prompt. For example, if the image information representing illumination perception extracted from the sample image is a V-channel image, then the V-channel image is the second image prompt, and the first image prompt can be obtained using the formula: "First image prompt = 1 - V-channel image".

[0174] In this embodiment, by analyzing the first text prompt information and the first image prompt information, accurate generation of image restoration prompt information for restoring the enhanced image to the sample image can be achieved, which is beneficial to improving the image restoration effect.

[0175] In an exemplary embodiment, the initial image restoration model includes a second cue information encoder, an image restoration encoder, a multi-scale restoration feature extraction network, and an image restoration decoder; the initial image restoration model is used to restore the enhanced image based on the image restoration cue information to obtain the restored image, including:

[0176] The image restoration prompt information is encoded by the second prompt information encoder to obtain the third encoding feature, and the enhanced image is encoded by the image restoration encoder in the initial image restoration model to obtain the fourth encoding feature;

[0177] By using a multi-scale restoration feature extraction network, multi-scale feature extraction is performed on the fourth coding feature based on the third coding feature to obtain the restored image features;

[0178] The restored image is obtained by decoding the features of the restored image using an image restoration decoder.

[0179] The system comprises the following components: a second prompt information encoder, which compresses the input image restoration prompt information into a low-dimensional feature representation (i.e., the third encoded feature); an image restoration encoder, which compresses the input enhanced image into a low-dimensional feature representation (i.e., the fourth encoded feature); a multi-scale restoration feature extraction network, which combines multi-scale feature extraction techniques to enhance the model's understanding of the details and global picture of the enhanced image through feature fusion at different scales; and an image restoration decoder, which recovers the restored image from the low-dimensional restored image features.

[0180] For example, the initial image restoration model includes a second prompt information encoder, an image restoration encoder, a multi-scale restoration feature extraction network, and an image restoration decoder, as shown in Figure 7. When restoring the enhanced image through the initial image restoration model, the terminal first encodes the image restoration prompt information through the second prompt information encoder to obtain the third encoded feature, and then encodes the enhanced image through the image restoration encoder in the initial image restoration model to obtain the fourth encoded feature. Then, the multi-scale restoration feature extraction network extracts multi-scale features based on the third encoded feature to obtain the restored image features. Finally, the image restoration decoder decodes the restored image features to obtain the restored image.

[0181] In specific applications, an image restoration decoder, specifically an encoder in a variational autoencoder, is used to map an enhanced image to a low-dimensional latent space and learn a compressed representation of the enhanced image, while an image restoration decoder, specifically a decoder in a variational autoencoder, is used to generate a restored image from the latent space.

[0182] In specific applications, the image restoration prompt information includes second text prompt information and second image prompt information. The second prompt information is encoded by the second prompt information encoder to obtain the second encoded feature. The second text prompt information is encoded by the second text encoder in the second prompt information encoder to obtain the second text encoded feature. The second image prompt information is encoded by the second image encoder in the second prompt information encoder to obtain the second image encoded feature. The second text encoded feature and the second image encoded feature are used as the third encoded feature.

[0183] In practical applications, the multi-scale restoration feature extraction network includes N levels of second coding layers, N levels of second decoding layers, and a second bridging layer, where N is a positive integer greater than 1. The method for multi-scale feature extraction of the fourth coding feature is as follows: through the N levels of second coding layers, based on the second text coding feature and the second image coding feature, multi-scale feature extraction is performed on the fourth coding feature to obtain the restored coding feature output by the Nth level of second coding layers. Through the second bridging layer, based on the second text coding feature and the second image coding feature, feature extraction is performed on the restored coding feature output by the Nth level of second coding layers to obtain the restored coding feature output by the second bridging layer. Through the N levels of second decoding layers, based on the second text coding feature and the second image coding feature, multi-scale feature extraction is performed on the restored coding feature output by the second bridging layer to obtain the restored decoding feature output by the Nth level of second decoding layers. The restored decoding feature output by the Nth level of second decoding layers is used as the restored image feature.

[0184] It is understandable that the model structure of the initial image restoration model is similar to that of the initial image enhancement model. Therefore, the way in which the multi-scale restoration feature extraction network extracts multi-scale features from the fourth coded feature based on the third coded feature to obtain the restored image features is similar to the way in which the multi-scale enhancement feature extraction network extracts multi-scale features from the second coded feature based on the first coded feature.

[0185] In a specific application, when obtaining the restored coding features output by the Nth level second coding layer, for the nth level second coding layer among the N levels, when n equals 1, the fourth coding feature can be used as part of the input data of the nth level second coding layer. When n is greater than 1 and less than or equal to N, the restored coding features output by the (n-1)th level second coding layer can be used as part of the input data of the nth level second coding layer. In the nth level second coding layer, feature extraction is performed on part of the input data of the nth level second coding layer to obtain the third scale feature. Cross-attention interaction is performed on the third scale feature, the second text coding feature, and the second image coding feature to obtain the restored coding features output by the nth level second coding layer.

[0186] In a specific application, when obtaining the restored decoding features output by the Nth level second decoding layer, for the nth level second decoding layer among the N levels, when n equals 1, the restored coding features output by the second bridging layer can be used as part of the input data of the nth level second decoder. When n is greater than 1 and less than or equal to N, the restored decoding features output by the (n-1)th level second decoding layer are used as part of the input data of the nth level second decoding layer. In the nth level second decoding layer, feature extraction is performed on part of the input data of the nth level second decoding layer to obtain the fourth scale feature. The fourth scale feature and the restored coding features output by the second coding layer corresponding to the nth level second decoding layer are fused to obtain the fused feature. Cross-attention interaction is performed on the fused feature, the second text coding feature and the second image coding feature to obtain the restored decoding features output by the nth level second decoding layer.

[0187] In this embodiment, by encoding the image restoration prompt information, the amount of data can be compressed, and the image restoration prompt information can be represented by the third encoding feature with a smaller amount of data. By encoding the enhanced image, the amount of data can be compressed, and the enhanced image can be represented by the fourth encoding feature with a smaller amount of data. Then, by performing multi-scale feature extraction based on the third encoding feature and the fourth encoding feature, the image restoration can be performed with the third encoding feature as the control condition to obtain the restored image features. Thus, the restored image can be obtained by decoding the restored image features.

[0188] In one exemplary embodiment, the image enhancement model training method further includes:

[0189] The sample image is input into the pre-trained image description model, and the image content of the sample image is described by the pre-trained image description model to obtain the image content description information of the sample image;

[0190] Image content features are obtained by encoding image content description information through a pre-trained text encoder.

[0191] The sample image is input into the image enhancement encoder in the initial image enhancement model. The sample image is then encoded by the image enhancement encoder to obtain the fifth encoded feature.

[0192] The second image enhancement feature is obtained by using the multi-scale enhancement feature extraction network in the initial image enhancement model to extract multi-scale features from the fifth encoded feature based on the image content features.

[0193] By comparing the differences between the enhanced features of the second image and the features of the restored image, the content consistency loss value of the sample image is determined.

[0194] Based on the sample training loss values ​​of multiple sample images, the model parameters of the initial image enhancement model are adjusted to obtain a trained image enhancement model, including:

[0195] Based on the sample training loss value and content consistency loss value of multiple sample images, the model parameters of the initial image enhancement model are adjusted to obtain a trained image enhancement model.

[0196] Among them, a pre-trained image description model refers to a model that has been pre-trained to describe the content of an input image. This model analyzes the input image and outputs descriptive information about its content. Pre-trained image description models can be configured according to specific application scenarios. For example, a specific pre-trained image description model could be a BLIP (Bootstrapping Language-Image Pre-training) model, a multimodal vision-language pre-training model designed to improve the efficiency and performance of joint image and text learning by combining contrastive learning and generative tasks.

[0197] The content consistency loss value is used to evaluate whether the content of an image is consistent. In this embodiment, it is mainly used to determine whether the restored image is consistent with the sample image in terms of content; that is, it optimizes the model's performance by measuring the similarity between the restored image and the sample image. In this embodiment, feature matching loss is mainly used to measure content consistency.

[0198] For example, as shown in Figure 8, the terminal can obtain image content description information of the sample image by inputting the sample image into a pre-trained image description model. By encoding the image content description information through a pre-trained text encoder, image content features can be obtained. Then, by inputting the sample image into the image enhancement encoder in the initial image enhancement model and encoding the sample image through the image enhancement encoder to obtain the fifth encoded feature, the multi-scale enhancement feature extraction network in the initial image enhancement model can perform multi-scale feature extraction on the fifth encoded feature based on the image content features to obtain the second image enhancement feature. Thus, the content consistency loss value of the sample image can be determined by comparing the second image enhancement feature and the restored image feature.

[0199] For example, based on the determined content consistency loss value, the terminal can adjust the model parameters of the initial image enhancement model based on the individual sample training loss values ​​and content consistency loss values ​​of multiple sample images to obtain a trained image enhancement model. Specifically, the terminal can superimpose the individual sample training loss values ​​and content consistency loss values ​​of multiple sample images, calculate the total loss value of the model during training, calculate the gradient of each parameter in the initial image enhancement model using the backpropagation algorithm based on the total loss value, and then use an optimization algorithm to update the parameters based on the gradients, thereby adjusting the model parameters of the initial image enhancement model.

[0200] In practical applications, the content consistency loss value can be obtained by comparing the cosine similarity between the enhanced features of the second image and the features of the restored image. The specific calculation formula is as follows:

[0201] in, U represents cosine similarity. init (E l (I l ),C cap ) represents the second image enhancement feature, I l E represents the sample image. l (I l ) represents the fifth coding feature, C cap U represents image content description information. init () indicates multi-scale feature extraction, U d (E d (I n ),C d ) represents the features of the restored image, I n E represents an enhanced image. d (I n ) represents the fourth coding feature, C d U indicates a prompt for image restoration. d () indicates multi-scale feature extraction.

[0202] In this embodiment, image description consistency is proposed as a method for model training. It utilizes the description of the input sample image to maintain the high-level abstract semantic consistency of the latent features during the iterative generation process, making the model pay more attention to high-level semantic information, thereby improving the image enhancement effect.

[0203] In one exemplary embodiment, the image enhancement model training method further includes:

[0204] The first enhanced image features obtained through the initial image enhancement model are input into the pre-trained first reflection decoder for decoding to obtain the first reflection map, and the restored image features are input into the pre-trained second reflection decoder for decoding to obtain the second reflection map;

[0205] By comparing the differences between the first and second reflection maps, the reflection map consistency loss value is obtained;

[0206] The sample image is input into a pre-trained image decomposition model for decomposition to obtain the third reflection map of the sample image. The difference between the first reflection map and the third reflection map is compared to obtain the reconstruction loss value.

[0207] Based on the reflection map consistency loss value and the reconstruction loss value, the supervision loss value of the sample image is obtained;

[0208] Based on the sample training loss values ​​of multiple sample images, the model parameters of the initial image enhancement model are adjusted to obtain a trained image enhancement model, including:

[0209] Based on the sample training loss and supervision loss values ​​of multiple sample images, the model parameters of the initial image augmentation model are adjusted to obtain a trained image augmentation model.

[0210] The pre-trained first and second reflection decoders are pre-trained models used to output reflection maps based on input data. They analyze the input data to output reflection maps. A reflection map is the portion of the image that remains unchanged under varying lighting conditions. Together with the illumination map, it constitutes the original image and is a crucial component of image decomposition. It primarily reflects the material properties of an object's surface, rather than the influence of lighting conditions. The reflection map is the original image after removing highlights, better preserving the object's color information and material details. The pre-trained first and second reflection decoders can be configured according to the specific application scenario.

[0211] Among them, the pre-trained image decomposition model refers to a model that is pre-trained to output a reflection map (i.e., a third reflection map) of a sample image by decomposing it. It can be configured according to the actual application scenario. For example, the pre-trained image decomposition model can be a RetinexNet model, which can output a reflection map based on the input image through the decomposition network in the RetinexNet model.

[0212] The reflectance consistency loss value is used to evaluate whether the reflectance maps of the enhanced and restored images are consistent. It can be used to determine whether the enhanced and restored images are consistent in the image parts that need to be kept unchanged. The reconstruction loss value is used to evaluate whether the reflectance maps of the enhanced and sample images are consistent. It can be used to determine whether the enhanced and sample images are consistent in the image parts that need to be kept unchanged. The supervised loss value combines the reflectance consistency loss value and the reconstruction loss value to evaluate whether the reflectance maps are consistent throughout the entire loop generation process. It can be used to evaluate whether the input sample image, the enhanced image after image enhancement, and the restored image after image restoration are consistent in the image parts that need to be kept unchanged throughout the entire loop generation process.

[0213] For example, as shown in Figure 9, the terminal inputs the first enhanced image features obtained through the initial image enhancement model into a pre-trained first reflection decoder for decoding to obtain a first reflection map, and inputs the restored image features into a pre-trained second reflection decoder for decoding to obtain a second reflection map. By comparing the differences between the first and second reflection maps, a reflection map consistency loss value is obtained. Simultaneously, the sample image is input into a pre-trained image decomposition model for decomposition to obtain a third reflection map of the sample image. By comparing the differences between the first and third reflection maps, a reconstruction loss value is obtained. Based on the reflection map consistency loss value and the reconstruction loss value, a supervised loss value of the sample image is obtained. Based on the determined supervised loss value of the sample image, the terminal can adjust the model parameters of the initial image enhancement model based on the individual sample training loss value and supervised loss value of multiple sample images to obtain a trained image enhancement model.

[0214] In practical applications, by comparing the differences between the first and second reflection maps, the terminal can calculate the mean squared error loss between the two maps to obtain the reflection map consistency loss value. Similarly, by comparing the differences between the first and third reflection maps, the terminal can calculate the L1 loss between them to obtain the reconstruction loss value.

[0215] In a specific application, the supervision loss value can be the sum of the reflection map consistency loss value and the reconstruction loss value. The supervision loss value can then be calculated using the following formula:

[0216] Z l =U l (E l (I l ),C l ), is the first enhanced image feature, I l For the sample image, C l To enhance the prompt information for the first image, Zd =U d (E d (I n ),C d ), is the feature of the restored image, I ref This represents the third reflection image. It is the mean squared error loss value. It is L1 loss.

[0217] In practical applications, the terminal can overlay the training loss values ​​and supervision loss values ​​of multiple sample images to calculate the total loss value of the model during training. Based on the total loss value, the gradient of each parameter in the initial image enhancement model is calculated through the backpropagation algorithm. Then, the optimization algorithm is used to update the parameters based on the gradient, thereby adjusting the model parameters of the initial image enhancement model.

[0218] In this embodiment, a reflection map consistency is proposed to be applied to model training. By utilizing the reflection map consistency, robust and consistent spatial semantic features can be learned during the iterative generation process, enabling the model to be trained with spatial semantics as a guide, thereby improving the image enhancement effect.

[0219] In one exemplary embodiment, the image enhancement model training method further includes:

[0220] For each sample image, the sample image and the enhanced image are input into the first discriminator. The first discriminator classifies the sample image and the enhanced image to obtain the first classification result.

[0221] The sample image and the restored image are input into the second discriminator. The second discriminator classifies the sample image and the restored image to obtain the second classification result.

[0222] Based on the first classification result and the second classification result, the classification loss value of the sample image is obtained;

[0223] Based on the sample training loss values ​​of multiple sample images, the model parameters of the initial image enhancement model are adjusted to obtain a trained image enhancement model, including:

[0224] Based on the sample training loss and classification loss values ​​of multiple sample images, the model parameters of the initial image augmentation model are adjusted to obtain a trained image augmentation model.

[0225] The first discriminator is used to classify the sample image and the enhanced image; it can distinguish whether the input image is a real image or a generated image. Understandably, the sample image is a real image, and the enhanced image is a generated image. The second discriminator is used to classify the sample image and the restored image; it can distinguish whether the input image is a real image or a generated image. Understandably, the sample image is still a real image, and the restored image is a generated image.

[0226] For example, for each sample image, the terminal inputs the sample image and the enhanced image into a first discriminator, which classifies them to obtain a first classification result. Then, the terminal inputs the sample image and the restored image into a second discriminator, which classifies them to obtain a second classification result. Based on the first and second classification results, a classification loss value for the sample image is obtained. Using this classification loss value, the terminal can adjust the model parameters of the initial image enhancement model based on the individual sample training loss values ​​and classification loss values ​​of multiple sample images, thus obtaining a trained image enhancement model.

[0227] In practical applications, the terminal can overlay the sample training loss values ​​and classification loss values ​​of multiple sample images to calculate the total loss value of the model during training. Based on the total loss value, the gradient of each parameter in the initial image enhancement model is calculated through the backpropagation algorithm. Then, the optimization algorithm is used to update the parameters based on the gradient, thereby adjusting the model parameters of the initial image enhancement model.

[0228] In practical applications, the method for calculating the classification loss value can be configured according to the actual application scenario. For example, the specific method for calculating the classification loss value can be binary cross-entropy loss:

[0229] in, This is the Sigmoid function, representing the probability that the discriminator predicts the current sample belongs to the positive class. Its value ranges from 0 to 1. x represents the input image, also known as input data, which can be a sample image, an enhanced image, or a restored image. t is the true label of the input image, and the values ​​of t are shown in Table 1 below.

[0230] Table 1

[0231] As shown in Table 1, the discriminator is mainly used to determine whether the input data is real or generated data. If the input data is real, it indicates that the input data is a positive sample, and its true label is 1. If the input data is generated, it indicates that the input data is a negative sample, and its true label is 0. It can be understood that for the discriminator, real data refers to the sample image, and generated data refers to the enhanced or restored image. The generator is mainly used to determine whether the input is generated data. If the input data is generated, it indicates that the input data is a positive sample, and its true label is 1. It can be understood that for the generator, generated data refers to the enhanced or restored image.

[0232] In this embodiment, two discriminators are used to classify the sample image and the enhanced image, as well as the sample image and the restored image, respectively. The classification results can be used to calculate the classification loss value, which can then provide rich feedback information to help the initial image enhancement model adjust its parameters more effectively, thereby improving the image enhancement effect.

[0233] In an exemplary embodiment, based on the sample training loss values ​​and classification loss values ​​of multiple sample images, the model parameters of the initial image augmentation model are adjusted to obtain a trained image augmentation model, including:

[0234] For each sample image, the sample image and the first image enhancement prompt information are input into the initial image restoration model. The initial image restoration model extracts multi-scale features from the sample image based on the first image enhancement prompt information to obtain the features of the image to be decoded. The features of the image to be decoded are then decoded to obtain the decoded image.

[0235] By comparing the differences between the decoded image and the sample image, the image consistency loss value is determined, and the second image enhancement features and reconstruction loss value of the sample image are obtained.

[0236] Based on the image consistency loss value, the second image enhancement feature, the reconstruction loss value, and the features of the image to be decoded, the regularization loss value of the sample image is determined.

[0237] The first loss value is obtained by weighted summing of the regularization loss value and the classification loss value.

[0238] Based on the sample training loss values ​​and first loss values ​​of multiple sample images, the model parameters of the initial image enhancement model are adjusted to obtain a trained image enhancement model.

[0239] For example, during model training, the terminal needs to regularize the network used for multi-scale feature extraction while considering the semantic consistency of the image to improve the generalization and stability of the image enhancement model. Therefore, the terminal inputs the sample image and the first image enhancement prompt information into the initial image restoration model. Through the initial image restoration model, using the first image enhancement prompt information as the condition information to guide image restoration, multi-scale feature extraction is performed on the sample image to obtain the features of the image to be decoded. The features of the image to be decoded are then decoded to obtain the decoded image. Finally, the regularization loss value of the sample image is calculated using the features of the image to be decoded and the decoded image.

[0240] For example, the terminal determines the image consistency loss value by comparing the differences between the decoded image and the sample image, and obtains the second image enhancement feature and reconstruction loss value of the sample image. It then compares the differences between the second image enhancement feature and the features of the image to be decoded to determine the regularized content consistency loss value. The image consistency loss value, reconstruction loss value, and regularized content consistency loss value are then superimposed to obtain the regularized loss value. Furthermore, the regularized loss value and the classification loss value are weighted and summed to obtain the first loss value. Based on the sample training loss value and the first loss value of multiple sample images, the model parameters of the initial image enhancement model are adjusted to obtain the trained image enhancement model.

[0241] In practical applications, the image consistency loss, specifically the L1 loss, is obtained by comparing the pixel values ​​of each pixel in the decoded image and the sample image. The regularized content consistency loss is obtained by calculating the cosine similarity between the second image enhancement feature and the features of the image to be decoded. The second image enhancement feature is obtained by using the image content description information of the sample image to perform multi-scale feature extraction on the encoded features of the encoded sample image. The reconstruction loss is obtained by comparing the first reflection map obtained based on the first enhanced image feature and the third reflection map of the sample image determined through image decomposition.

[0242] In a specific application, the regularization loss value can be calculated using the following formula:

[0243] in, This refers to L1 loss, also known as image consistency loss, I. l For the sample image, C d To enhance the prompt information for the first image, f d () represents the initial image restoration model, f d (I l C d ) refers to decoding images. This refers to the loss value of regularized content consistency, E. d (Il () refers to the sixth encoded feature obtained by encoding the sample image through the image restoration encoder in the initial image restoration model, U l (E d (I l ),C d This refers to the image features to be decoded obtained by using the multi-scale restoration feature extraction network in the initial image restoration model to extract features from the sixth encoded feature based on the first image enhancement prompt information. This refers to the reconstruction loss value, I ref For the third reflection diagram, D r (Z d () is the second reflection map.

[0244] In practical applications, to improve the generalization and stability of the model, unpaired sample images and augmented images can be used simultaneously during model training. However, these image pairs are not directly matched. Based on this, the model's performance can be optimized using independent sample images and augmented images, enabling it to better handle image augmentation tasks. Therefore, the terminal can perform a weighted sum of the regularization loss and classification loss to obtain a first loss value, which can then be used to evaluate the model's training progress.

[0245] In practical applications, the total loss value of the model during training can be calculated by superimposing the sample training loss value and the first loss value of multiple sample images. Based on the total loss value, the gradient of each parameter in the initial image enhancement model is calculated by the backpropagation algorithm. Then, the optimization algorithm is used to update the parameters based on the gradient, thereby adjusting the model parameters of the initial image enhancement model.

[0246] In this embodiment, the network used for multi-scale feature extraction can be regularized while considering semantic consistency, and the regularization loss value can be determined. Then, the first loss value can be obtained by weighting the regularization loss value and the classification loss value. The first loss value can be used to adjust the initial image enhancement model to improve the generalization and stability of the image enhancement model.

[0247] In an exemplary embodiment, based on the sample training loss values ​​and first loss values ​​of multiple sample images, the model parameters of the initial image enhancement model are adjusted to obtain a trained image enhancement model, including:

[0248] Determine the supervised loss value and content consistency loss value of the sample image, and then superimpose the first loss value, the sample training loss value, the content consistency loss value, and the supervised loss value to obtain the second loss value of the sample image;

[0249] Based on the second loss values ​​of multiple sample images, the model parameters of the initial image enhancement model are adjusted to obtain a trained image enhancement model.

[0250] For example, when it is necessary to adjust the model parameters of the initial image augmentation model, the terminal determines the supervised loss value and content consistency loss value of the sample image, and superimposes the first loss value, the sample training loss value, the content consistency loss value and the supervised loss value to obtain the second loss value of the sample image. Based on the second loss values ​​of multiple sample images, the model parameters of the initial image augmentation model are adjusted to obtain the trained image augmentation model.

[0251] In practical applications, the second loss value can be calculated using the following formula:

[0252] in, This refers to the training loss value of the samples. This refers to the content consistency loss value. This refers to the monitoring loss value. This refers to the first loss value. This is the regularization loss value. Let λ be the classification loss value. idt and λ GAN These are the weighting coefficients used when weighting the regularization loss value and the classification loss value, respectively. They can be configured according to the actual application scenario or obtained through learning.

[0253] In practical applications, the total loss value of the model during training can be calculated by superimposing the second loss values ​​of multiple sample images. Based on the total loss value, the gradient of each parameter in the initial image enhancement model can be calculated using the backpropagation algorithm. Then, the optimization algorithm can be used to update the parameters based on the gradient, thereby adjusting the model parameters of the initial image enhancement model.

[0254] In this embodiment, by combining loss values ​​from multiple dimensions, the initial image enhancement model can be adjusted to improve the accuracy, generalization, and stability of the image enhancement model, thereby enhancing the image enhancement effect.

[0255] In an exemplary embodiment, the image enhancement method provided in this application can be applied to a terminal. The terminal acquires an image to be enhanced and second image enhancement prompt information for that image. It then inputs the image to be enhanced and the second image enhancement prompt information into a trained image enhancement model. Using the trained image enhancement model and the second image enhancement prompt information as guiding conditional information for image enhancement, the image to be enhanced is performed to obtain the enhanced image. Taking a dark light image as the image to be enhanced and a normal light image as the enhanced image as an example, the image enhancement effect of this application's image enhancement method can be shown in Figure 10. By using the image enhancement model and performing image enhancement on the dark light image based on the second image enhancement prompt information, a normal light image can be obtained. It should be noted that in Figure 10, dark light is represented by black and gray fill colors, and normal light is represented by white fill color.

[0256] The terminals can be, but are not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle systems, and projection devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted displays. Head-mounted displays can include virtual reality (VR) devices, augmented reality (AR) devices, and smart glasses.

[0257] In an exemplary embodiment, as shown in FIG11, an image enhancement method is provided. This embodiment illustrates the application of this method to a terminal. It is understood that the method can also be applied to a server, or to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. In this embodiment, the method includes steps 1102 to 1104. Wherein:

[0258] Step 1102: Obtain the image to be enhanced and the second image enhancement prompt information for the image to be enhanced.

[0259] Step 1104: Input the image to be enhanced and the second image enhancement prompt information into the trained image enhancement model. Through the trained image enhancement model, using the second image enhancement prompt information as the condition information to guide image enhancement, perform image enhancement on the image to be enhanced to obtain the enhanced image. The trained image enhancement model is obtained by training the image enhancement model training method described above.

[0260] For example, when image enhancement is required, the terminal obtains the image to be enhanced and the second image enhancement prompt information of the image to be enhanced. The image to be enhanced and the second image enhancement prompt information are input into the trained image enhancement model. The trained image enhancement model uses the second image enhancement prompt information as the condition information to guide image enhancement, and performs image enhancement on the image to be enhanced to obtain the enhanced image. The trained image enhancement model is trained by the above image enhancement model training method.

[0261] In practical applications, the trained image enhancement model includes a trained prompt information encoder, a trained image enhancement encoder, a trained enhancement feature extraction network, and a trained image enhancement decoder. After the image to be enhanced and the second image enhancement prompt information are input into the trained image enhancement model, the terminal encodes the second image enhancement prompt information using the trained prompt information encoder to obtain the first target encoded feature, and encodes the image to be enhanced using the trained image enhancement encoder to obtain the second target encoded feature. The enhanced feature extraction network performs multi-scale feature extraction on the second target encoded feature based on the first target encoded feature to obtain the enhanced image feature. Finally, the trained image enhancement decoder decodes the enhanced image feature to obtain the enhanced image.

[0262] In a specific application, the second image enhancement prompt information includes text prompt information and image prompt information. The text prompt information can be automatically generated by the terminal through analysis of the image to be enhanced, or it can be entered by the user. The image prompt information is image information representing illumination perception extracted from the image to be enhanced. The trained first prompt information encoder includes a trained text encoder and a trained image encoder, where the trained text encoder is used to encode the text prompt information, and the trained image encoder is used to encode the image prompt information.

[0263] The aforementioned image enhancement method, based on obtaining the image to be enhanced and its second image enhancement prompt information, inputs the image to be enhanced and the second image enhancement prompt information into a trained image enhancement model. This allows the trained image enhancement model to perform image enhancement on the image to be enhanced, using the second image enhancement prompt information as the conditional information to guide image enhancement, and obtain the enhanced image. Since the trained image enhancement model is obtained through a cyclical generation process of image enhancement and image restoration, it achieves unsupervised image enhancement model training, avoiding the problem of overfitting in the image enhancement model, thereby improving the image enhancement effect.

[0264] In an exemplary embodiment, the image enhancement model training method and image enhancement method of this application are applied to low-light enhancement as an example to illustrate the image enhancement model training method and image enhancement method of this application.

[0265] The inventors believe that most traditional image enhancement methods focus on improving the visual quality of images while neglecting the requirements of downstream tasks for the semantic quality of images. This makes them unfriendly to high-level vision tasks (including mainstream vision tasks such as classification, detection, and segmentation), resulting in low accuracy in high-level vision task processing.

[0266] Based on this, this application designs a semantically consistent unsupervised low-light image enhancement method, which can improve the semantic fidelity of enhanced images and be used for high-level vision tasks, directly improving the performance of zero-shot high-level vision tasks. Specifically, it includes the following parts: First, this application designs a low-light image enhancement method for zero-shot high-level vision, proposing a semantically consistent unsupervised low-light image enhancement framework. It defines a cyclic generation process including brightening and darkening stages, using a pre-trained diffusion model as the image generator to achieve unsupervised image enhancement. Second, this application introduces illumination-aware image cues to explicitly guide image generation and proposes a cyclic attention adapter to fully utilize the illumination-aware semantic features of the image cues and promote the semantic representation learning of latent features. Third, this application introduces title cues (i.e., image content description information) to describe the content of the input image as a supplement to the text cues, and proposes image description consistency to ensure high-level abstract semantic consistency in cyclic generation. Finally, this application introduces reflection maps to learn robust spatial semantic feature representations and proposes reflection consistency to ensure spatial semantic consistency in cyclic generation. It is understandable that image description consistency and reflectance consistency constitute semantic consistency. By learning high-level abstract semantics and spatial semantics in the cyclic generation, respectively, semantic degradation in the cyclic generation is reduced and the performance of the model on high-level vision tasks is improved, which can improve the semantic fidelity of high-level vision images.

[0267] The low-light enhancement method of this application can improve the model performance of downstream tasks by enhancing low-light images. It can be used in fields such as palmprint detection and face recognition under low-light conditions, and has strong application value. Palmprint detection is a technique for identity recognition by analyzing and comparing palmprint features. It achieves identity recognition by acquiring palmprint images, extracting palmprint features, and comparing them. When the low-light enhancement method of this application is applied to palmprint detection under low-light conditions, it can enhance the brightness of the low-light palmprint images acquired under low-light conditions. By performing feature extraction and identity recognition on the enhanced palmprint images, it can effectively improve the accuracy of palmprint feature extraction and recognition effect. It is understandable that palmprint detection has wide application value in scenarios requiring identity recognition through palm scanning. For example, palmprint detection can be applied to scenarios such as palm scan payment, palm scan access control, and palm scan attendance. The following describes the application of the image enhancement model training method of this application to low-light enhancement model training. The model framework for low-light enhancement model training is shown in Figure 12.

[0268] In an exemplary embodiment, as shown in Figure 12, the model framework for training the unsupervised dark light enhancement model in this application is defined as a loop generation process. Given a dark light image I... l The goal is to use an image generator to convert it into a normal light image. n The image generator consists of three components: an image enhancement encoder (also known as a brightening encoder), a multi-scale enhancement feature extraction network, and an image enhancement decoder (also known as a brightening decoder). Specifically, the image generator can be a generator based on SD-Turbo (an efficient single-step text-to-image generation model), and the multi-scale enhancement feature extraction network can be a U-net network.

[0269] Therefore, we define two objective functions, namely the brightening function f. l (I l C l ):I l →I n and the darkening function f d (I d C d ):I n →I l And corresponding brightening encoders, brightening decoders, darkening encoders (i.e., image restoration encoders) and darkening decoders (i.e., image restoration decoders) to achieve cyclic generation, where the conditional input C l For example, the image enhancement (brightening) prompt information input to the brightening function can be a first image enhancement prompt information, including a first text prompt information and a first image prompt information, C. dFor example, the image restoration (darkening) prompt information input to the darkening function can be an image restoration prompt information that includes a second text prompt and a second image prompt. The iterative generation process can be represented as: I′ l =f d (f l (I l C l ),C d );

[0270] Among them, I l Given the input sample image, f l (I l C l () refers to the enhanced image obtained through image enhancement (brightening), which is I in Figure 12. n , I′ l This is the restored image obtained through image restoration (darkening).

[0271] In a specific application, as shown in Figure 12, the first text encoding feature can be obtained by encoding the first text prompt information using a pre-trained first text encoder, and the first image encoding feature can be obtained by encoding the first image prompt information using a pre-trained first image encoder. The first text encoding feature and the first image encoding feature will be input into the recurrent attention adapter (CA-Adapter) proposed in this application to enrich the latent semantic features and further maintain semantic consistency with multi-query information.

[0272] In a specific application, the recurrent attention adapter is part of a multi-scale augmented feature extraction network. The first text-encoded feature and the first image-encoded feature are applied to each layer of the multi-scale augmented feature extraction network. Taking the U-net network as an example, as shown in Figure 13, the first text-encoded feature and the first image-encoded feature are used as inputs to the U-net network and applied to each layer of the U-net network.

[0273] Furthermore, as shown in Figure 12, this application also proposes Image Description Consistency (CC) and Reflection Map Consistency (RC). Image Description Consistency utilizes the description of the input image to maintain high-level abstract semantic consistency of latent features during recurrent generation. Reflection Map Consistency learns robust and consistent spatial semantic features during recurrent generation. The recurrent attention adapter, Image Description Consistency, Reflection Map Consistency, and the unsupervised training concept are explained below.

[0274] In a specific application, as shown in Figure 12, when a dark image I... lWhen using sample images as input, it is necessary to first generate first image enhancement prompts for the low-light image, and then, based on these first image enhancement prompts, generate image restoration prompts to restore the enhanced image to the low-light image. Specifically, during model training, the first text prompts in the first image enhancement prompts ("normal light image" as shown in Figure 12) and the second text prompts in the image restoration prompts ("low-light image" as shown in Figure 12) can be pre-generated. For example, they can be pre-input by the user. The first image prompts in the first image enhancement prompts can be obtained by inverting the pixel values ​​of each pixel in the image information representing illumination perception extracted from the low-light image, while the second image prompts in the image restoration prompts can be the image information representing illumination perception extracted from the low-light image. Let the image information representing illumination perception extracted from the low-light image be the V-channel image I. v Taking this as an example, the terminal will output the V channel image I... v This serves as the second image prompt, and its reverse image is calculated. The inverted image is used as the first image cue. The first and second image cuees are used in the brightening and brightening cycle.

[0275] In a specific application, as shown in Figure 12, this application proposes a recurrent attention adapter to query the first image encoded features after the first image cue information has been encoded by the image encoder and processed by the multilayer perceptron, and to provide the feature response to the first image encoded features through another cross-attention layer. Specifically, as shown in Figure 14, firstly, given the first-scale features in any first encoding layer or the fused features in any first decoding layer of the multi-scale enhancement feature extraction network, the output feature Z is used as the first-scale feature Z. u And give the first image coding feature c i The output feature Z of the first cross-attention layer i It can be defined by the following equation:

[0276] Q i =c i W′ q , is the first query vector, c i It is the first image coding feature, W′ q It is a learnable linear projection layer used to linearly map the encoded features of the first image, K. u =Z u , is the first key vector, V u =Z u , is the first value vector, d is the dimension of the first key vector, and the softmax function converts the attention scores into a probability distribution, representing the attention weight of the query vector to each key vector.

[0277] Then, in the second cross-attention layer, the output feature Z of the first cross-attention layer is... i Z serves as both query feature and output feature. u The feature response in the first attention stage (i.e., the response to the output feature Z) u The feature Z is obtained by performing cross-attention interaction with the first text encoding feature. t Feedback is given to the first image encoding feature c i The method is as follows:

[0278] First, the output feature Z of the first cross-attention layer is calculated using the following formula. i Perform cross-attention interaction with the encoded features of the first image:

[0279] Z i K is the second query vector. i =c i W′ k , is the second key vector, c i It is the first image coding feature, W′ k A learnable linear projection layer, used to linearly map the encoded features of the first image, V i =c i W′ v , is the second value vector, where W′ v A learnable linear projection layer is used to linearly map the encoded features of the first image, where d is the dimension of the second key vector. The softmax function transforms the attention scores into a probability distribution, representing the attention weight of the query vector to each key vector. Understandably, by using two cross-attention layers, illumination-aware image cues can be fully utilized.

[0280] Secondly, the output feature Z will be... u The feature Z is obtained by performing cross-attention interaction with the first text encoding feature. t , and the output feature Z of the first cross-attention layer i The feature Z′ is obtained by performing cross-attention interaction with the first image encoding features. i The components are concatenated to obtain the final output, which is the output feature Z. u The output of the first encoding layer or the first decoding layer. As shown in Figure 14, the final output of the decoupled cross-attention becomes the original text cross-attention feature Z. t With Z′ i The sum of.

[0281] In a specific application, regarding the image description consistency part, as shown in Figure 12, taking the input low-light image as an example, we use a pre-trained image description model to generate image content description information for the low-light image (as shown in Figure 12, "The view of the city can be seen from the window of a building"), and ensure that the second image enhancement feature output by the multi-scale enhancement feature extraction network in the initial image enhancement model is consistent with the restored image features output by the multi-scale restoration feature extraction network in the initial image restoration model. This process can be represented as:

[0282] in, U represents cosine similarity. init (E l (I l ),C cap ) represents the second image enhancement feature, I l E represents a low-light image. l (I l C represents the encoded feature (i.e., the fifth encoded feature) obtained by encoding the low-light image through the image enhancement encoder in the initial image enhancement model. cap U represents image content description information. init () indicates multi-scale feature extraction, U d (E d (I n ),C d ) represents the features of the restored image, I n E represents an enhanced image. d (I n ) represents the fourth coding feature, C d U indicates a prompt for image restoration. d () indicates multi-scale feature extraction.

[0283] As can be understood, as shown in Figure 12, in the image description consistency part, the initial image enhancement model is used asynchronously, that is, this process is independent of the iterative generation process, and the second image enhancement feature needs to be obtained by multi-scale feature extraction of the fifth coding feature based on the image content feature.

[0284] In this way, the consistency of features extracted by the brightening encoder and the darkening encoder can be guaranteed, allowing the recurrent attention adapter to focus more on high-level semantic information.

[0285] In a specific application, regarding the reflectance consistency part, as shown in Figure 12, taking the input dark-light image as an example, the supervised loss is divided into reflectance consistency loss. and reconstruction losses The two parts can be represented as:

[0286] Z l =U l (E l (I l ),C l Z is the first enhanced image feature. d =U d (E d (I n ),C d ), is the feature of the restored image, I ref This represents the third reflection image. It is the mean squared error loss value. It is L1 loss.

[0287] Specifically, as shown in Figure 12, the terminal inputs the first enhanced image features obtained through the initial image enhancement model into the pre-trained first reflection decoder for decoding to obtain the first reflection map, and inputs the restored image features into the pre-trained second reflection decoder for decoding to obtain the second reflection map. By comparing the first reflection map and the second reflection map, a reflection map consistency loss value is obtained. At the same time, the dark light image is input into the pre-trained image decomposition model for decomposition to obtain the third reflection map of the dark light image. By comparing the first reflection map and the third reflection map, a reconstruction loss value is obtained. Based on the reflection map consistency loss value and the reconstruction loss value, a supervised loss value for the dark light image is obtained.

[0288] In a specific application, during training, we use fine-tuning adapters (such as the LoRA adapter) to fine-tune the pre-trained encoder and decoder to construct the latent spaces for brightening and darkening. Using a low-light image I... l For example, the cycle consistency loss (i.e., the sample training loss value) can be expressed as:

[0289] in This represents L1 loss. l (I l C l ) represents an enhanced image, C l To enhance the prompt information for the first image, f d (f l (I l C l ),C d ) represents the restored image, C d Provides a prompt for image restoration.

[0290] Furthermore, we use two discriminators to classify the brightening and darkening results output by the model, as well as the real image, respectively. The GAN loss (i.e., the classification loss value) is... Specifically, the terminal inputs the low-light image and the enhanced image into a first discriminator, which classifies them to obtain a first classification result. Then, it inputs the low-light image and the restored image into a second discriminator, which classifies them to obtain a second classification result. Based on the first and second classification results, a classification loss value for the low-light image is obtained. The classification loss value can be calculated using a binary cross-entropy loss method.

[0291] in, is the Sigmoid function, which represents the probability that the discriminator predicts the current sample belongs to the positive class. The value range is between 0 and 1. x represents the input image, which can also be called input data. Specifically, it can be a sample image, an enhanced image, or a restored image. t is the true label of the input image. The value of t can be as shown in Table 1.

[0292] Furthermore, for the regularization of the feature network, we also considered its semantic consistency. Taking low-light images as an example, the regularization loss value can be expressed as:

[0293] in, This refers to L1 loss, I l For low-light images, C d To enhance the prompt information for the first image, f d () represents the initial image restoration model, f d (I l C d ) refers to the decoded image, E d (I l ) refers to the sixth encoded feature obtained by encoding a low-light image through the image restoration encoder in the initial image restoration model, U l (E d (I l ),C d This refers to the image features to be decoded obtained by using the multi-scale restoration feature extraction network in the initial image restoration model to extract features from the sixth encoded feature based on the first image enhancement prompt information. This refers to the reconstruction loss value, I ref For the third reflection diagram, D r (Z d ( ) represents the second reflection image. It is understandable that for regularization of the feature network, the dark light image and the first image enhancement prompt information need to be input into the initial image restoration model, allowing the initial image restoration model to restore the normal light image. Then, the normal light image restored by the initial image restoration model is used to calculate the regularization loss value. Therefore, when performing image restoration on the enhanced image, C... dThe prompts for image restoration are different, so C will respond differently. d Enhance the prompt information for the first image.

[0294] In practical applications, to improve the generalization and stability of the model, unpaired sample images and augmented images can be used simultaneously during model training. However, these image pairs are not directly matched. Based on this, the model's performance can be optimized using independent sample images and augmented images, enabling it to better handle image augmentation tasks. Therefore, the terminal can perform a weighted sum of the regularization loss and classification loss to obtain a first loss value, which is used to evaluate the model training. The total loss value (i.e., the second loss value) of the dark light enhancement model in this application during training can be calculated using the following formula:

[0295] in, This refers to the training loss value of the samples. This refers to the content consistency loss value. This refers to the monitoring loss value. This refers to the first loss value. This is the regularization loss value. Let λ be the classification loss value. idt and λ GAN These are the weighting coefficients used when weighting the regularization loss value and the classification loss value, respectively. They can be configured according to the actual application scenario or obtained through learning.

[0296] It should be noted that, in this embodiment, when adjusting the parameters of modules or layers in the model framework, only the parameters of the parts marked as being trained can be adjusted, while the parameters of the parts marked as frozen modules are not adjusted.

[0297] Extensive experiments demonstrate that our proposed method not only outperforms state-of-the-art methods in terms of image visual quality but also excels in advanced visual tasks, including nighttime image classification, dark face detection, and nighttime semantic segmentation. The following section provides a comprehensive comparison of our image enhancement method applied to low-light enhancement with traditional low-light enhancement methods (T), fully supervised methods (S), and unsupervised methods (U), illustrating that our proposed method surpasses current methods in both image visual quality comparison and advanced visual tasks. The low-light enhancement methods include LIME (Low-light Image Enhancement via Illumination) and DUAL (Dual Illumination Estimation). Fully supervised methods include RetinexNet and Retinexformer (a single-stage Transformer low-light image enhancement method based on Retinex theory).Unsupervised methods include EnligtenGan (an unsupervised low-light image enhancement method based on Generative Adversarial Networks (GANs)), Zero-DCE (Zero-Reference Deep Curve Estimation for Low-Light Image Enhancement, a deep learning method for low-light image enhancement), Zero-DCE++, RUAS (Retinex-inspired Unrolling with Cooperative Prior Architecture Search, a low-light image enhancement method), SCI (Self-Calibrated Illumination, a low-light image enhancement method), PairLIE (an unsupervised low-light image enhancement method based on paired low-light images), SADG (Segment Any Dynamic Gaussian Without Object Trackers, a method for dynamic Gaussian segmentation), CLIP-LIT (an unsupervised backlight image enhancement method), NeRCo (implicit neural representations for collaborative low-light image enhancement), QuadPrior (Zero-Reference Low-Light Enhancement via Physical Quadruple Priors, an unsupervised backlight image enhancement method), and ZERO-IG-LSRW (Zero-Shot). (Illumination-Guided Joint Denoising and Adaptive Enhancement for Low-Light Images, a zero-shot method for low-light image enhancement and denoising), ZERO-IG-LOL (ZERO-IG: Zero-Shot Illumination-Guided Joint Denoising and Adaptive Enhancement for Low-Light Images, a zero-shot method for low-light image enhancement and denoising), and LigtenDiffusion (an unsupervised low-light image enhancement framework based on a diffusion model).

[0298] First, the visual quality of different methods on the LSRW (Large-Scale Real-World paired low / normal-light images dataset, specifically designed for low-light image enhancement tasks) and LOL (Low-Light Dataset, a benchmark dataset for low-light image enhancement tasks containing pairs of low-light and normal-light images) datasets is compared, as shown in Table 2 below. PSNR (Peak Signal-to-Noise Ratio) is a commonly used objective evaluation metric in image processing and image enhancement tasks, used to measure image quality. It evaluates the effectiveness of image enhancement by comparing the difference between the enhanced image and the original image (or reference image). SSIM (Structural Similarity Index) is a metric used to measure the similarity between two images in terms of brightness, contrast, and structure. Unlike the traditional PSNR, SSIM is closer to the perception of the human visual system and can more accurately reflect image quality. LPIPS (Learned Perceptual Image Patch Similarity) is a deep learning-based image quality assessment metric used to measure the perceptual similarity between two images. Compared to traditional pixel-level assessment metrics (such as PSNR and SSIM), LPIPS can more closely approximate human visual perception in evaluating image quality.

[0299] Table 2 Comparison of visual quality of different methods on the LSRW and LOL datasets.

[0300] As shown in Table 2, compared with traditional methods, the low-light enhancement method of this application demonstrates better visual quality on the dataset in terms of peak signal-to-noise ratio, structural similarity index, and perceptual similarity. Furthermore, as shown in Figure 15, comparisons with Ground Truth (GT) and various traditional unsupervised methods also show that the visual quality of the low-light enhancement method of this application is superior to many unsupervised methods, exhibiting better enhancement effects on low-light images. It effectively enhances low-light images (Figure 15 uses white and gray fills to represent brightening, where gray represents a lower degree of brightening than white). This indicates that the low-light enhancement method of this application achieves a higher degree of brightening, thus improving the image enhancement effect.

[0301] Secondly, for advanced vision tasks, this method was tested on low-light image classification (CODaN (Common Objects Day and Night, a dataset for low-light image classification)), low-light face detection (DARK FACE, a dataset specifically for face detection under low-light conditions), and nighttime image segmentation (BDD100k-night, a large-scale autonomous driving image dataset). As shown in Table 3, this method achieved the best performance in all these tasks. Top-1 accuracy refers to the proportion of the highest probability class predicted by the model that matches the actual class. mAP (mean Average Precision) is a commonly used evaluation metric in object detection, representing the average precision of the model across all classes. mIoU (Mean Intersection over Union) is an important metric in computer vision used to evaluate the performance of image segmentation models.

[0302] Table 3 Comparison of results of different methods in advanced vision tasks

[0303] As shown in Table 3, compared with traditional methods, the dark light enhancement method of this application outperforms traditional dark light enhancement methods in terms of Top-1 accuracy, mean precision, and mean intersection-union ratio in advanced vision tasks, achieving the best performance.

[0304] For example, as shown in Figure 16, taking low-light face detection as an example of a high-level vision task, and using the ground truth (which can be a well-enhanced image as shown in Figure 16) as a reference, it can be seen from the comparison with various traditional unsupervised methods that the enhanced image obtained by using the low-light enhancement method of this application can detect more facial regions than the enhanced images obtained by various traditional unsupervised methods. That is, the accuracy of low-light face detection is the highest.

[0305] For example, as shown in Figure 17, taking nighttime image segmentation as an example of a high-level vision task, it can be seen from the comparison with GT and various traditional unsupervised methods that by performing image segmentation on the enhanced image obtained after image enhancement using the dark light enhancement method of this application, a more accurate image segmentation result can be obtained, segmenting the input dark light image into multiple components with clear boundaries.

[0306] To illustrate further, as shown in Figure 18, taking nighttime image segmentation as an example of a high-level vision task, a comparison with ground truth (GT) and various traditional unsupervised methods reveals that image segmentation using the enhanced image obtained after image enhancement using the low-light enhancement method of this application yields more accurate results, separating the vehicle region, road region, and environment region from the input low-light image. In Figure 18, different fill patterns represent the vehicle region and road region, while blank spaces represent the environment region.

[0307] Finally, our proposed method was compared with traditional zero-shot day-night domain adaptation methods. While not requiring downstream tasks for learning, our method still achieved the best performance, as shown in Table 4. Traditional zero-shot day-night domain adaptation methods include: CIConv (Color Invariant Convolution, a convolutional layer for zero-shot day-night domain adaptation), Sim-MinMax (a method for zero-shot domain adaptation), and DAI-Net (DArk-Illuminated Network, a method for zero-shot day-night domain adaptation).

[0308] Table 4 compares the results with traditional zero-shot day / night domain adaptation methods.

[0309] As shown in Table 4, compared with the traditional zero-shot day-night domain adaptation method, the dark light enhancement method of this application outperforms the traditional zero-shot day-night domain adaptation method in terms of mean accuracy and mean cross-union ratio, achieving the best performance.

[0310] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0311] Based on the same inventive concept, this application also provides an image enhancement model training apparatus for implementing the image enhancement model training method described above, and an image enhancement apparatus for implementing the image enhancement method designed above. The solution provided by this apparatus is similar to the implementation scheme described in the above method. Therefore, the specific limitations of one or more image enhancement model training apparatuses and image enhancement apparatus embodiments provided below can be found in the limitations of the image enhancement model training method and image enhancement method described above, and will not be repeated here.

[0312] In an exemplary embodiment, as shown in FIG19, an image enhancement model training device is provided, including: a sample data acquisition module 1902, an image enhancement module 1904, an image restoration module 1906, a loss calculation module 1908, and a parameter adjustment module 1910, wherein:

[0313] The sample data acquisition module 1902 is used to acquire multiple sample images and the first image enhancement prompt information for each sample image;

[0314] The image enhancement module 1904 is used to input the sample image and the first image enhancement prompt information into the initial image enhancement model for each sample image. Through the initial image enhancement model, the first image enhancement prompt information is used as the condition information to guide image enhancement, and the sample image is enhanced to obtain an enhanced image.

[0315] The image restoration module 1906 is used to generate image restoration prompts to restore the enhanced image to the sample image based on the first image enhancement prompts. The enhanced image and the image restoration prompts are input into the initial image restoration model. The image restoration is performed on the enhanced image through the initial image restoration model, using the image restoration prompts as the condition information to guide the image restoration, to obtain the restored image.

[0316] The loss calculation module 1908 is used to compare the differences between the sample image and the restored image to determine the sample training loss value of the sample image.

[0317] The parameter adjustment module 1910 is used to adjust the model parameters of the initial image enhancement model based on the sample training loss values ​​of multiple sample images, so as to obtain a trained image enhancement model.

[0318] The aforementioned image enhancement model training device, upon acquiring multiple sample images and first image enhancement prompts for each sample image, inputs the sample image and the first image enhancement prompts into an initial image enhancement model for each sample image. This enables the initial image enhancement model to enhance the sample image using the first image enhancement prompts as guiding conditions, resulting in an enhanced image. Based on the first image enhancement prompts, image restoration prompts are generated to restore the enhanced image back to the sample image. Inputting the enhanced image and the image restoration prompts into an initial image restoration model enables the initial image restoration model to restore the enhanced image using the image restoration prompts as guiding conditions, resulting in a restored image. By comparing the sample image and the restored image, the sample training loss value of the sample image can be determined. Furthermore, based on the sample training loss values ​​of multiple sample images, the model parameters of the initial image enhancement model are adjusted to obtain a trained image enhancement model. The entire process defines a loop generation process that includes image enhancement and image restoration, enabling unsupervised training of the image enhancement model. This avoids the problem of overfitting in the image enhancement model, thus obtaining an image enhancement model that can support improved image enhancement effects.

[0319] In an exemplary embodiment, the initial image enhancement model includes a first prompt information encoder, an image enhancement encoder, a multi-scale enhancement feature extraction network, and an image enhancement decoder. The image enhancement module is further configured to encode a first image enhancement prompt information using the first prompt information encoder to obtain a first encoded feature, and to encode a sample image using the image enhancement encoder to obtain a second encoded feature. The multi-scale enhancement feature extraction network is then used to extract multi-scale features from the second encoded feature based on the first encoded feature to obtain a first enhanced image feature. Finally, the image enhancement decoder is used to decode the first enhanced image feature to obtain an enhanced image.

[0320] In an exemplary embodiment, the first image enhancement prompt information includes first text prompt information and first image prompt information. The image enhancement module is further configured to encode the first text prompt information using the first text encoder in the first prompt information encoder to obtain first text encoding features, and to encode the first image prompt information using the first image encoder in the first prompt information encoder to obtain first image encoding features. The first text encoding features and the first image encoding features are used as first encoding features. A multi-scale enhancement feature extraction network is used to extract second encoding features based on the first text encoding features and the first image encoding features to obtain first enhanced image features.

[0321] In an exemplary embodiment, the image enhancement module is further configured to extract image information representing illumination perception from the sample image, and to invert the pixel value of each pixel in the image information representing illumination perception to obtain the first image prompt information.

[0322] In an exemplary embodiment, the multi-scale enhanced feature extraction network includes N-level first encoding layers, N-level first decoding layers, and a first bridging layer, where N is a positive integer greater than 1. The image enhancement module is further configured to perform multi-scale feature extraction on the second encoding features based on the first text encoding features and the first image encoding features through the N-level first encoding layers to obtain the enhanced encoding features output by the N-level first encoding layer; to perform feature extraction on the enhanced encoding features output by the N-level first encoding layer based on the first text encoding features and the first image encoding features through the first bridging layer to obtain the enhanced encoding features output by the first bridging layer; to perform multi-scale feature extraction on the enhanced encoding features output by the first bridging layer based on the first text encoding features and the first image encoding features through the N-level first decoding layers to obtain the enhanced decoding features output by the N-level first decoding layer; and to use the enhanced decoding features output by the N-level first decoding layer as the first enhanced image features.

[0323] In an exemplary embodiment, the image enhancement module is further configured to, for the nth level of the first coding layer in N levels, when n equals 1, use the second coding feature as part of the input data of the nth level of the first coding layer; when n is greater than 1 and less than or equal to N, use the enhanced coding feature output by the (n-1)th first coding layer as part of the input data of the nth level of the first coding layer; in the nth level of the first coding layer, perform feature extraction on part of the input data of the nth level of the first coding layer to obtain the first scale feature; and perform cross-attention interaction on the first scale feature, the first text coding feature, and the first image coding feature to obtain the enhanced coding feature output by the nth level of the first coding layer.

[0324] In an exemplary embodiment, the image enhancement module is further configured to perform cross-attention interaction on the first scale feature and the first text encoding feature to obtain a first interactive feature, and perform cross-attention interaction on the first scale feature and the first image encoding feature to obtain a second interactive feature, and concatenate the first interactive feature and the second interactive feature to obtain the enhanced encoding feature output by the first encoding layer at the nth level.

[0325] In an exemplary embodiment, the image enhancement module is further configured to perform linear mapping on the first image encoding features to obtain a first query vector, and use the first scale features as a first key vector and a first value vector, perform cross-attention interaction on the first query vector, the first key vector, and the first value vector to obtain a third interaction feature, use the third interaction feature as a second query vector, and perform linear mapping on the first image encoding features to obtain a second key vector and a second value vector, and perform cross-attention interaction on the second query vector, the second key vector, and the second value vector to obtain a second interaction feature.

[0326] In an exemplary embodiment, the image enhancement module is further configured to, for the nth level of the first decoding layer in N levels, when n equals 1, use the enhanced coding features output by the first bridging layer as part of the input data of the nth first decoder; when n is greater than 1 and less than or equal to N, use the enhanced decoding features output by the (n-1)th first decoding layer as part of the input data of the nth level first decoding layer; in the nth level first decoding layer, perform feature extraction on part of the input data of the nth level first decoding layer to obtain second-scale features; fuse the second-scale features and the enhanced coding features output by the first encoding layer corresponding to the nth level first decoding layer to obtain fused features; and perform cross-attention interaction on the fused features, the first text encoding features, and the first image encoding features to obtain the enhanced decoding features output by the nth level first decoding layer.

[0327] In an exemplary embodiment, the first image enhancement prompt information includes a first text prompt information and a first image prompt information. The image restoration module is further configured to generate a second text prompt information that restores the enhanced image to the sample image based on the first text prompt information, and to invert the pixel value of each pixel in the first image prompt information to obtain the second image prompt information. The second text prompt information and the second image prompt information are used as the image restoration prompt information.

[0328] In an exemplary embodiment, the initial image restoration model includes a second prompt information encoder, an image restoration encoder, a multi-scale restoration feature extraction network, and an image restoration decoder. The image restoration module is further configured to encode the image restoration prompt information using the second prompt information encoder to obtain a third encoded feature, and to encode the enhanced image using the image restoration encoder in the initial image restoration model to obtain a fourth encoded feature. The multi-scale restoration feature extraction network then performs multi-scale feature extraction on the fourth encoded feature based on the third encoded feature to obtain restored image features. Finally, the image restoration decoder decodes the restored image features to obtain the restored image.

[0329] In an exemplary embodiment, the loss calculation module is further configured to input the sample image into a pre-trained image description model to obtain image content description information of the sample image; encode the image content description information using a pre-trained text encoder to obtain image content features; input the sample image into the image enhancement encoder in the initial image enhancement model to encode the sample image to obtain a fifth encoded feature; extract the fifth encoded feature using a multi-scale enhancement feature extraction network in the initial image enhancement model based on the image content features to obtain a second image enhancement feature; compare the difference between the second image enhancement feature and the restored image feature to determine the content consistency loss value of the sample image; and adjust the model parameters of the initial image enhancement model based on the sample training loss value and content consistency loss value of each of the multiple sample images to obtain a trained image enhancement model.

[0330] In an exemplary embodiment, the loss calculation module is further configured to input the first enhanced image features obtained through the initial image enhancement model into a pre-trained first reflection decoder for decoding to obtain a first reflection map, and input the restored image features into a pre-trained second reflection decoder for decoding to obtain a second reflection map, compare the differences between the first reflection map and the second reflection map to obtain a reflection map consistency loss value, input the sample image into a pre-trained image decomposition model for decomposition to obtain a third reflection map of the sample image, compare the differences between the first reflection map and the third reflection map to obtain a reconstruction loss value, obtain a supervised loss value of the sample image based on the reflection map consistency loss value and the reconstruction loss value, and adjust the model parameters of the initial image enhancement model based on the sample training loss value and supervised loss value of each of the multiple sample images to obtain a trained image enhancement model.

[0331] In an exemplary embodiment, the loss calculation module is further configured to, for each sample image, input the sample image and the enhanced image into a first discriminator, classify the sample image and the enhanced image through the first discriminator to obtain a first classification result, input the sample image and the restored image into a second discriminator, classify the sample image and the restored image through the second discriminator to obtain a second classification result, obtain a classification loss value for the sample image based on the first classification result and the second classification result, and adjust the model parameters of the initial image enhancement model based on the sample training loss value and classification loss value of each of the multiple sample images to obtain a trained image enhancement model.

[0332] In one exemplary embodiment, the loss calculation module is further configured to, for each sample image, input the sample image and the first image enhancement prompt information into the initial image restoration model, extract multi-scale features from the sample image based on the first image enhancement prompt information through the initial image restoration model to obtain the features of the image to be decoded, decode the features of the image to be decoded to obtain the decoded image, compare the differences between the decoded image and the sample image to determine the image consistency loss value, and obtain the second image enhancement features and reconstruction loss value of the sample image, determine the regularization loss value of the sample image based on the image consistency loss value, the second image enhancement features, the reconstruction loss value, and the features of the image to be decoded, perform a weighted summation of the regularization loss value and the classification loss value to obtain the first loss value, and adjust the model parameters of the initial image enhancement model based on the sample training loss value and the first loss value of each of the multiple sample images to obtain the trained image enhancement model.

[0333] In an exemplary embodiment, the loss calculation module is further configured to determine the supervised loss value and content consistency loss value of the sample image, and to superimpose the first loss value, the sample training loss value, the content consistency loss value and the supervised loss value to obtain the second loss value of the sample image. Based on the second loss values ​​of multiple sample images, the model parameters of the initial image enhancement model are adjusted to obtain the trained image enhancement model.

[0334] In an exemplary embodiment, as shown in FIG20, an image enhancement device is provided, including: a data acquisition module 2002 and a processing module 2004, wherein:

[0335] The data acquisition module 2002 is used to acquire the image to be enhanced and the second image enhancement prompt information of the image to be enhanced;

[0336] The processing module 2004 is used to input the image to be enhanced and the second image enhancement prompt information into the trained image enhancement model. Through the trained image enhancement model, the second image enhancement prompt information is used as the condition information to guide image enhancement, and the image to be enhanced is enhanced to obtain the enhanced image. The trained image enhancement model is trained by the above image enhancement model training method.

[0337] The aforementioned image enhancement device, based on acquiring the image to be enhanced and the second image enhancement prompt information for the image to be enhanced, inputs the image to be enhanced and the second image enhancement prompt information into a trained image enhancement model. This enables the trained image enhancement model to enhance the image to be enhanced based on the second image enhancement prompt information, resulting in an enhanced image. Since the trained image enhancement model is obtained through a cyclical generation process of image enhancement and image restoration, it achieves unsupervised image enhancement model training, thus avoiding the problem of overfitting in the image enhancement model and improving the image enhancement effect.

[0338] The aforementioned image enhancement model training device and image enhancement device, each module of which can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0339] In an exemplary embodiment, a computer device is provided. This computer device can be a terminal or a server. Taking the computer device as a terminal as an example, its internal structure diagram is shown in Figure 21. The computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used for exchanging information between the processor and external devices. The communication interface of the computer device is used for wired or wireless communication with external terminals. Wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements an image enhancement model training method and an image enhancement method. The display unit of this computer device is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of this computer device can be a touch layer covering the display screen, or buttons, a trackball, or a touchpad set on the casing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0340] Those skilled in the art will understand that the structure shown in Figure 21 is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0341] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0342] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0343] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0344] It should be noted that the data involved in this application (including but not limited to data used for analysis, data stored, data displayed, etc.) are all data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0345] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0346] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0347] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A method for training an image enhancement model, executed by an electronic device, the method comprising: Acquire multiple sample images and a first image enhancement prompt message for each of the sample images; For each sample image, the sample image and the first image enhancement prompt information are input into an initial image enhancement model. The initial image enhancement model uses the first image enhancement prompt information as conditional information to guide image enhancement, and the sample image is enhanced to obtain an enhanced image. Based on the first image enhancement prompt information, image restoration prompt information is generated to restore the enhanced image to the sample image. The enhanced image and the image restoration prompt information are input into an initial image restoration model. Through the initial image restoration model, using the image restoration prompt information as conditional information to guide image restoration, the enhanced image is restored to obtain the restored image. By comparing the differences between the sample image and the restored image, the sample training loss value of the sample image is determined; and Based on the sample training loss values ​​of each of the multiple sample images, the model parameters of the initial image enhancement model are adjusted to obtain a trained image enhancement model.

2. The method according to claim 1, wherein the initial image enhancement model comprises a first cue information encoder, an image enhancement encoder, a multi-scale enhancement feature extraction network, and an image enhancement decoder; the step of enhancing the sample image using the initial image enhancement model, with the first image enhancement cue information as conditional information for guiding image enhancement, to obtain an enhanced image, comprises: The first image enhancement prompt information is encoded using the first prompt information encoder to obtain a first encoding feature, and the sample image is encoded using the image enhancement encoder to obtain a second encoding feature; The first enhanced image feature is obtained by using the multi-scale enhanced feature extraction network to extract multi-scale features from the second encoded feature based on the first encoded feature. and The image enhancement decoder decodes the features of the first enhanced image to obtain the enhanced image.

3. The method according to claim 2, wherein the first image enhancement prompt information includes first text prompt information and first image prompt information; The step of encoding the first image-enhanced prompt information using the first prompt information encoder to obtain the first encoded feature includes: The first text prompt information is encoded by the first text encoder in the first prompt information encoder to obtain the first text encoding feature, and the first image prompt information is encoded by the first image encoder in the first prompt information encoder to obtain the first image encoding feature. The first text encoding feature and the first image encoding feature are used as the first encoding feature; and The step of extracting first enhanced image features by performing multi-scale feature extraction on the second encoded features based on the first encoded features through the multi-scale enhanced feature extraction network includes: The first enhanced image feature is obtained by performing multi-scale feature extraction on the second encoding feature based on the first text encoding feature and the first image encoding feature through the multi-scale enhanced feature extraction network.

4. The method according to claim 3, further comprising: Image information representing illumination perception is extracted from the sample image, and the pixel value of each pixel in the image information representing illumination perception is inverted to obtain the first image prompt information.

5. The method according to claim 3, wherein the multi-scale enhanced feature extraction network comprises N-level first encoding layers, N-level first decoding layers, and a first bridging layer; N is a positive integer greater than 1; the step of performing multi-scale feature extraction on the second encoding features based on the first text encoding features and the first image encoding features through the multi-scale enhanced feature extraction network to obtain the first enhanced image features includes: Through the N-level first coding layers, based on the first text coding features and the first image coding features, multi-scale feature extraction is performed on the second coding features to obtain the enhanced coding features output by the N-level first coding layer; Through the first bridging layer, based on the first text encoding features and the first image encoding features, feature extraction is performed on the enhanced encoding features output by the first encoding layer of the Nth level to obtain the enhanced encoding features output by the first bridging layer; Through the N-level first decoding layers, based on the first text encoding features and the first image encoding features, multi-scale feature extraction is performed on the enhanced encoding features output by the first bridging layer to obtain the enhanced decoding features output by the N-level first decoding layer. and The enhanced decoding features output by the first decoding layer of the Nth level are used as the first enhanced image features.

6. The method according to claim 5, wherein the step of performing multi-scale feature extraction on the second coding features based on the first text coding features and the first image coding features through the N-level first coding layers to obtain the enhanced coding features output by the nth-level first coding layer includes: For the nth level of the first coding layer in the N levels, when n equals 1, the second coding feature is used as part of the input data of the nth level of the first coding layer; when n is greater than 1 and less than or equal to N, the enhanced coding feature output by the (n-1)th first coding layer is used as part of the input data of the nth level of the first coding layer. In the first encoding layer of the nth level, feature extraction is performed on a portion of the input data of the first encoding layer of the nth level to obtain the first scale feature. Cross-attention interaction is performed on the first scale feature, the first text encoding feature and the first image encoding feature to obtain the enhanced encoding feature output by the first encoding layer of the nth level.

7. The method according to claim 6, wherein performing cross-attention interaction on the first scale feature, the first text encoding feature, and the first image encoding feature to obtain the enhanced encoding feature output by the nth level first encoding layer includes: Cross-attention interaction is performed on the first scale feature and the first text encoding feature to obtain the first interactive feature, and cross-attention interaction is performed on the first scale feature and the first image encoding feature to obtain the second interactive feature. and By concatenating the first interaction feature and the second interaction feature, the enhanced coding feature output by the first coding layer of the nth level is obtained.

8. The method according to claim 7, wherein performing cross-attention interaction on the first scale feature and the first image coding feature to obtain the second interaction feature includes: The first image encoding features are linearly mapped to obtain the first query vector, and the first scale features are used as the first key vector and the first value vector. A third interaction feature is obtained by performing cross-attention interaction on the first query vector, the first key vector, and the first value vector. The third interactive feature is used as the second query vector, and the first image encoding feature is linearly mapped to obtain the second key vector and the second value vector. and The second query vector, the second key vector, and the second value vector are subjected to cross-attention interaction to obtain the second interaction feature.

9. The method according to claim 5, wherein the step of performing multi-scale feature extraction on the enhanced coding features output by the first bridging layer based on the first text coding features and the first image coding features through the N-level first decoding layers to obtain the enhanced decoding features output by the N-level first decoding layer includes: For the nth level of the first decoding layer in the N levels, when n equals 1, the enhanced coding features output by the first bridging layer are used as part of the input data of the nth first decoder; when n is greater than 1 and less than or equal to N, the enhanced decoding features output by the (n-1)th first decoding layer are used as part of the input data of the nth level of the first decoding layer. and In the first decoding layer of the nth level, feature extraction is performed on a portion of the input data of the first decoding layer of the nth level to obtain second-scale features. The second-scale features and the enhanced coding features output by the first coding layer corresponding to the first decoding layer of the nth level are fused to obtain fused features. Cross-attention interaction is performed on the fused features, the first text coding features and the first image coding features to obtain the enhanced decoding features output by the first decoding layer of the nth level.

10. The method according to any one of claims 1 to 9, wherein the first image enhancement prompt information includes first text prompt information and first image prompt information; The step of generating image restoration prompt information based on the first image enhancement prompt information to restore the enhanced image to the sample image includes: Based on the first text prompt information, a second text prompt information is generated to restore the enhanced image to the sample image, and the pixel value of each pixel in the first image prompt information is reversed to obtain the second image prompt information; and The second text prompt and the second image prompt are used as image restoration prompts.

11. The method according to any one of claims 1 to 10, wherein the initial image restoration model comprises a second cue information encoder, an image restoration encoder, a multi-scale restoration feature extraction network, and an image restoration decoder; the step of using the initial image restoration model, with the image restoration cue information as conditional information for guiding image restoration, to perform image restoration on the enhanced image to obtain a restored image includes: The image restoration prompt information is encoded by the second prompt information encoder to obtain the third encoding feature, and the enhanced image is encoded by the image restoration encoder in the initial image restoration model to obtain the fourth encoding feature; The multi-scale restoration feature extraction network extracts features from the fourth coding feature based on the third coding feature to obtain the restored image features; and The image restoration decoder decodes the features of the restored image to obtain the restored image.

12. The method according to claim 11, further comprising: The sample image is input into a pre-trained image description model, and the image content of the sample image is described by the pre-trained image description model to obtain the image content description information of the sample image; The image content description information is encoded using a pre-trained text encoder to obtain image content features; The sample image is input into the image enhancement encoder in the initial image enhancement model, and the sample image is encoded by the image enhancement encoder to obtain the fifth encoded feature; The second image enhancement feature is obtained by using the multi-scale enhancement feature extraction network in the initial image enhancement model to extract multi-scale features from the fifth encoded feature based on the image content features. By comparing the differences between the second image enhancement features and the restored image features, the content consistency loss value of the sample image is determined; and The step of adjusting the model parameters of the initial image enhancement model based on the sample training loss values ​​of the multiple sample images to obtain a trained image enhancement model includes: Based on the sample training loss value and content consistency loss value of each of the multiple sample images, the model parameters of the initial image enhancement model are adjusted to obtain a trained image enhancement model.

13. The method according to claim 11, further comprising: The first enhanced image features obtained through the initial image enhancement model are input into a pre-trained first reflection decoder for decoding to obtain a first reflection map, and the restored image features are input into a pre-trained second reflection decoder for decoding to obtain a second reflection map; By comparing the differences between the first and second reflection maps, the reflection map consistency loss value is obtained; The sample image is input into a pre-trained image decomposition model for decomposition to obtain a third reflection map of the sample image. The difference between the first reflection map and the third reflection map is compared to obtain the reconstruction loss value. Based on the reflectance map consistency loss value and the reconstruction loss value, the supervised loss value of the sample image is obtained; and The step of adjusting the model parameters of the initial image enhancement model based on the sample training loss values ​​of the multiple sample images to obtain a trained image enhancement model includes: Based on the sample training loss value and supervision loss value of each of the multiple sample images, the model parameters of the initial image enhancement model are adjusted to obtain a trained image enhancement model.

14. The method according to any one of claims 1 to 11, wherein the method further comprises: For each sample image, the sample image and the enhanced image are input into a first discriminator, and the first discriminator classifies the sample image and the enhanced image to obtain a first classification result; The sample image and the restored image are input into a second discriminator, which classifies the sample image and the restored image to obtain a second classification result. Based on the first classification result and the second classification result, the classification loss value of the sample image is obtained; and The step of adjusting the model parameters of the initial image enhancement model based on the sample training loss values ​​of the multiple sample images to obtain a trained image enhancement model includes: Based on the sample training loss value and classification loss value of each of the multiple sample images, the model parameters of the initial image enhancement model are adjusted to obtain a trained image enhancement model.

15. The method according to claim 14, wherein adjusting the model parameters of the initial image enhancement model based on the sample training loss value and classification loss value of each of the plurality of sample images to obtain a trained image enhancement model comprises: For each sample image, the sample image and the first image enhancement prompt information are input into the initial image restoration model. The initial image restoration model extracts multi-scale features from the sample image based on the first image enhancement prompt information to obtain the features of the image to be decoded. The features of the image to be decoded are then decoded to obtain the decoded image. By comparing the differences between the decoded image and the sample image, the image consistency loss value is determined, and the second image enhancement feature and reconstruction loss value of the sample image are obtained. Based on the image consistency loss value, the second image enhancement feature, the reconstruction loss value, and the features of the image to be decoded, the regularization loss value of the sample image is determined; The first loss value is obtained by weighted summation of the regularization loss value and the classification loss value; and Based on the sample training loss value and the first loss value of each of the multiple sample images, the model parameters of the initial image enhancement model are adjusted to obtain a trained image enhancement model.

16. The method according to claim 15, wherein adjusting the model parameters of the initial image enhancement model based on the sample training loss value and the first loss value of each of the plurality of sample images to obtain a trained image enhancement model comprises: Determine the supervised loss value and content consistency loss value of the sample image, and then superimpose the first loss value, the sample training loss value, the content consistency loss value, and the supervised loss value to obtain the second loss value of the sample image; Based on the second loss values ​​of each of the multiple sample images, the model parameters of the initial image enhancement model are adjusted to obtain a trained image enhancement model.

17. An image enhancement method performed by an electronic device, the method comprising: Obtain the image to be enhanced and the second image enhancement prompt information for the image to be enhanced; and The image to be enhanced and the second image enhancement prompt information are input into a trained image enhancement model. The trained image enhancement model uses the second image enhancement prompt information as conditional information to guide image enhancement, and performs image enhancement on the image to be enhanced to obtain an enhanced image. The trained image enhancement model is trained by the method described in any one of claims 1 to 16.

18. An image enhancement model training apparatus, the apparatus comprising: The sample data acquisition module is used to acquire multiple sample images and first image enhancement prompt information for each of the sample images; An image enhancement module is used to input the sample image and the first image enhancement prompt information into an initial image enhancement model for each sample image, and to enhance the sample image using the first image enhancement prompt information as condition information to guide image enhancement, thereby obtaining an enhanced image. The image restoration module is used to generate image restoration prompts to restore the enhanced image to the sample image based on the first image enhancement prompts, input the enhanced image and the image restoration prompts into an initial image restoration model, and restore the enhanced image to the sample image through the initial image restoration model, using the image restoration prompts as conditional information to guide image restoration, to obtain the restored image. The loss calculation module is used to compare the differences between the sample image and the restored image to determine the sample training loss value of the sample image; and The parameter adjustment module is used to adjust the model parameters of the initial image enhancement model based on the sample training loss values ​​of the multiple sample images, so as to obtain a trained image enhancement model.

19. A computer device comprising a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the method according to any one of claims 1 to 17.

20. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 17.

21. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1 to 17.