Image enhancement method and device and electronic equipment

The method enhances visible light images by combining visible and non-visible light images using a control network and feature fusion to improve lighting and geometric details, addressing poor image quality in low-light conditions.

CN120318134APending Publication Date: 2025-07-15HANGZHOU EZVIZ SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510336522.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Due to poor lighting, the visible light images obtained during image acquisition perform poorly in brightness, contrast and details, resulting in poor display results.

Method used

By acquiring auxiliary images of visible and non-visible images, image enhancement is performed using a pre-trained image enhancement model. The model includes a first control network and a second control network, and the diffusion model is controlled based on auxiliary feature contents of the illumination information dimension and spatial geometric information dimension, respectively, and the enhanced image is generated in combination with the feature fusion module.

Benefits of technology

The display effect of visible light images is improved, making them close to the image quality acquired under good lighting conditions, and meets the needs of image enhancement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318134A_ABST
    Figure CN120318134A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image enhancement method and device and electronic equipment, and relates to the technical field of image processing, and the method comprises the steps: obtaining a to-be-enhanced target visible light image, and obtaining a first auxiliary image belonging to a non-visible light image; determining a target image pair from a plurality of preset reference image pairs based on the first auxiliary image; extracting illumination information of a visible light image in the target image pair, and based on the extracted illumination information, determining auxiliary feature content corresponding to the target visible light image and related to an illumination information dimension as first auxiliary feature content; and inputting the target visible light image and the first auxiliary feature content into a pre-trained image enhancement model to obtain an enhanced image of the target visible light image. Through the scheme, the display effect of the visible light image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technologies, and particularly to an image enhancement method, apparatus, and electronic device. Background Art

[0002] In an image acquisition scenario, due to poor lighting, visible light images with poor display effects are often obtained during the image acquisition process. Such visible light images with poor display effects often perform poorly in one or more aspects such as brightness, contrast, and details.

[0003] In order to improve the display effect of images, there is a need to perform image enhancement processing on visible light images, so as to achieve the display effect of images acquired under good lighting conditions. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide an image enhancement method, apparatus, and electronic device to improve the display effect of visible light images. The specific technical solutions are as follows:

[0005] In a first aspect, the embodiments of this application provide an image enhancement method, and the method includes:

[0006] Obtain a target visible light image to be enhanced, and obtain a first auxiliary image belonging to a non-visible light image; wherein, the target visible light image and the first auxiliary image are images obtained by an image acquisition device performing image acquisitions on visible light and non-visible light respectively for the same scene content;

[0007] Based on the first auxiliary image, determine a target image pair from a preset plurality of reference image pairs; wherein, each reference image pair includes: a visible light image and a non-visible light image of the same scene content obtained by the image acquisition device performing image acquisitions on visible light and non-visible light respectively, and the visible light image in each reference image pair is an image acquired under a scene that meets the lighting conditions; the target image pair is: the reference image pair whose included non-visible light image matches the first auxiliary image;

[0008] Extract the lighting information of the visible light image in the target image pair, and based on the extracted lighting information, determine the auxiliary feature content corresponding to the target visible light image in terms of the lighting information dimension as the first auxiliary feature content;

[0009] Input the target visible light image and the first auxiliary feature content into a pre-trained image enhancement model to obtain an enhanced image of the target visible light image; wherein, the image enhancement model is a diffusion model for image generation equipped with a first control network, and the first control network is a network that controls the diffusion model according to the auxiliary feature content corresponding to the image to be enhanced and related to the dimension of illumination information.

[0010] Optionally, in one implementation, the image enhancement model is further provided with a second control network; wherein, the second control network is: a network that controls the diffusion model according to the auxiliary feature content corresponding to the image to be enhanced and related to the dimension of spatial geometric information;

[0011] The method further includes:

[0012] Extract the normal map of the first auxiliary image, and based on the extracted normal map, determine the auxiliary feature content corresponding to the target visible light image and related to the dimension of spatial geometric information as the second auxiliary feature content;

[0013] Input the second auxiliary feature content into a pre-trained image enhancement model.

[0014] Optionally, in one implementation, the image enhancement model further includes a feature fusion module; wherein, the feature fusion module is used to fuse the feature content output by the first control network as the control condition and the feature content output by the second control network as the control condition according to their respective fusion weights, and input the fused feature content into the diffusion model;

[0015] The diffusion model is specifically used to generate an enhanced image of the target visible light image based on the fused feature content;

[0016] The fusion weights of the first control network and the second control network are set according to the required image enhancement effect, and the enhancement effect includes: an effect where the enhancement of detail texture is higher than that of illumination or an effect where the enhancement of illumination is higher than that of detail texture.

[0017] Optionally, in one implementation, the image acquisition device is a rotatable acquisition device;

[0018] The determining of the target image pair from a preset plurality of reference image pairs based on the first auxiliary image includes:

[0019] Select the viewing angle range to which the viewing angle of the first auxiliary image belongs from a plurality of viewing angle ranges; wherein, the plurality of viewing angle ranges are different viewing angle ranges of the image acquisition device, and each viewing angle range corresponds to each reference image pair among the plurality of reference image pairs whose viewing angles belong to this viewing angle range;

[0020] Determine each pair of reference images corresponding to the selected viewing angle range;

[0021] Calculate the similarity between the first auxiliary image and the non-visible light images in each determined pair of reference images;

[0022] Based on the obtained similarities, determine a target pair of images from each of the determined pairs of reference images.

[0023] Optionally, in one implementation, the extracting the illumination information of the visible light image in the target pair of images includes:

[0024] Extract the illumination information of the entire visible light image in the target pair of images;

[0025] Or,

[0026] Perform extraction on the visible light image in the target pair of images with respect to the image background region to obtain an image background to be utilized, and extract the illumination component of the image background to be utilized as the illumination information;

[0027] Or,

[0028] Extract the illumination environment map of the visible light image in the target pair of images as the illumination information.

[0029] Optionally, in one implementation, before the determining, based on the extracted illumination information, the auxiliary feature content corresponding to the target visible light image with respect to the illumination information dimension as the first auxiliary feature content, the method further includes:

[0030] Extract the illumination information of the target visible light image;

[0031] The determining, based on the extracted illumination information, the auxiliary feature content corresponding to the target visible light image with respect to the illumination information dimension as the first auxiliary feature content includes:

[0032] Generate, based on the illumination information of the visible light image in the target pair of images and the illumination information of the target visible light image, the auxiliary feature content corresponding to the target visible light image with respect to the illumination information dimension as the first auxiliary feature content.

[0033] Optionally, in one implementation, the training process of the image enhancement model includes:

[0034] Obtain a training set; wherein, the training set includes a plurality of sample image pairs; each sample image pair includes a first image to be subjected to image enhancement and a second image that is the image to be obtained after enhancing the first image. The first image and the second image in each sample image pair are both visible light images, and the second image is an image collected under a scene that meets the lighting conditions;

[0035] For each sample image pair, extract the lighting information of the second image in the sample image pair. Based on the extracted lighting information, determine the auxiliary feature content corresponding to the first image in the sample image pair in the dimension of lighting information as the first feature content corresponding to the sample image pair;

[0036] For each sample image pair, extract the normal map of the first image in the sample image pair. Based on the extracted normal map, determine the auxiliary feature content corresponding to the first image in the sample image pair in the dimension of spatial geometric information as the second feature content corresponding to the sample image pair;

[0037] With the model parameters of the diffusion model in the image enhancement model fixed, train the first control network and the second control network of the image enhancement model based on each sample image pair in the training set and the corresponding first feature content and second feature content.

[0038] Optionally, in one implementation, the training of the first control network and the second control network of the image enhancement model based on each sample image pair in the training set and the corresponding first feature content and second feature content includes:

[0039] Select a plurality of sample image pairs from the training set to obtain a plurality of first sample image pairs to be utilized, and use the second images in the plurality of first sample image pairs as the ground truth content. Based on the first images and the corresponding first feature content in the selected plurality of first sample image pairs, train the first control network of the image enhancement model;

[0040] In response to the completion of the training of the first control network, select a plurality of sample image pairs from the training set to obtain a plurality of second sample image pairs to be utilized, and use the second images in the selected plurality of second sample image pairs as the ground truth content. Based on the first images and the corresponding second feature content in the selected plurality of second sample image pairs, train the second control network of the image enhancement model;

[0041] In response to the completion of the training of the second control network, multiple sample image pairs in the training set are selected to obtain multiple third sample image pairs. Using the second image in the selected multiple third sample image pairs as the ground truth content, the first control network of the image enhancement model is retrained based on the first image in the selected multiple third sample image pairs, the corresponding first feature content, and the second feature content.

[0042] In a second aspect, an embodiment of the present application provides an image enhancement device, which includes:

[0043] An image acquisition module, configured to acquire a target visible light image to be enhanced and acquire a first auxiliary image belonging to a non-visible light image; wherein, the target visible light image and the first auxiliary image are images obtained by an image acquisition device respectively performing image acquisition of visible light and non-visible light for the same scene content.

[0044] An image pair determination module, configured to determine a target image pair from a preset multiple reference image pairs based on the first auxiliary image; wherein each reference image pair includes: a visible light image and a non-visible light image of the same scene content obtained by the image acquisition device respectively performing image acquisition of visible light and non-visible light, and the visible light image in each reference image pair is an image acquired under a scene that meets the illumination conditions; the target image pair is: the reference image pair in which the included non-visible light image matches the first auxiliary image.

[0045] A first determination module, configured to extract the illumination information of the visible light image in the target image pair and determine, based on the extracted illumination information, auxiliary feature content corresponding to the target visible light image in the dimension of illumination information as the first auxiliary feature content.

[0046] An image enhancement module, configured to input the target visible light image and the first auxiliary feature content into a pre-trained image enhancement model to obtain an enhanced image of the target visible light image; wherein, the image enhancement model is a diffusion model for image generation provided with a first control network, and the first control network is a network for controlling the diffusion model according to the auxiliary feature content corresponding to the image to be enhanced in the dimension of illumination information.

[0047] Optionally, in an implementation, the image enhancement model is further provided with a second control network; wherein, the second control network is: a network for controlling the diffusion model according to the auxiliary feature content corresponding to the image to be enhanced in the dimension of spatial geometric information.

[0048] The device further includes:

[0049] A second determination module, configured to extract the normal map of the first auxiliary image, and based on the extracted normal map, determine the auxiliary feature content corresponding to the target visible light image in terms of the spatial geometric information dimension as the second auxiliary feature content;

[0050] An information input module, configured to input the second auxiliary feature content into a pre-trained image enhancement model.

[0051] Optionally, in one implementation, the image enhancement model further includes a feature fusion module; wherein, the feature fusion module is configured to fuse the feature content output by the first control network as a control condition and the feature content output by the second control network as a control condition according to their respective fusion weights, and input the fused feature content into the diffusion model;

[0052] The diffusion model is specifically configured to generate an enhanced image of the target visible light image based on the fused feature content;

[0053] The respective fusion weights of the first control network and the second control network are set according to the required image enhancement effect, and the enhancement effect includes: an effect where the detail texture enhancement is higher than the illumination enhancement or an effect where the illumination enhancement is higher than the detail texture enhancement.

[0054] Optionally, in one implementation, the image acquisition device is a rotatable acquisition device;

[0055] The image pair determination module is specifically configured to:

[0056] Select the view range to which the shooting view of the first auxiliary image belongs from multiple view ranges; wherein, the multiple view ranges are different view ranges of the image acquisition device, and each view range corresponds to each reference image pair among the multiple reference image pairs whose shooting views belong to this view range;

[0057] Determine the respective reference image pairs corresponding to the selected view range;

[0058] Calculate the similarity between the first auxiliary image and the non-visible light images in the determined respective reference image pairs;

[0059] Based on the obtained similarity, determine the target image pair from the determined respective reference image pairs.

[0060] Optionally, in one implementation, the first determination module is specifically configured to:

[0061] Extract the illumination information of the entire image of the visible light image in the target image pair;

[0062] Or,

[0063] Perform extraction on the visible light image in the target image pair with respect to the image background region to obtain the image background to be utilized, and extract the illumination component of the image background to be utilized as illumination information;

[0064] Or,

[0065] Extract the illumination environment map of the visible light image in the target image pair as illumination information.

[0066] Optionally, in one implementation manner, the apparatus further includes:

[0067] An extraction module, configured to extract the illumination information of the target visible light image before determining, based on the extracted illumination information, the auxiliary feature content corresponding to the target visible light image with respect to the illumination information dimension as the first auxiliary feature content;

[0068] The first determination module is specifically configured to:

[0069] Generate, based on the illumination information of the visible light image in the target image pair and the illumination information of the target visible light image, the auxiliary feature content corresponding to the target visible light image with respect to the illumination information dimension as the first auxiliary feature content.

[0070] Optionally, in one implementation manner, the following modules are used to train the image enhancement model:

[0071] An acquisition sub-module, configured to acquire a training set; wherein, the training set includes a plurality of sample image pairs; each sample image pair includes a first image to be subjected to image enhancement and a second image as the image required to be obtained after the enhancement of the first image, and the first image and the second image in each sample image pair are both visible light images and the second image is an image acquired under a scene that meets the illumination conditions;

[0072] A first determination sub-module, configured to, for each sample image pair, extract the illumination information of the second image in the sample image pair, and determine, based on the extracted illumination information, the auxiliary feature content corresponding to the first image in the sample image pair with respect to the illumination information dimension as the first feature content corresponding to the sample image pair;

[0073] A second determination sub-module, configured to, for each sample image pair, extract the normal map of the first image in the sample image pair, and determine, based on the extracted normal map, the auxiliary feature content corresponding to the first image in the sample image pair with respect to the spatial geometry information dimension as the second feature content corresponding to the sample image pair;

[0074] A model training sub-module, configured to train a first control network and a second control network of the image enhancement model based on each sample image pair in the training set, and the corresponding first feature content and second feature content, when the model parameters of the diffusion model in the image enhancement model are fixed.

[0075] Optionally, in one implementation, the model training sub-module is specifically configured to:

[0076] Select multiple sample image pairs from the training set to obtain multiple first sample image pairs to be utilized, and use the second image in the multiple first sample image pairs as the ground truth content. Based on the first image in the selected multiple first sample image pairs and the corresponding first feature content, train the first control network of the image enhancement model;

[0077] In response to the completion of the training of the first control network, select multiple sample image pairs from the training set to obtain multiple second sample image pairs to be utilized, and use the second image in the selected multiple second sample image pairs as the ground truth content. Based on the first image in the selected multiple second sample image pairs and the corresponding second feature content, train the second control network of the image enhancement model;

[0078] In response to the completion of the training of the second control network, select multiple sample image pairs from the training set to obtain multiple third sample image pairs, use the second image in the selected multiple third sample image pairs as the ground truth content, and based on the first image in the selected multiple third sample image pairs, the corresponding first feature content and second feature content, retrain the first control network of the image enhancement model.

[0079] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory for storing a computer program; a processor for implementing any of the image enhancement methods when executing the program stored in the memory.

[0080] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computing program is stored, and the computer program implements any of the image enhancement methods when executed by a processor.

[0081] An embodiment of the present application further provides a computer program product including instructions, which when running on a computer, causes the computer to execute any of the above-mentioned image enhancement methods.

[0082] Beneficial effects of the embodiments of the present application:

[0083] As can be seen above, in the solution provided by the embodiments of the present application, multiple reference image pairs can be obtained in advance. Each reference image pair includes: a visible light image and a non-visible light image of the same scene content obtained by an image acquisition device respectively performing image acquisition of visible light and non-visible light, and the visible light image in each reference image pair is an image acquired under a scene that meets the illumination conditions.

[0084] When performing image enhancement, an image obtained by the image acquisition device performing image acquisition of visible light for the same scene content can be obtained as the target visible light image to be enhanced, and an image obtained by performing image acquisition of non-visible light for the same scene content can be obtained as the first auxiliary image. Thus, based on the first auxiliary image, from the above multiple reference image pairs, a reference image pair whose included non-visible light image matches the first auxiliary image is determined as the target image pair. Then, by extracting the illumination information of the visible light image in the target image pair and based on the extracted illumination information, the auxiliary feature content corresponding to the target visible light image in the dimension of illumination information is determined as the first auxiliary feature content, and the target visible light image and the first auxiliary feature content are input into a pre-trained image enhancement model to obtain the enhanced image of the target visible light image.

[0085] Among them, the above image enhancement model is a diffusion model for image generation equipped with a first control network, and this first control network is a network that controls the diffusion model according to the auxiliary feature content in the dimension of illumination information corresponding to the image to be enhanced. That is to say, in the process of using the above image enhancement model for image enhancement, the above first control network guides the image enhancement process based on the obtained first auxiliary feature content, so that the obtained enhanced image is close to the image acquired under a scene that meets the illumination conditions, so as to meet the requirements for image enhancement processing of visible light images and achieve the display effect of images acquired under good illumination. BRIEF DESCRIPTION OF THE DRAWINGS

[0086] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other embodiments according to these drawings.

[0087] Figure 1 It is a schematic flowchart of an image enhancement method provided by an embodiment of the present application;

[0088] Figure 2 It is a schematic flowchart of another image enhancement method provided by an embodiment of the present application;

[0089] Figure 3 Schematic structural diagram of an image enhancement model provided by an embodiment of the present application;

[0090] Figure 4 Schematic flowchart of the training process of an image enhancement model provided by an embodiment of the present application;

[0091] Figure 5 Schematic flowchart of a specific embodiment provided by an embodiment of the present application;

[0092] Figure 6 Schematic structural diagram of a specific embodiment provided by an embodiment of the present application;

[0093] Figure 7 Schematic structural diagram of an image enhancement device provided by an embodiment of the present application;

[0094] Figure 8 Schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0095] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the scope of protection of the present application.

[0096] To facilitate the understanding of the solution provided by the embodiment of the present application, the professional terms involved in the embodiment of the present application will be introduced first.

[0097] Stable Diffusion: Abbreviated as SD, it is a diffusion model based on latent. It introduces text condition in U-net (U-shaped Network, a deep learning network model for image segmentation) to realize image generation based on text.

[0098] Control Net: It is a neural network structure that controls the diffusion model by adding additional conditions.

[0099] ORB (Oriented FAST and Rotated BRIEF) algorithm: A computer vision algorithm for feature point detection and description, specifically used for image feature extraction and matching. It combines the FAST (Feature from Accelerated Segment Test) feature detector and the BRIEF (Binary Robust Independent Elementary Features, a binary encoding method for describing feature points in images), and improves them to enable faster image processing while ensuring good rotational invariance and scale invariance.

[0100] To solve the above technical problems, the embodiments of the present application provide an image enhancement method, apparatus, and electronic device.

[0101] Among them, this method can be applicable to various application scenarios for image enhancement processing of visible light images. For example, image enhancement processing of visible light images collected by a camera at night, or image enhancement processing of visible light images collected by a mobile phone on a cloudy day, etc. And this method can be applied to various electronic devices such as laptop computers and servers, hereinafter referred to as electronic devices. The embodiments of the present application do not specifically limit the execution subject and application scenario of this method.

[0102] An image enhancement method provided by the embodiments of the present application may include the following steps:

[0103] Obtain a target visible light image to be enhanced, and obtain a first auxiliary image belonging to a non-visible light image; wherein, the target visible light image and the first auxiliary image are images obtained by an image acquisition device for respectively performing image acquisition of visible light and non-visible light on the same scene content.

[0104] Based on the first auxiliary image, determine a target image pair from a preset plurality of reference image pairs; wherein, each reference image pair includes: a visible light image and a non-visible light image of the same scene content obtained by the image acquisition device for respectively performing image acquisition of visible light and non-visible light, and the visible light image in each reference image pair is an image acquired under a scene that meets the illumination conditions; the target image pair is: the reference image pair in which the included non-visible light image matches the first auxiliary image.

[0105] Extract the illumination information of the visible light image in the target image pair, and based on the extracted illumination information, determine the auxiliary feature content corresponding to the target visible light image in terms of the illumination information dimension as the first auxiliary feature content.

[0106] Input the target visible light image and the first auxiliary feature content into a pre-trained image enhancement model to obtain an enhanced image of the target visible light image. Among them, the image enhancement model is a diffusion model for image generation with a first control network added, and the first control network is a network that controls the diffusion model according to the auxiliary feature content regarding the illumination information dimension corresponding to the image to be enhanced.

[0107] As can be seen above, in the solution provided by the embodiments of the present application, multiple reference image pairs can be obtained in advance. Each reference image pair includes: a visible light image and a non-visible light image of the same scene content obtained by an image acquisition device respectively performing image acquisition regarding visible light and non-visible light, and the visible light image in each reference image pair is an image acquired under a scene that meets the illumination conditions.

[0108] When performing image enhancement, an image obtained by an image acquisition device performing image acquisition regarding visible light for the same scene content can be obtained as the target visible light image to be enhanced, and an image obtained by performing image acquisition regarding non-visible light for the same scene content can be obtained as the first auxiliary image. Thus, based on the first auxiliary image, from the above multiple reference image pairs, a reference image pair whose included non-visible light image matches the first auxiliary image is determined as the target image pair. Then, by extracting the illumination information of the visible light image in the target image pair and based on the extracted illumination information, the auxiliary feature content regarding the illumination information dimension corresponding to the target visible light image is determined as the first auxiliary feature content, and the target visible light image and the first auxiliary feature content are input into a pre-trained image enhancement model to obtain an enhanced image of the target visible light image.

[0109] Among them, the above image enhancement model is a diffusion model for image generation with a first control network added, and moreover, the first control network is a network that controls the diffusion model according to the auxiliary feature content regarding the illumination information dimension corresponding to the image to be enhanced. That is to say, in the process of using the above image enhancement model for image enhancement, the first control network guides the image enhancement process based on the obtained first auxiliary feature content, so that the obtained enhanced image is close to the image acquired under a scene that meets the illumination conditions, to meet the requirement of image enhancement processing for visible light images and achieve the display effect of images acquired under good illumination.

[0110] Next, in combination with the accompanying drawings, a specific description of an image enhancement method provided by the embodiments of the present application will be given.

[0111] Figure 1 It is a schematic flowchart of an image enhancement method provided by the embodiments of the present application. As Figure 1 shown, the method includes:

[0112] S101: Obtain a target visible light image to be enhanced, and obtain a first auxiliary image belonging to a non-visible light image;

[0113] Among them, the target visible light image and the first auxiliary image are images obtained by an image acquisition device respectively performing visible light and non-visible light image acquisitions on the same scene content.

[0114] In this application, when performing image enhancement, a target visible light image to be enhanced can be obtained. Exemplarily, the above target visible light image can be referred to as a low-light image collected in a low-light environment, that is, the target visible light image is an image with a poor image display effect due to insufficient light.

[0115] Then, take the non-visible light image that is collected simultaneously with the target visible light image and has the same scene content as the target visible light image as the first auxiliary image of the target visible light image. That is, the obtained target visible light image and the first auxiliary image are images obtained by an image acquisition device respectively performing visible light and non-visible light image acquisitions on the same scene content, and it can be understood that these two images are the same image belonging to two different image types. For example, for the object p in scene P, the two images are a visible light image and a non-visible light image respectively captured from the same shooting angle.

[0116] Among them, the image acquisition device for visible light acquisition and non-visible light acquisition can be the same device or different devices; when the image acquisition devices for visible light acquisition and non-visible light acquisition are different devices, it is necessary to ensure that the two devices can capture the same scene content.

[0117] Exemplarily, when the image acquisition device for visible light acquisition and non-visible light acquisition is the same device, the image acquisition device can be an RGB-NIR (Red, Green, Blue-Near Infrared) camera; when they are different devices, the device for visible light acquisition can be an RGB camera, and the image acquisition device for non-visible light acquisition can be an NIR camera, etc., which are all reasonable.

[0118] S102: Based on the first auxiliary image, determine a target image pair from a plurality of preset reference image pairs;

[0119] Among them, each reference image pair includes: a visible light image and a non-visible light image of the same scene content obtained by an image acquisition device respectively performing image acquisition on visible light and non-visible light, and the visible light image in each reference image pair is an image acquired under a scene that meets the illumination conditions; the target image pair is: a reference image pair in which the included non-visible light image matches the first auxiliary image.

[0120] In this application, multiple reference image pairs can be obtained in advance. Each reference image pair includes a visible light image and a non-visible light image of the same scene content obtained by an image acquisition device respectively performing image acquisition on visible light and non-visible light. Considering that when performing image enhancement later, the illumination information of the visible light image in the determined target image pair is used as the determination benchmark for the auxiliary feature content, and the purpose of image enhancement is to make the enhanced image achieve the display effect of an image acquired under good illumination conditions. Therefore, the visible light image in each obtained reference image pair is an image under a scene that meets the illumination conditions. The above-mentioned scene that meets the illumination conditions can be understood as a scene with good illumination conditions. There is no problem of poor image display effect caused by insufficient illumination in the image acquired under this scene. Exemplarily, the above-mentioned scene that meets the illumination conditions can be a scene during the day with sufficient illumination. The image acquired under this scene does not need to be enhanced in terms of illumination dimension; and the image acquired under the above-mentioned scene that meets the illumination conditions can be an image that is not affected by insufficient illumination and has a good image display effect without the need for enhancement.

[0121] Thus, based on the first auxiliary image, a reference image pair in which the included non-visible light image matches the first auxiliary image can be determined from the above-mentioned multiple reference image pairs as the target image pair.

[0122] Optionally, determine the similarity between the first auxiliary image and the non-visible light image in each reference image pair, and based on the obtained similarity, determine the target image pair from the above-mentioned multiple reference image pairs.

[0123] Optionally, perform feature extraction on the first auxiliary image to obtain the image features of the first auxiliary image, and through feature matching, determine a reference image pair in which the image features of the included non-visible light image match the first auxiliary image from the above-mentioned multiple reference image pairs as the target image pair.

[0124] Specifically, perform image preprocessing on the first auxiliary image. For example, perform processing such as image denoising and enhancing image contrast to improve the accuracy of feature extraction. Then, use a predetermined feature point detection algorithm to detect feature points in the processed first auxiliary image, and perform feature description on the detected feature points to obtain feature description information. After that, use a predetermined feature matching algorithm to match the feature description information corresponding to the first auxiliary image with the feature description information corresponding to the non-visible light images in multiple reference image pairs to determine the reference image pairs including non-visible light images that match the first auxiliary image as target image pairs.

[0125] Among them, the above-mentioned predetermined feature point detection algorithm may include the SURF algorithm (Speeded Up Robust Features, a robust image recognition and description algorithm), the ORB algorithm, etc. The above-mentioned algorithm for feature description may be the ORB algorithm, FREAK (Fast Retina Key point), etc. The above-mentioned predetermined feature matching algorithm may be the nearest neighbor matching or the fast matching algorithm based on the feature tree, etc. In this regard, the embodiments of the present application do not make specific limitations.

[0126] Considering that there may be multiple matching reference image pairs during feature matching, and even incorrect matching reference image pairs, optionally, use a predetermined verification algorithm to verify the matched reference image pairs to remove the incorrect matching reference image pairs and improve the accuracy of the determined target image pairs. Among them, the above-mentioned predetermined detection algorithm may be the RANSAC (RANdom SAmple Consensus) algorithm, etc. In this regard, the embodiments of the present application do not make specific limitations.

[0127] It should be noted that the feature extraction process for the non-visible light images in the reference image pairs is similar to that for the first auxiliary image and will not be elaborated here.

[0128] S103: Extract the illumination information of the visible light image in the target image pair, and based on the extracted illumination information, determine the auxiliary feature content corresponding to the target visible light image in terms of the illumination information dimension as the first auxiliary feature content.

[0129]

[0130] ​In this application, after determining the target image pair, the visible light image in the target image pair can be obtained. Thus, the illumination information of the obtained visible light image is extracted. Then, the extracted illumination information is encoded, and the encoded data is determined as the auxiliary feature content regarding the illumination information dimension corresponding to the target visible light image, that is, the data content obtained by encoding the illumination information is determined as the first auxiliary feature content.

[0131] Optionally, after extracting the illumination information of the visible light image, the extracted illumination information is encoded using an illumination feature encoder, and the data content obtained after encoding is determined as the first auxiliary feature content.

[0132] Optionally, in a specific implementation, in the above step S103, extracting the illumination information of the visible light image in the target image pair may include the following steps:

[0133] Step A1: Extract the illumination information of the entire visible light image in the target image pair.

[0134] In this specific implementation, when extracting the illumination information of the visible light image in the target image pair, the entire visible light image can be used as the extraction object, and the illumination information of the entire extracted image is used as the illumination information of the visible light image.

[0135] In this specific implementation, taking the entire visible light image as the extraction object for extracting illumination information can improve the integrity of the extracted illumination information, so as to provide comprehensive illumination information about the visible light image for subsequent determination of the first auxiliary feature content.

[0136] Optionally, in a specific implementation, in the above step S103, extracting the illumination information of the visible light image in the target image pair may include the following steps:

[0137] Step A2: Perform extraction on the background region of the visible light image in the target image pair to obtain the image background to be utilized, and extract the illumination component of the image background to be utilized as the illumination information.

[0138] In a specific application, compared with the illumination information of the foreground region in the image, the illumination information of the background region in the image is relatively stable. That is to say, in the images collected from the same shooting angle, the illumination information of the foreground region in the image is constantly changing. For example, the colors of the objects appearing in the captured image of the image are different, resulting in large differences in the illumination information included in the foreground region, while the illumination information included in the background region can be considered fixed and unchanged. When performing image enhancement, it can provide effective illumination information assistance and has high guiding ability.

[0139] Therefore, in this specific implementation, based on the characteristic that the light information contained in the background region is relatively stable, when extracting the light information of the visible light image in the target image pair, the background region of the image can be extracted from the visible light image to obtain the image background to be utilized. Thus, the light component extracted from the image background to be utilized is used as the light information of the visible light image.

[0140] Optionally, a Gaussian mixture model is used to distinguish the image background region and the moving target in the visible light image, and the image background region in the visible light image is extracted.

[0141] Optionally, BiRefNet (image segmentation model) for performing foreground and background separation operations is used to determine and extract the image background in the visible light image, and the light information of the image background is calculated to obtain the light information of the visible light image. Among them, the Retinex algorithm can be used to calculate the light information of the image background. Specifically, the Retinex algorithm decomposes the region image corresponding to the above image background into a reflection component map and a light component map. The light component map represents the brightness information generated by the light source irradiation in the image, while the reflection component map reflects the true color and texture of the object surface in the image and is not affected by the light change. Therefore, after the light component map is decomposed, the data information in the light component map can be used as the light information of the visible light image.

[0142] In this specific implementation, combined with the characteristic that the light information of the image background region is relatively stable, the light component extracted from the image background in the visible light image is used as the light information of the visible light image to improve the stability of the extracted light information, and further improve the guiding ability of the determined first auxiliary feature content in the image enhancement process.

[0143] Optionally, in a specific implementation, in the above step S103, extracting the light information of the visible light image in the target image pair may include the following steps:

[0144] Step A3: Extract the light environment map of the visible light image in the target image pair as the light information.

[0145] In specific applications, the light environment map is a mathematical description information of the light intensity distribution in the entire scene. That is to say, the data information contained in the light environment map can reflect the brightness distribution characteristics of different regions in the image scene in the form of data, thereby more standardly representing the light information in the image.

[0146] In this specific implementation, after obtaining the visible light image in the target image pair, the light environment map of the visible light image can be extracted, and the light environment map is used as the light information of the visible light image.

[0147] Among them, the process of extracting the illumination environment map from the visible light image can be understood as the process of recovering the information of the environmental light source from the image data of the visible light image. Exemplarily, a preset deep learning model is used to obtain the normal map of the visible light image, and based on the reflectivity of the light in the image, the albedo map of the visible light image is obtained by using the Retinex algorithm. Then, the visible light image is divided by the above reflectivity to calculate the shadow map, and the shadow map is projected backward onto the Normal Sphere, that is, through the normal direction of the pixels in the shadow map, the shadow map is projected backward onto the normal sphere, and the high dynamic range (HDR, High Dynamic Range Imaging, also known as high dynamic range imaging) of the image is maintained; among them, on the normal sphere, if multiple values are mapped to the same point at a certain position, the average value is taken. After completing the above steps, the direction and color data (in HDR form) on the normal sphere are mapped onto the Environment Map to obtain the illumination environment map.

[0148] Optionally, an open-source Diffusion light (diffusion illumination) model is used to extract the illumination environment map.

[0149] It should be noted that the embodiments of the present application do not specifically limit the method for extracting the illumination environment map from the image, and any method capable of extracting the illumination environment map from the image is applicable to the present application.

[0150] In this specific implementation manner, by virtue of the characteristic that the illumination environment map itself reflects the brightness distribution characteristics of different regions in the scene of the image in the form of data, the accuracy of the determined illumination information can be improved.

[0151] S104: Input the target visible light image and the first auxiliary feature content into the pre-trained image enhancement model to obtain the enhanced image of the target visible light image;

[0152] Among them, the image enhancement model is a diffusion model for image generation equipped with a first control network, and the first control network is a network that controls the diffusion model according to the auxiliary feature content corresponding to the image to be enhanced in the dimension of illumination information.

[0153] An image enhancement model provided by this application is a diffusion model for image generation with a first control network added. Among them, the first control network is a network that controls the diffusion model according to the auxiliary feature content regarding the dimension of illumination information corresponding to the image to be enhanced. That is to say, in the process of using the image enhancement model for image enhancement, through the above-mentioned first control network based on the obtained first auxiliary feature content, the image enhancement process is guided in terms of the dimension of illumination information, so that the display effect of the obtained enhanced image conforms to the display effect actually required by the user. This application does not limit the structure of the diffusion model, and any diffusion model with any structure for image generation existing in related technologies can be applied to the embodiments of this application.

[0154] Therefore, after obtaining the first auxiliary feature content, the target visible light image and the first auxiliary feature content are input into the pre-trained image enhancement model, and the first auxiliary feature content is used as the control condition of the first control network to guide the enhancement process of the entire image enhancement model for the target visible light image, so that the enhanced image of the target visible light image is close to the image collected under the scene that conforms to the illumination conditions, meets the requirement of image enhancement processing for visible light images, and achieves the display effect of the image collected under good illumination.

[0155] Among them, for the sake of convenient writing, for the specific implementation manner of the training process of the above image enhancement model, please refer to the following steps S401 - S404, which will not be elaborated here.

[0156] As can be seen above, in the solution provided by the embodiments of this application, multiple reference image pairs can be obtained in advance. Each reference image pair includes: a visible light image and a non-visible light image of the same scene content obtained by an image acquisition device respectively collecting images of visible light and non-visible light, and the visible light image in each reference image pair is an image collected under a scene that conforms to the illumination conditions.

[0157] When performing image enhancement, an image obtained by an image acquisition device through visible light image acquisition for the same scene content can be acquired as the target visible light image to be enhanced, and an image obtained by non-visible light image acquisition for the same scene content can be acquired as the first auxiliary image. Thus, based on the first auxiliary image, from the above-mentioned multiple reference image pairs, a reference image pair including a non-visible light image that matches the first auxiliary image is determined as the target image pair. Then, by extracting the illumination information of the visible light image in the target image pair and based on the extracted illumination information, the auxiliary feature content corresponding to the target visible light image in the dimension of illumination information is determined as the first auxiliary feature content, and the target visible light image and the first auxiliary feature content are input into a pre-trained image enhancement model to obtain the enhanced image of the target visible light image.

[0158] Among them, the above-mentioned image enhancement model is a diffusion model for image generation equipped with a first control network, and this first control network is a network that controls the diffusion model according to the auxiliary feature content in the dimension of illumination information corresponding to the image to be enhanced. That is to say, in the process of using the above-mentioned image enhancement model for image enhancement, the above-mentioned first control network guides the image enhancement process based on the obtained first auxiliary feature content, so that the obtained enhanced image is close to the image acquired under a scene that meets the illumination conditions, so as to meet the requirements for image enhancement processing of visible light images and achieve the display effect of images acquired under good illumination.

[0159] Optionally, in one embodiment, the image enhancement model is further provided with a second control network; among them, the second control network is: a network that controls the diffusion model according to the auxiliary feature content corresponding to the image to be enhanced in the dimension of spatial geometric information; Figure 2 The flowchart of another image enhancement method provided by the embodiments of the present application is shown as Figure 2 shown, and this method further includes the following steps:

[0160] S201: Extract the normal map of the first auxiliary image, and based on the extracted normal map, determine the auxiliary feature content corresponding to the target visible light image in the dimension of spatial geometric information as the second auxiliary feature content;

[0161] S202: Input the second auxiliary feature content into the pre-trained image enhancement model.

[0162] In specific applications, the purpose of image enhancement is to improve the visual effect of the image, making the enhanced image easier to observe with the human eye or facilitating subsequent analysis and processing by an image processing system. Usually, when performing image enhancement, parameters such as the brightness and contrast of the image can be adjusted. In some application scenarios, it is also necessary to emphasize specific information or features in the image, such as spatial geometric information like edges and textures in the scene content of the image, to make the enhanced image clearer.

[0163] In this embodiment, on the basis of being additionally provided with a first control network, the pre-trained image enhancement model is further provided with a second control network, which is used to control the diffusion model according to the auxiliary feature content corresponding to the image to be enhanced and regarding the dimension of spatial geometric information; that is, in the process of using the image enhancement model to perform image enhancement, through the above-mentioned first control network, based on the obtained first auxiliary feature content, the image enhancement process is guided in terms of the dimension of illumination information, and through the above-mentioned second control network, based on the obtained second auxiliary feature content, the image enhancement process is guided in terms of the dimension of spatial geometric information, so as to further improve the display effect of the enhanced image.

[0164] Therefore, after obtaining the first auxiliary image, extract the normal map of the first auxiliary image, and then encode the extracted normal map, and determine the encoded data as the auxiliary feature content corresponding to the target visible light image and regarding the dimension of spatial geometric information, that is, determine the data content of the feature information encoded in the normal map as the second auxiliary feature content, and then, while inputting the target visible light image and the first auxiliary feature content into the pre-trained image enhancement model, also input the second auxiliary feature content into the pre-trained image enhancement model. In this way, when the image enhancement model performs image enhancement on the target visible light image, it can combine the first auxiliary feature content and the second auxiliary feature content at the same time to perform image enhancement on the target visible light image, so as to improve the display effect of the enhanced image of the target visible light image.

[0165] Optionally, after extracting the normal map of the first auxiliary image, use a normal feature encoder to encode the feature information in the extracted normal map, and determine the encoded data content as the second auxiliary feature content.

[0166] In this embodiment, through the second control network additionally added to the image enhancement model, and using the second auxiliary feature content as the control condition of the second control network, a guidance regarding the dimension of spatial geometric information is added to the entire image enhancement process. Thus, guidance of auxiliary feature content in multiple dimensions is provided for the image enhancement process, so as to further improve the display effect of the enhanced image.

[0167] Optionally, in one embodiment, Figure 3 FIG. Figure 3 is a schematic structural diagram of an image enhancement model provided by an embodiment of the present application. The image enhancement model further includes a feature fusion module 303;

[0168] The feature fusion module 303 is configured to fuse the feature content used as a control condition output by the first control network 301 and the feature content used as a control condition output by the second control network 302 according to their respective fusion weights, and input the fused feature content into the diffusion model 304;

[0169] The diffusion model 304 is specifically configured to generate an enhanced image of the target visible light image based on the fused feature content;

[0170] The respective fusion weights of the first control network 301 and the second control network 302 are set according to the required image enhancement effect. The enhancement effects include: the effect of enhancing detail texture is higher than the effect of enhancing illumination, or the effect of enhancing illumination is higher than the effect of enhancing detail texture.

[0171] In this embodiment, considering that there are auxiliary guides in multiple dimensions during the image enhancement process, different emphases on different dimensions will result in different image enhancement effects. Specifically: the effect of enhancing detail texture is higher than the effect of enhancing illumination, or the effect of enhancing illumination is higher than the effect of enhancing detail texture.

[0172] Therefore, the respective fusion weights of the first control network and the second control network can be set in advance according to the required image enhancement effect of the user. For example, when the fusion weight of the first control network is greater than the fusion weight of the second control network, the achievable enhancement effect is that the effect of enhancing illumination is higher than the effect of enhancing detail texture; when the fusion weight of the first control network is less than the fusion weight of the second control network, the achievable enhancement effect is that the effect of enhancing detail texture is higher than the effect of enhancing illumination; when the fusion weight of the first control network is equal to the fusion weight of the second control network, the display effect of the enhanced image will not emphasize the enhancement of detail texture or illumination, but will equally enhance the detail texture and illumination of the image.

[0173] Thus, after the feature fusion module obtains the feature content used as a control condition output by the first control network and the feature content used as a control condition output by the second control network, it can fuse the two feature contents according to the respective fusion weights of the first control network and the second control network, and then input the fused feature content into the diffusion model, so that the diffusion model generates an enhanced image of the target visible light image based on the fused feature content.

[0174] In this embodiment, through the setting of the fusion weights, when there is auxiliary feature content in multiple dimensions, according to the needs of the user, the enhancement effect of the image can focus on enhancing the detail texture more than the lighting enhancement effect, or, focus on enhancing the lighting more than the detail texture enhancement effect. Thus, the flexibility of the image enhancement model during image enhancement is improved to meet the requirements for image enhancement processing of visible light images and achieve the display effect of images collected under good lighting.

[0175] Optionally, in one embodiment, the image acquisition device is a rotatable acquisition device; the above step S102, based on the first auxiliary image, determining a target image pair from a preset plurality of reference image pairs may include the following steps:

[0176] Step B1: Select the viewing angle range to which the viewing angle of the first auxiliary image belongs from a plurality of viewing angle ranges;

[0177] Wherein, the plurality of viewing angle ranges are different viewing angle ranges of the image acquisition device, and each viewing angle range corresponds to each reference image pair among the plurality of reference image pairs whose viewing angle belongs to this viewing angle range;

[0178] Step B2: Determine the respective reference image pairs corresponding to the selected viewing angle range;

[0179] Step B3: Calculate the similarity between the first auxiliary image and the non-visible light images in the respective reference image pairs determined;

[0180] Step B4: Based on the obtained similarity, determine a target image pair from the respective reference image pairs determined.

[0181] In this embodiment, the image acquisition device is a rotatable acquisition device, that is to say, the viewing angle of the image acquired by this image acquisition device is variable within a certain area range, that is, each image acquired by this image acquisition device may be multiple images within a viewing angle range or multiple images in different viewing angle ranges.

[0182] Therefore, when determining a target image pair from a preset plurality of reference image pairs based on the first auxiliary image, first, the viewing angle range to which the viewing angle of this first auxiliary image belongs can be determined, and thus, select the viewing angle range to which the viewing angle of this first auxiliary image belongs from the plurality of viewing angle ranges, wherein each viewing angle range corresponds to each reference image pair among the plurality of reference image pairs whose viewing angle belongs to this viewing angle range. Therefore, after determining the respective reference image pairs corresponding to the selected viewing angle range, the similarity between the first auxiliary image and the non-visible light images in the respective reference image pairs determined can be calculated, and thus, based on the obtained similarity, determine a target image pair from the respective reference image pairs determined.

[0183] Optionally, a similarity threshold is preset, and the reference image pairs with a similarity higher than the similarity threshold are determined as target image pairs.

[0184] Optionally, based on the obtained similarity, the reference image pair with the highest similarity is directly determined as the target image pair.

[0185] Optionally, the above-mentioned multiple reference image pairs are stored in an image database or data storage space established in advance for the image acquisition device, and the collected reference image pairs are grouped and stored according to different viewing angle ranges. Among them, the above-mentioned image database or data storage space can be set in the image acquisition device itself or in other devices capable of communicating with the image acquisition device. In this regard, the embodiments of the present application do not make specific limitations.

[0186] Optionally, the above-mentioned image database or data storage space may also store the feature description information corresponding to the non-visible light images in the reference image pairs. In this way, when determining the reference image pair including the non-visible light image and the first auxiliary image from the above-mentioned multiple reference image pairs based on the first auxiliary image, it is not necessary to perform the determination operation of the feature description information for the non-visible light images in each reference image pair, and the associated feature description information can be directly obtained, further improving the determination efficiency of the target image pair.

[0187] Among them, the above-mentioned viewing angle range can be divided according to the Euler angle state of the camera when the image acquisition device acquires images, or the viewing angle range of the camera can be mapped to the spherical coordinate system, and the viewing angle range to which the images acquired by the camera belong can be reversely determined through data division in the spherical coordinate system, etc. In this regard, the embodiments of the present application do not make specific limitations.

[0188] In this embodiment, considering that the image acquisition device is a rotatable acquisition device, when such an acquisition device acquires images, the acquired images do not have a fixed shooting angle. Therefore, when determining the target image pair, the viewing angle range to which the shooting angle of the first auxiliary image belongs can be first determined, and then the target image pair can be determined from the reference image pairs corresponding to this viewing angle range, thereby reducing the computing resources and time costs required for similarity calculation. Moreover, the illumination information included in the visible light images within the same viewing angle range can provide effective reference information for subsequent determination of the first auxiliary feature content, thereby improving the effectiveness and rationality of the first auxiliary feature content.

[0189] Optionally, in one embodiment, in an image enhancement method provided by the present application, before the above step S103 of determining the auxiliary feature content corresponding to the target visible light image in the dimension of illumination information based on the extracted illumination information as the first auxiliary feature content, the following steps are further included:

[0190] Step C1: Extract the illumination information of the target visible light image;

[0191] In the above step S103, based on the extracted illumination information, determining the auxiliary feature content corresponding to the target visible light image in the dimension of illumination information as the first auxiliary feature content may include the following steps:

[0192] Step C2: Generate the auxiliary feature content corresponding to the target visible light image in the dimension of illumination information based on the illumination information of the visible light image in the target image pair and the illumination information of the target visible light image, and use it as the first auxiliary feature content.

[0193] In this embodiment, considering that the enhancement effect required by the user may be: the illumination effect is enhanced, but the display effect of the enhanced image does not need to conform to the display effect of the image collected under the illumination conditions. For example, the target visible light image is a low-illumination image. After enhancing this image, the image is still a low-illumination image, but the overall display effect of the image is higher than that of the original image. However, the visible light image in the determined target image pair is an image collected under the illumination conditions that meet the requirements. Therefore, when using only the illumination information of the visible light image in the target image pair to guide the image enhancement process, the display effect of the enhanced image will not reach the enhancement effect required by the user.

[0194] Based on this, the illumination information of the target visible light image itself can be extracted, and based on its own illumination information and the illumination information of the visible light image in the target image pair, the auxiliary feature content corresponding to the target visible light image in the dimension of illumination information is generated as the first auxiliary feature content.

[0195] That is to say, after extracting the illumination information of the target visible light image itself and the illumination information of the visible light image in the target image pair, the two types of illumination information are fused and encoded, and the encoded data is determined as the first auxiliary feature content.

[0196] Optionally, after extracting the first illumination information of the target visible light image itself and the second illumination information of the visible light image in the target image pair, the first illumination information and the second illumination information are input into the illumination feature encoder, so that the illumination feature encoder fuses and encodes the first illumination information and the second illumination information, and determines the encoded data content as the first auxiliary feature content.

[0197] In this embodiment, through the fusion of the illumination information of the target visible light image with the illumination information of the visible light image in the target image pair, the generated first auxiliary feature content will not enhance the display effect of the enhanced image to the display effect of the image captured under the illumination condition that meets the illumination condition during the subsequent image enhancement process, so as to meet the enhanced effect required by the user.

[0198] Optionally, in one embodiment, Figure 4 It is a schematic flowchart of the training process of an image enhancement model provided by an embodiment of the present application. The training process may include the following steps:

[0199] S401: Obtain a training set;

[0200] Among them, the training set includes multiple sample image pairs; each sample image pair includes a first image to be enhanced, and a second image that is the image required to be obtained after the enhancement of the first image. The first image and the second image in each sample image pair are both visible light images, and the second image is an image captured under a scene that meets the illumination condition;

[0201] S402: For each sample image pair, extract the illumination information of the second image in the sample image pair. Based on the extracted illumination information, determine the auxiliary feature content corresponding to the first image in the sample image pair in the dimension of illumination information as the first feature content corresponding to the sample image pair;

[0202] S403: For each sample image pair, extract the normal map of the first image in the sample image pair. Based on the extracted normal map, determine the auxiliary feature content corresponding to the first image in the sample image pair in the dimension of spatial geometric information as the second feature content corresponding to the sample image pair;

[0203] S404: With the model parameters of the diffusion model in the image enhancement model fixed, based on each sample image pair in the training set, as well as the corresponding first feature content and second feature content, train the first control network and the second control network of the image enhancement model.

[0204] In this embodiment, the obtained training set includes multiple pairs of sample image pairs, and each sample image pair includes a first image and a second image that belong to visible light images; among them, the first image is the image to be enhanced, and the second image is the image required to be obtained after the enhancement of the first image. For example, the first image and the second image are images taken of the same scene content, but the first image is an image captured when the light is poor, that is, a low-light image, while the second image is an image captured when the light is good, and the second image can be understood as an image captured under a scene that meets the illumination condition, that is, a good-illumination image.

[0205] Then, for each pair of sample images, extract the illumination information of the second image in the pair of sample images. Then, encode the extracted illumination information, and determine the data obtained by encoding as the auxiliary feature content corresponding to the first image in the pair of sample images in terms of the illumination information dimension. That is, determine the data content after encoding the illumination information as the first feature content corresponding to the pair of sample images.

[0206] After that, for each pair of sample images, extract the normal map of the first image in the pair of sample images. Then, encode the extracted normal map, and determine the data obtained by encoding as the auxiliary feature content corresponding to the first image in the pair of sample images in terms of the spatial geometric information dimension. That is, determine the data content after encoding the feature information in the normal map as the second feature content corresponding to the pair of sample images.

[0207] Thus, with the model parameters of the diffusion model in the image enhancement model fixed, based on each pair of sample images in the training set, as well as the corresponding first feature content and second feature content, train the first control network and the second control network of the image enhancement model to obtain a trained image enhancement model.

[0208] Optionally, in one implementation, in the above step S404, training the first control network and the second control network of the image enhancement model based on each pair of sample images in the training set, as well as the corresponding first feature content and second feature content, may include the following steps:

[0209] Step D1: Select multiple pairs of sample images from the training set to obtain multiple first sample image pairs to be utilized, and use the second image in the multiple first sample image pairs as the ground truth content. Based on the first images in the selected multiple first sample image pairs and the corresponding first feature content, train the first control network of the image enhancement model;

[0210] Step D2: In response to the completion of the training of the first control network, select multiple pairs of sample images from the training set to obtain multiple second sample image pairs to be utilized, and use the second image in the selected multiple second sample image pairs as the ground truth content. Based on the first images in the selected multiple second sample image pairs and the corresponding second feature content, train the second control network of the image enhancement model;

[0211] Step D3: In response to the completion of the training of the second control network, multiple sample image pairs in the training set are selected to obtain multiple third sample image pairs. Using the second images in the selected multiple third sample image pairs as the ground truth content, the first control network of the image enhancement model is retrained based on the first images in the selected multiple third sample image pairs, the corresponding first feature content, and the second feature content.

[0212] In this implementation, multiple sample image pairs in the training set are selected to obtain multiple first sample image pairs to be utilized. Then, for each first sample image pair, the first image in the sample image pair and the corresponding illumination information are input into the image enhancement model. The first control network takes the output content determined based on the illumination information corresponding to the first image and used as the first feature content as the control condition and inputs it into the diffusion model. In this way, the diffusion model processes the received first image based on the received control condition to obtain the model output result corresponding to the sample image pair. Thus, based on the differences between the second images in each first sample image pair and the corresponding model output results, a first type of loss value is calculated, and when it is determined based on the first type of loss value that the image enhancement model has not converged, with the model parameters of the diffusion model in the image enhancement model fixed, the network parameters of the first control network are adjusted.

[0213] Optionally, to improve the control effect (guidance effect) of the first control network of the image enhancement model during the inference stage, during the training process of the first control network, an intermediate convergence value is set. That is, when the convergence value of the first control network reaches the intermediate convergence value, the foreground and background separation operation can be performed on the first and second images in each first sample image pair that did not participate in the training process. Using the background images of the first image and the second image as a new first sample image pair, the first control network is trained.

[0214] After the first control network converges and the training is completed, second sample image pairs to be utilized are selected from the multiple sample image pairs in the above training set. Then, for each second sample image pair, the first image in the sample image pair and the corresponding normal map are input into the image enhancement model. The second control network takes the output content determined based on the normal map corresponding to the first image and used as the second feature content as the control condition and inputs it into the diffusion model. In this way, the diffusion model processes the received first image based on the received control condition to obtain the model output result corresponding to the sample image pair. Thus, based on the differences between the second images in each second sample image pair and the corresponding model output results, a second type of loss value is calculated, and when it is determined based on the second type of loss value that the image enhancement model has not converged, with the model parameters of the diffusion model in the image enhancement model fixed, the network parameters of the second control network are adjusted.

[0215] Optionally, to facilitate the training of the second control network, a pre-trained image-to-text model is used to generate description information for the second image in each second sample image pair as the description information corresponding to the sample image pair, and the description information is input into the diffusion model, so that the diffusion model processes the received first image based on the received control conditions and the description information to obtain the model output result corresponding to the sample image. Among them, the above description information can also be referred to as a prompt.

[0216] So far, the training processes of the first control network and the second control network are completed respectively. However, considering that the image enhancement model can, through the setting of fusion weights, perform focused enhancement on the input image in different dimensions to achieve the effect that the enhancement of detailed texture is higher than that of illumination, or the effect that the enhancement of illumination is higher than that of detailed texture. Therefore, in the process of image enhancement, the cooperation between the first control network and the second control network is also involved. Thus, after the second control network is trained, the first control network can be retrained so that the retrained first control network can adapt to the second control network.

[0217] Based on this, in response to the completion of the training of the second control network, multiple sample image pairs in the training set are selected to obtain multiple third sample image pairs. Then, for each third sample image pair, the first image in the sample image pair, as well as the corresponding illumination information and normal map, are input into the image enhancement model. The first control network inputs the output content determined as the first feature content based on the illumination information corresponding to the first image as the control condition into the diffusion model, and the second control network inputs the output content determined as the second feature content based on the normal map corresponding to the first image as the control condition into the diffusion model. In this way, the diffusion model fuses according to the respective corresponding fusion weights based on the two received control conditions, and based on the fused content as the control condition to be utilized, processes the received first image to obtain the model output result corresponding to the sample image pair. Thus, based on the difference between the second image in each third sample image pair and the corresponding model output result, a third type of loss value is calculated, and when it is determined that the image enhancement model has not converged based on the third type of loss value, the model parameters of the diffusion model in the image enhancement model and the network parameters of the second control network are fixed, and the network parameters of the first control network are adjusted again.

[0218] In this embodiment, by fixing the model parameters of the diffusion model in the image enhancement model and adjusting the network parameters of the first control network and the second control network, the first control network and the second control network can determine the control conditions for controlling the image enhancement process based on the obtained feature content. In this way, when the image enhancement model processes the image to be enhanced, it can control the image enhancement process according to the above control conditions. Moreover, when performing image enhancement in multiple dimensions on the image, through the above control conditions, the image enhancement model can generate an image that meets the enhanced effect required by the user. That is to say, the setting of the control conditions makes the image enhancement process more flexible.

[0219] For ease of understanding, a specific embodiment is hereinafter combined to specifically illustrate an image enhancement method provided by an embodiment of the present application.

[0220] In the related art, when performing image enhancement, after decomposing a low-light image (target visible light image or first image, also referred to as a low-light image) into an illumination image and a reflection image, it is used as physical prior information and input into a diffusion model to generate a low-light image enhancement result. This solution does not utilize the illumination scene under good lighting during the day and does not have a flexible adjustment strategy. Therefore, it is impossible to efficiently and flexibly adjust the generation effect of low-light image enhancement.

[0221] Based on this, this specific embodiment provides a solution for enhancing a low-light image by combining a good-light image (second image, also referred to as a normal-light image) of the same viewing angle scene under a good lighting scene, and can flexibly adjust the image enhancement process.

[0222] The following combines Figure 5 and Figure 6 to specifically introduce the solution provided by this specific embodiment. Among them, Figure 5 is a schematic flowchart of a specific embodiment provided by an embodiment of the present application, Figure 6 is a schematic structural diagram of a specific embodiment provided by an embodiment of the present application.

[0223] In this specific embodiment, an RGB-NIR camera is used as the image acquisition device, and the image acquisition device is also connected to a storage database (image database or data storage space) for storing data. Next, the overall process of this solution will be described in conjunction with Figure 5 The overall process of this solution can be divided into three stages: a model training stage, a database establishment and image retrieval matching stage, and a model inference stage.

[0224] Step1. The model training stage includes:

[0225] Step1.1. Data preparation.

[0226] Prepare two types of data, a low-light / normal-light image paired dataset (i.e., the training set in the embodiments of the present application, where the low-light image is the first image and the normal-light image is the second image), which are image pairs taken in real scenarios and contain various environments and objects. Then, preprocess the above dataset. Use a preprocessing normal preprocessor to extract the normal map of the low-light dataset, that is, extract the normal map of the low-light images in the above dataset; use the Retinex algorithm to calculate the illumination component maps of the low-light image and the normal-light image respectively to obtain the illumination information of the low-light image and the normal-light image.

[0227] Step1.2. Train a control network based on illumination information; among them, the control network based on illumination information is the first control network in the embodiments of the present application.

[0228] Input the low-light image into the SD model (i.e., the trained image enhancement model in the embodiments of the present application). Use a pair of images, namely the illumination component map (illumination information) corresponding to the low-light image and the illumination component map (illumination information) corresponding to the normal-light image, as the input conditions, and use the normal-light image as the training target to train the control network based on illumination information. Among them, when training the control network based on illumination information, fix the weight part of the SD model (fix the model parameters of the diffusion model in the image enhancement model), and train the corresponding partial weight parameters (the network parameters of the first control network) of the control network based on illumination information. At the same time, in order to make the control effect more accurate in the model inference stage, perform a foreground-background separation operation on the illumination component maps and the low-light / normal-light image pairs participating in the training with a probability of 40% in the second half of the training. Only use the mask (mask) of the background area as the input for training (that is, when the convergence value of the first control network reaches the intermediate convergence value, perform a foreground-background separation operation on the first image and the second image in each first sample image pair that has not participated in the training process, and use the background images of the first image and the second image as a new first sample image pair to train the first control network).

[0229] Step1.3. Train a control network based on normal information; among them, the control network based on normal information is the second control network in the embodiments of the present application, and the above control network based on normal information can also be called a control network based on normal map.

[0230] Input the normal image corresponding to the low-light image as the control condition of the SD model, use a pre-trained image-to-text model to generate the prompt corresponding to the normal-light image, and use the normal-light image as the target image for training. Then, fix the weight part of the SD model and train the corresponding partial weight parameters (the network parameters of the second control network) of the control network based on normal information.

[0231] 2. Database establishment and image retrieval and matching stage, including:

[0232] Step2.1. Establish a multi-view normal illumination RGB-NIR database:

[0233] Given an RGB-NIR camera, during the daytime with good illumination, capture representative images at different lens angles and ensure that all possible night views are covered.

[0234] Group and store the collected RGB images (visible light images) and NIR images (non-visible light images) according to different view ranges. For example, different view regions can be divided according to Euler angles.

[0235] Then, extract and describe the features of the NIR images for subsequent matching. Specifically, first, preprocess the NIR images, perform denoising and contrast enhancement to improve the accuracy of feature extraction. Then, detect feature points in the preprocessed NIR images, such as SURF and ORB feature point detection algorithms. Next, describe the extracted features, such as using algorithms like ORB and FREAK to describe the extracted features. Among them, the feature description of the image will be associated with the image and stored in the storage database.

[0236] Step2.2. Input the low-light RGB image and low-light NIR image to be processed, and retrieve and match the RGB-NIR image pairs with the same view:

[0237] The RGB-NIR camera can capture a low-light RGB image (target visible light image) and a low-light NIR image (first auxiliary image) at a certain view at night, and use the low-light NIR image to retrieve and match the corresponding RGB-NIR image pair (target image pair) under good illumination.

[0238] First, determine the view region (view range) to be retrieved according to the Euler angle state of the camera during image acquisition. Then, similar to Step2.1 above, preprocess the low-light NIR image first, then extract features, and then describe the features. Finally, enter the feature matching stage, where the nearest neighbor matching or fast matching algorithm based on the feature tree can be used to match the feature points of the image to be matched with the feature points in the database. It should be noted that the RANSAC algorithm can also be used to remove incorrect matches.

[0239] After the above process is executed, the low-light RGB-NIR image pair to be processed can be retrieved and matched with the RGB-NIR image pair with the same view under good illumination.

[0240] 3. Model inference stage, including:

[0241] Assume that a low-light RGB image to be enhanced has been obtained and the corresponding low-light NIR image and a well-lit RGB image has been retrieved (the visible light image in the target image pair) and the corresponding well-lit NIR image (the non-visible light image in the target image pair).

[0242] Step3.1. Extract spatial geometric information: Input the low-light NIR image into a pre-trained normal preprocessor to obtain the corresponding normal map, and determine the spatial geometric information of the image (the second auxiliary feature content).

[0243] Exemplarily, as Figure 6 shown, input the low-light NIR image into the normal preprocessor to obtain the normal map. Then, input the normal map into the normal feature encoder to encode the image information in the normal map, and based on the encoded data obtained, determine the spatial geometric information of the low-light NIR image. Then, input the obtained spatial geometric information into the control network based on normal information.

[0244] Step3.2. Extract illumination information: Obtain reference illumination information (the first auxiliary feature content) from the well-lit RGB image where, in order to obtain effective illumination information, the present solution provides the following three methods to extract the background image from the image and then obtain the illumination information from the background image:

[0245] First, use the Gaussian mixture model method to distinguish the fixed background and moving objects, extract the

[0246] background object area, and obtain the illumination information from this background object area; Second, use the BiRefNet foreground-background separation model to extract the background area. Then, after extracting the

[0247] background object area, calculate the illumination information for this area to obtain the illumination component map corresponding to this background object area as the illumination information obtained from this background object area; among them, the Retinex algorithm can be used to calculate the illumination component map; Third, use the Diffusion light model to extract the illumination environment map, where, when adopting this method, the illumination component map used for training the corresponding control network based on illumination information is the illumination environment map.

[0248] Optionally, it is possible to separately obtain from the well-lit RGB image

[0249] from the well-lit RGB image and low-light RGB images Obtain their respective illumination information, and then fuse the two illumination information to obtain reference illumination information.

[0250] Exemplarily, as Figure 6 shown, the low-light RGB image and the well-illuminated RGB image are input into the illumination feature extractor to respectively extract the low-light feature map corresponding to the low-light RGB image and the well-illuminated feature map corresponding to the well-illuminated RGB image through the illumination feature extractor, and the low-light feature map and the well-illuminated feature map are input into the illumination feature encoder to fuse and encode the image information in the low-light feature map and the well-illuminated feature map to obtain reference illumination information. Then, the obtained reference illumination information is input into the control network based on illumination information.

[0251] Step3.3. Generate an enhanced image of the low-light image through the SD model:

[0252] Similar to the image generation process in the SD model, after inputting the low-light RGB image into the SD model, first, use the Variational Auto-Encoder (VAE encoder) to encode the low-light RGB image to be enhanced, and then use the latent space features to compress the encoded features in the pixel dimension, retaining the key structural information of the image (such as edges, textures, etc.), thereby reducing the computational complexity during subsequent image enhancement. Then, add noise to the encoded features after pixel dimension compression and input them into the U-Net generation model for denoising. At the same time, use two trained control networks to control the SD model to generate images respectively based on the feature content corresponding to the illumination information and the feature content corresponding to the spatial geometric information. Then, use the Variational Auto-Decoder (VAE decoder) to decode the generated image to obtain the illumination-enhanced RGB image.

[0253] Moreover, by adjusting the control intensity weight parameters W1 and W2 (fusion weights), fuse the feature content corresponding to the reference illumination information and the feature content corresponding to the spatial geometric information in the feature fusion module, so as to flexibly adjust the image enhancement effect according to specific requirements. For example, increasing the weight parameter of the control network based on normal information makes the enhanced image pay more attention to detail textures, but the illumination brightening effect may be reduced at this time. Another example is to increase the weight parameter of the control network based on illumination information, making the enhanced image pay more attention to the illumination brightening effect, but the detail textures of the image may be reduced at this time. Among them, the weight parameter W1 is the weight corresponding to the control network based on illumination information, and the weight parameter W2 is the weight corresponding to the control network based on normal information.

[0254] As can be seen above, a suitable reference image (the visible light image in the target image pair) is retrieved through image feature matching, and by analyzing and processing the reference image, an effective reference target object (reference illumination information) is extracted as the reference information for denoising by the diffusion model (the illumination information of the visible light image in the target image pair). Using the illumination component features (the first auxiliary feature content) extracted from the normal light image under a good illumination scene for reference (the image collected under a scene that meets the illumination conditions) and the normal map features (the second auxiliary feature content) extracted from the low-light image to be enhanced as the control signal (control condition) guidance, the enhanced normal light image is obtained through the diffusion model. This control mode has higher flexibility and can adjust the emphasis on low-light image enhancement according to different usage requirements by setting the control coefficient (fusion weight), that is, whether to be more inclined to the consistency of the original low-light image (the effect of enhancing detail texture is higher than that of enhancing illumination) or the illumination effect (the effect of enhancing illumination is higher than that of enhancing detail texture).

[0255] Corresponding to the above method embodiment, an embodiment of the present application also provides an image enhancement device, as Figure 7 shown. The device includes:

[0256] An image acquisition module 710, configured to acquire a target visible light image to be enhanced and acquire a first auxiliary image that belongs to a non-visible light image; wherein, the target visible light image and the first auxiliary image are images obtained by an image acquisition device performing image acquisitions on visible light and non-visible light respectively for the same scene content;

[0257] An image pair determination module 720, configured to determine a target image pair from a plurality of preset reference image pairs based on the first auxiliary image; wherein, each reference image pair includes: a visible light image and a non-visible light image obtained by the image acquisition device performing image acquisitions on visible light and non-visible light respectively for the same scene content, and the visible light image in each reference image pair is an image collected under a scene that meets the illumination conditions; the target image pair is: the reference image pair whose included non-visible light image matches the first auxiliary image;

[0258] A first determination module 730, configured to extract the illumination information of the visible light image in the target image pair and determine, based on the extracted illumination information, the auxiliary feature content corresponding to the target visible light image in terms of the illumination information dimension as the first auxiliary feature content;

[0259] An image enhancement module 740 is configured to input the target visible light image and the first auxiliary feature content into a pre-trained image enhancement model to obtain an enhanced image of the target visible light image. The image enhancement model is a diffusion model for image generation with a first control network added, and the first control network is a network that controls the diffusion model according to the auxiliary feature content regarding the illumination information dimension corresponding to the image to be enhanced.

[0260] As can be seen above, in the solution provided by the embodiments of the present application, multiple reference image pairs can be obtained in advance. Each reference image pair includes: a visible light image and a non-visible light image of the same scene content obtained by an image acquisition device through image acquisition of visible light and non-visible light respectively, and the visible light image in each reference image pair is an image acquired under a scene that meets the illumination conditions.

[0261] When performing image enhancement, an image obtained by an image acquisition device through image acquisition of visible light for the same scene content can be obtained as the target visible light image to be enhanced, and an image obtained by the image acquisition device through image acquisition of non-visible light for the same scene content can be obtained as the first auxiliary image. Thus, based on the first auxiliary image, from the above multiple reference image pairs, a reference image pair whose included non-visible light image matches the first auxiliary image is determined as the target image pair. Then, by extracting the illumination information of the visible light image in the target image pair and based on the extracted illumination information, the auxiliary feature content regarding the illumination information dimension corresponding to the target visible light image is determined as the first auxiliary feature content, and the target visible light image and the first auxiliary feature content are input into the pre-trained image enhancement model to obtain an enhanced image of the target visible light image.

[0262] The image enhancement model is a diffusion model for image generation with a first control network added, and the first control network is a network that controls the diffusion model according to the auxiliary feature content regarding the illumination information dimension corresponding to the image to be enhanced. That is to say, in the process of using the above image enhancement model for image enhancement, the first control network guides the image enhancement process based on the obtained first auxiliary feature content, so that the obtained enhanced image is close to the image acquired under a scene that meets the illumination conditions, to meet the requirement of image enhancement processing for visible light images and achieve the display effect of images acquired under good illumination.

[0263] Optionally, in one implementation, the image enhancement model is further provided with a second control network. The second control network is a network that controls the diffusion model according to the auxiliary feature content regarding the spatial geometric information dimension corresponding to the image to be enhanced.

[0264] The device further includes:

[0265] A second determination module, configured to extract the normal map of the first auxiliary image, and based on the extracted normal map, determine the auxiliary feature content corresponding to the target visible light image in the dimension of spatial geometric information as the second auxiliary feature content;

[0266] An information input module, configured to input the second auxiliary feature content into a pre-trained image enhancement model.

[0267] Optionally, in one implementation manner, the image enhancement model further includes a feature fusion module; wherein, the feature fusion module is configured to fuse the feature content used as a control condition output by the first control network and the feature content used as a control condition output by the second control network according to their respective fusion weights, and input the fused feature content into the diffusion model;

[0268] The diffusion model is specifically configured to generate an enhanced image of the target visible light image based on the fused feature content;

[0269] The respective fusion weights of the first control network and the second control network are set according to the required image enhancement effect, and the enhancement effect includes: an effect where the enhancement of detail texture is higher than the enhancement of illumination or an effect where the enhancement of illumination is higher than the enhancement of detail texture.

[0270] Optionally, in one implementation manner, the image acquisition device is a rotatable acquisition device;

[0271] The image pair determination module 720 is specifically configured to:

[0272] Select the view range to which the shooting view of the first auxiliary image belongs from multiple view ranges; wherein, the multiple view ranges are different view ranges of the image acquisition device, and each view range corresponds to each reference image pair among the multiple reference image pairs whose shooting view belongs to this view range;

[0273] Determine each reference image pair corresponding to the selected view range;

[0274] Calculate the similarity between the first auxiliary image and the non-visible light images in each of the determined reference image pairs;

[0275] Based on the obtained similarity, determine the target image pair from the determined reference image pairs.

[0276] Optionally, in one implementation manner, the first determination module 730 is specifically configured to:

[0277] Extract the illumination information of the entire image of the visible light image in the target image pair;

[0278] Or,

[0279] Extract the image background region from the visible light image in the target image pair to obtain the image background to be utilized, and extract the illumination component of the image background to be utilized as the illumination information;

[0280] Or,

[0281] Extract the illumination environment map of the visible light image in the target image pair as the illumination information.

[0282] Optionally, in one implementation, the apparatus further includes:

[0283] An extraction module, configured to extract the illumination information of the target visible light image before determining the auxiliary feature content corresponding to the target visible light image in the dimension of illumination information based on the extracted illumination information as the first auxiliary feature content;

[0284] The first determination module 730 is specifically configured to:

[0285] Generate the auxiliary feature content corresponding to the target visible light image in the dimension of illumination information based on the illumination information of the visible light image in the target image pair and the illumination information of the target visible light image as the first auxiliary feature content.

[0286] Optionally, in one implementation, the following modules are used to train the image enhancement model:

[0287] An acquisition sub-module, configured to acquire a training set; wherein, the training set includes a plurality of sample image pairs; each sample image pair includes a first image to be enhanced and a second image that is the image to be obtained after enhancing the first image, and the first image and the second image in each sample image pair are both visible light images and the second image is an image collected under a scene that meets the illumination conditions;

[0288] A first determination sub-module, configured to, for each sample image pair, extract the illumination information of the second image in the sample image pair, and determine the auxiliary feature content corresponding to the first image in the sample image pair in the dimension of illumination information as the first feature content corresponding to the sample image pair;

[0289] A second determination sub-module, configured to, for each sample image pair, extract the normal map of the first image in the sample image pair, and determine the auxiliary feature content corresponding to the first image in the sample image pair in the dimension of spatial geometric information as the second feature content corresponding to the sample image pair;

[0290] A model training sub-module, configured to train a first control network and a second control network of the image enhancement model based on each sample image pair in the training set, and the corresponding first feature content and second feature content, when the model parameters of the diffusion model in the image enhancement model are fixed.

[0291] Optionally, in one implementation, the model training sub-module is specifically configured to:

[0292] Select multiple sample image pairs from the training set to obtain multiple first sample image pairs to be utilized, and use the second image in the multiple first sample image pairs as the ground truth content. Based on the first image in the selected multiple first sample image pairs and the corresponding first feature content, train the first control network of the image enhancement model;

[0293] In response to the completion of the training of the first control network, select multiple sample image pairs from the training set to obtain multiple second sample image pairs to be utilized, and use the second image in the selected multiple second sample image pairs as the ground truth content. Based on the first image in the selected multiple second sample image pairs and the corresponding second feature content, train the second control network of the image enhancement model;

[0294] In response to the completion of the training of the second control network, select multiple sample image pairs from the training set to obtain multiple third sample image pairs, use the second image in the selected multiple third sample image pairs as the ground truth content, and based on the first image in the selected multiple third sample image pairs, the corresponding first feature content and second feature content, retrain the first control network of the image enhancement model.

[0295] An embodiment of the present application further provides an electronic device, as Figure 8 shown, including:

[0296] A memory 801, configured to store a computer program;

[0297] A processor 802, configured to implement the image enhancement method described in any one of the above embodiments when executing the program stored in the memory 801.

[0298] And the above electronic device may further include a communication bus and / or a communication interface. The processor 802, the communication interface, and the memory 801 complete communication with each other through the communication bus.

[0299] The communication bus mentioned in the above-mentioned electronic device can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0300] The communication interface is used for communication between the above-mentioned electronic device and other devices.

[0301] The memory can include a Random Access Memory (RAM), and can also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory can also be at least one storage device located far from the aforementioned processor.

[0302] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0303] In another embodiment provided by this application, a computer-readable storage medium is also provided. A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the steps of any of the above image enhancement methods are implemented.

[0304] In another embodiment provided by this application, a computer program product containing instructions is also provided. When it runs on a computer, it causes the computer to execute any of the image enhancement methods in the above embodiments.

[0305] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a solid-state disk (SSD), etc.

[0306] It should be noted that, in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0307] Each embodiment in this specification is described in a related manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the device embodiments, electronic device embodiments, computer-readable storage medium embodiments, and computer program product embodiments, since they are basically similar to the method embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the partial descriptions of the method embodiments.

[0308] The above are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are all included in the protection scope of the present application.

Claims

1. An image enhancement method, characterized in that, The method includes: Obtaining a target visible light image to be enhanced and obtaining a first auxiliary image belonging to a non-visible light image; wherein, the target visible light image and the first auxiliary image are images obtained by an image acquisition device respectively performing image acquisition of visible light and non-visible light for the same scene content; Based on the first auxiliary image, determining a target image pair from a preset plurality of reference image pairs; wherein, each reference image pair includes: a visible light image and a non-visible light image of the same scene content obtained by the image acquisition device respectively performing image acquisition of visible light and non-visible light, and the visible light image in each reference image pair is an image acquired under a scene that meets the lighting conditions; the target image pair is: a reference image pair in which the included non-visible light image matches the first auxiliary image; Extracting the lighting information of the visible light image in the target image pair, and based on the extracted lighting information, determining auxiliary feature content corresponding to the target visible light image in terms of the lighting information dimension as the first auxiliary feature content; Inputting the target visible light image and the first auxiliary feature content into a pre-trained image enhancement model to obtain an enhanced image of the target visible light image; wherein, the image enhancement model is a diffusion model for image generation provided with a first control network, and the first control network is a network for controlling the diffusion model according to the auxiliary feature content corresponding to the image to be enhanced in terms of the lighting information dimension.

2. The method according to claim 1, characterized in that, The image enhancement model is further provided with a second control network; wherein, the second control network is: a network for controlling the diffusion model according to the auxiliary feature content corresponding to the image to be enhanced in terms of the spatial geometric information dimension; The method further includes: Extracting the normal map of the first auxiliary image, and based on the extracted normal map, determining auxiliary feature content corresponding to the target visible light image in terms of the spatial geometric information dimension as the second auxiliary feature content; Inputting the second auxiliary feature content into a pre-trained image enhancement model.

3. The method according to claim 2, characterized in that, The image enhancement model further includes a feature fusion module; wherein, the feature fusion module is used to fuse the feature content output by the first control network as a control condition and the feature content output by the second control network as a control condition according to their respective fusion weights, and input the fused feature content into the diffusion model; The diffusion model is specifically used to generate an enhanced image of the target visible light image based on the fused feature content; The fusion weights of the first control network and the second control network are set according to the required image enhancement effect, and the enhancement effect includes: an effect of enhancing detail texture higher than enhancing lighting or an effect of enhancing lighting higher than enhancing detail texture.

4. The method according to any one of claims 1 to 3, characterized in that, The image acquisition device is a rotatable acquisition device; The determining a target image pair from a preset plurality of reference image pairs based on the first auxiliary image includes: Select the perspective range to which the shooting perspective of the first auxiliary image belongs from multiple perspective ranges; wherein, the multiple perspective ranges are different perspective ranges of the image acquisition device, and each perspective range corresponds to each reference image pair in the multiple reference image pairs whose shooting perspective belongs to this perspective range; Determine each reference image pair corresponding to the selected perspective range; Calculate the similarity between the first auxiliary image and the non-visible light images in each of the determined reference image pairs; Based on the obtained similarities, determine a target image pair from the determined reference image pairs.

5. The method according to any one of claims 1 to 3, characterized in that, The extracting the illumination information of the visible light image in the target image pair includes: Extract the illumination information of the entire image of the visible light image in the target image pair; Or, Perform extraction on the visible light image in the target image pair with respect to the image background area to obtain an image background to be utilized, and extract the illumination component of the image background to be utilized as the illumination information; Or, Extract the illumination environment map of the visible light image in the target image pair as the illumination information.

6. The method according to any one of claims 1-3, characterized in that Before the determining, based on the extracted illumination information, the auxiliary feature content corresponding to the target visible light image with respect to the illumination information dimension as the first auxiliary feature content, the method further includes: Extract the illumination information of the target visible light image; The determining, based on the extracted illumination information, the auxiliary feature content corresponding to the target visible light image with respect to the illumination information dimension as the first auxiliary feature content includes: Based on the illumination information of the visible light image in the target image pair and the illumination information of the target visible light image, generate the auxiliary feature content corresponding to the target visible light image with respect to the illumination information dimension as the first auxiliary feature content.

7. The method according to claim 2 or 3, characterized in that, The training process of the image enhancement model includes: Obtain a training set; wherein, the training set includes multiple sample image pairs; each sample image pair includes a first image to be subjected to image enhancement and a second image that is the image required to be obtained after the enhancement of the first image, and the first image and the second image in each sample image pair are both visible light images and the second image is an image acquired under a scene that meets the illumination conditions; For each sample image pair, extract the illumination information of the second image in the sample image pair, and based on the extracted illumination information, determine the auxiliary feature content corresponding to the first image in the sample image pair with respect to the illumination information dimension as the first feature content corresponding to the sample image pair; For each sample image pair, extract the normal map of the first image in the sample image pair, and based on the extracted normal map, determine the auxiliary feature content corresponding to the first image in the sample image pair with respect to the spatial geometric information dimension as the second feature content corresponding to the sample image pair; With the model parameters of the diffusion model in the image enhancement model fixed, train the first control network and the second control network of the image enhancement model based on each sample image pair in the training set and the corresponding first feature content and second feature content.

8. The method according to claim 7, wherein Training the first control network and the second control network of the image enhancement model based on each sample image pair in the training set, and the corresponding first feature content and second feature content, includes: Selecting multiple sample image pairs in the training set to obtain multiple first sample image pairs to be utilized, and using the second image in the multiple first sample image pairs as the ground truth content, training the first control network of the image enhancement model based on the first image in the selected multiple first sample image pairs and the corresponding first feature content; In response to the completion of the training of the first control network, selecting multiple sample image pairs in the training set to obtain multiple second sample image pairs to be utilized, and using the second image in the selected multiple second sample image pairs as the ground truth content, training the second control network of the image enhancement model based on the first image in the selected multiple second sample image pairs and the corresponding second feature content; In response to the completion of the training of the second control network, selecting multiple sample image pairs in the training set to obtain multiple third sample image pairs, using the second image in the selected multiple third sample image pairs as the ground truth content, and retraining the first control network of the image enhancement model based on the first image in the selected multiple third sample image pairs, the corresponding first feature content, and second feature content.

9. An image enhancement device, characterized in that, The device includes: An image acquisition module, configured to acquire a target visible light image to be enhanced and a first auxiliary image belonging to a non-visible light image; wherein, the target visible light image and the first auxiliary image are images obtained by an image acquisition device performing image acquisitions on the same scene content with respect to visible light and non-visible light respectively; An image pair determination module, configured to determine a target image pair from a preset plurality of reference image pairs based on the first auxiliary image; wherein each reference image pair includes: a visible light image and a non-visible light image of the same scene content obtained by the image acquisition device performing image acquisitions on visible light and non-visible light respectively, and the visible light image in each reference image pair is an image acquired under a scene that meets the illumination conditions; the target image pair is: the reference image pair whose included non-visible light image matches the first auxiliary image; A first determination module, configured to extract the illumination information of the visible light image in the target image pair, and based on the extracted illumination information, determine the auxiliary feature content corresponding to the target visible light image in terms of the illumination information dimension as the first auxiliary feature content; An image enhancement module, configured to input the target visible light image and the first auxiliary feature content into a pre-trained image enhancement model to obtain an enhanced image of the target visible light image; wherein, the image enhancement model is a diffusion model for image generation with a first control network added, and the first control network is a network for controlling the diffusion model according to the auxiliary feature content corresponding to the image to be enhanced in terms of the illumination information dimension.

10. An electronic device, characterized in that, Including: A memory, configured to store a computer program; A processor, when executing a program stored in a memory, implements the method according to any one of claims 1-8.

11. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the method according to any one of claims 1-8 is implemented.