Image Enhancement Method, Apparatus and Computer-Readable Storage Medium

By building an image enhancement model of a multi-level generative adversarial network, the network is trained using the low-frequency features and fusion images of high-quality images, and the problem of panoramic image quality degradation and artifacts is solved, and efficient image detail recovery and quality improvement is achieved.

CN114511449BActive Publication Date: 2025-05-30RICOH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011279491.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-16
Publication Date
2025-05-30
Estimated Expiration
2040-11-16

AI Technical Summary

Technical Problem

The panoramic images captured by panoramic cameras are relatively deteriorating in terms of sharpness, resolution and hue differences. The existing image enhancement and super-resolution methods are not effective in image detail recovery and are prone to artifacts, especially in extreme areas.

Method used

An image enhancement model including the first-level and second-level generative adversarial network is adopted. By acquiring multiple sets of training data, high-quality images are used as target images, and the network is trained based on low-frequency features and fusion images to generate high-quality enhanced images.

Benefits of technology

Effectively improve the quality of the image, reduce noise impact, reduce artifact generation, retain more texture details, and further optimize the image enhancement effect by introducing high-frequency similarity loss and replacement loss functions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114511449B_ABST
    Figure CN114511449B_ABST
Patent Text Reader

Abstract

The present invention provides an image enhancement method, apparatus, and computer-readable storage medium, belonging to the technical field of image processing. The image enhancement method includes: obtaining multiple sets of training data, each set of training data including a first image and a second image; constructing an image enhancement model including a first-level and a second-level generative adversarial network, and training the image enhancement model using the multiple sets of training data, wherein, taking the second image as the target image, training the first-level generative adversarial network with an enhanced low-frequency image generated based on the low-frequency features of the first image; taking the second image as the target image, training the second-level generative adversarial network with an enhanced image generated based on the fused image of the first image; the fused image is obtained by fusing the first image and the enhanced low-frequency image; inputting a third image to be enhanced into the trained image enhancement model, and outputting a fourth image after image enhancement. The present invention can improve the image quality of an image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to an image enhancement method, apparatus, and computer-readable storage medium. Background Art

[0002] The panoramic images captured by panoramic cameras usually have a field of view of 180 degrees or more. However, compared with the planar images captured by relatively high-quality cameras (such as digital single-lens reflex cameras or DSLRs), panoramic images are relatively poor in terms of sharpness, resolution, and chromatic aberration.

[0003] To solve the above problems, the prior art has proposed methods for enhancing and super-resolving (SR) images. These methods have achieved good improvements in image sharpening, denoising, deblurring, contrast improvement, and chromatic aberration correction, and have improved the image quality. However, the above methods generally do not well recover the image details, and at the same time, some artifacts are generated. Especially in the extreme regions of panoramic cameras, such artifacts become more serious as the super-resolution (sometimes also simply referred to as super-resolution in this article) multiple increases. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide an image enhancement method, apparatus, and computer-readable storage medium that can improve the quality of images.

[0005] To solve the above technical problems, the embodiments of the present invention provide the following technical solutions:

[0006] An embodiment of the present invention provides an image enhancement method, including:

[0007] Obtaining multiple sets of training data, each set of training data including a first image and a second image, where the image quality of the second image is better than that of the first image;

[0008] Constructing an image enhancement model including a first-level generative adversarial network and a second-level generative adversarial network, and training the image enhancement model using the multiple sets of training data to obtain a trained image enhancement model, where the first-level generative adversarial network is trained with the second image as the target image and an enhanced low-frequency image generated based on the low-frequency features of the first image; the second-level generative adversarial network is trained with the second image as the target image and an enhanced image generated based on the fused image of the first image; the fused image is obtained by fusing the first image and the enhanced low-frequency image;

[0009] Inputting a third image to be enhanced into the image enhancement model, and outputting a fourth image with enhanced image quality.

[0010] Optionally, training the first-level generative adversarial network using the enhanced low-frequency image generated based on the low-frequency features of the first image with the second image as the target image includes:

[0011] Extracting low-frequency features from the first image;

[0012] Performing image enhancement based on the low-frequency features to generate an enhanced low-frequency image of the first image;

[0013] Using the enhanced low-frequency image of the first image to train the first-level generative adversarial network with the second image as the target image until a preset training end condition is satisfied.

[0014] Optionally, the loss function Loss_G_1 of the first-level generative adversarial network is: Loss_G_1 = Lcobi 1 + λ 1 LG 1 + η 1 Lcolor 1 ;

[0015] Wherein, Lcobi 1 represents the context bilateral loss function for the enhanced low-frequency image and the second image; LG 1 represents the adversarial loss function for the enhanced low-frequency image and the second image; Lcolor 1 represents the color loss function for the enhanced low-frequency image and the second image; λ 1 and η 1 are both preset constants.

[0016] Optionally, training the second-level generative adversarial network using the enhanced image generated based on the fused image of the first image with the second image as the target image includes:

[0017] Adding the pixels at the same positions of the first image and the enhanced low-frequency image to obtain a fused image, and generating an enhanced image of the first image based on the fused image;

[0018] Using the enhanced image of the first image to train the second-level generative adversarial network with the second image as the target image until a preset training end condition is satisfied.

[0019] Optionally, the loss function Loss_G_2 of the second-level generative adversarial network is: Loss_G_2 = η 2 Lcobi-hf + η 3 Lcobi 2 + λ 2 LG 2+η 4 Lcolor 2 ;

[0020] Wherein, Lcobi 2 represents the context bilateral loss function for the enhanced image and the second image; Lcobi-hf represents the context bilateral loss function for the high-frequency features of the enhanced image and the high-frequency features of the second image; LG 2 represents the adversarial loss function for the enhanced image and the second image; Lcolor 2 represents the color loss function for the enhanced image and the second image; η 2 、η 3 、λ 2 and η 4 are all preset constants.

[0021] Optionally, the content captured by the first image and the second image in the same set of training data is the same.

[0022] Optionally, the first image is an equidistant cylindrical projection or a perspective view, and the second image is a perspective view.

[0023] Optionally, the image quality of the second image being better than the image quality of the first image includes at least one of the following:

[0024] The resolution of the second image is greater than the resolution of the first image;

[0025] The signal-to-noise ratio of the second image is higher than the signal-to-noise ratio of the first image;

[0026] The chromatic aberration of the second image is lower than the chromatic aberration of the first image.

[0027] An embodiment of the present invention further provides an image enhancement device, including:

[0028] An acquisition module, configured to acquire multiple sets of training data, each set of training data including a first image and a second image, wherein the image quality of the second image is better than the image quality of the first image;

[0029] A training module for constructing an image enhancement model including a first - level generative adversarial network and a second - level generative adversarial network, training the image enhancement model using the multiple sets of training data to obtain a trained image enhancement model. Among them, taking the second image as the target image, training the first - level generative adversarial network with the enhanced low - frequency image generated based on the low - frequency features of the first image; taking the second image as the target image, training the second - level generative adversarial network with the enhanced image generated based on the fused image of the first image; the fused image is obtained by fusing the first image and the enhanced low - frequency image;

[0030] An image processing module for inputting a third image to be enhanced into the image enhancement model and outputting a fourth image with enhanced image.

[0031] Optionally, the first - level generative adversarial network includes:

[0032] An octave convolution module for extracting low - frequency features from the first image;

[0033] A first generation network for enhancing an image based on the low - frequency features of the first image and generating an enhanced low - frequency image of the first image;

[0034] A first adversarial network for judging whether the enhanced low - frequency image is consistent with the second image;

[0035] The training module is further configured to use the enhanced low - frequency image of the first image to train the first - level generative adversarial network with the second image as the target image until a preset training end condition is met.

[0036] Optionally, the second - level generative adversarial network includes:

[0037] A fusion module for adding the pixels at the same positions of the first image and the enhanced low - frequency image to obtain a fused image;

[0038] A second generation network for generating an enhanced image of the first image based on the fused image;

[0039] A second adversarial network for judging whether the enhanced image is consistent with the second image;

[0040] The training module is further configured to use the enhanced image of the first image to train the second - level generative adversarial network with the second image as the target image until a preset training end condition is met.

[0041] An embodiment of the present invention further provides an image enhancement device, including:

[0042] A processor; and

[0043] A memory in which computer program instructions are stored

[0044] Wherein, when the computer program instructions are run by the processor, the processor is caused to perform the following steps:

[0045] Obtain multiple sets of training data, each set of training data including a first image and a second image, wherein the image quality of the second image is better than that of the first image;

[0046] Construct an image enhancement model including a first-level generative adversarial network and a second-level generative adversarial network, and use the multiple sets of training data to train the image enhancement model to obtain a trained image enhancement model. Wherein, using the second image as the target image, an enhanced low-frequency image generated based on the low-frequency features of the first image is used to train the first-level generative adversarial network; using the second image as the target image, an enhanced image generated based on the fused image of the first image is used to train the second-level generative adversarial network; the fused image is obtained by fusing the first image and the enhanced low-frequency image;

[0047] Input a third image to be enhanced into the image enhancement model, and output a fourth image after image enhancement.

[0048] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the steps of the image enhancement method as described above are implemented.

[0049] The embodiments of the present invention have the following beneficial effects:

[0050] In the embodiments of the present invention, low-frequency features are extracted in the first-level generative adversarial network and enhanced; then the original features of the image are added to the enhanced low-frequency features to reduce the influence of noise, and more texture details can be retained while reducing the generation of artifacts. In addition, the embodiments of the present invention introduce a high-frequency similarity loss, which only focuses on the high-frequency part of the generated image and can directly reduce the generation of artifacts. In addition, in the embodiments of the present invention, the L1 loss in the existing network loss function is replaced with a color loss function, and the color loss function pays more attention to the overall distribution of data, which helps to reduce the generation of artifacts. In addition, in the embodiments of the present invention, the perceptual loss in the existing network loss function is replaced with a CoBi loss, which is insensitive to data alignment and also helps to reduce the generation of artifacts. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 It is a schematic flowchart of an image enhancement method according to an embodiment of the present invention;

[0052] Figure 2 A schematic structural diagram of an image enhancement model according to an embodiment of the present invention;

[0053] Figure 3 A schematic structural diagram of an octave convolution module according to an embodiment of the present invention;

[0054] Figure 4 A structural block diagram of an image enhancement device according to an embodiment of the present invention;

[0055] Figure 5 Another structural block diagram of the image enhancement device according to an embodiment of the present invention. Detailed implementation manners

[0056] To make the technical problems, technical solutions and advantages to be solved by the embodiments of the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0057] The panoramic images captured by a panoramic camera usually have a field of view angle of 180 degrees or more. However, compared with the planar images captured by a relatively high-quality camera (such as a digital single-lens reflex camera), the panoramic images are relatively poor in terms of sharpness, resolution, and chromatic aberration.

[0058] Image enhancement and super-resolution (SR) methods can improve the quality of panoramic images. Image enhancement includes image sharpening, denoising, deblurring, contrast enhancement, and chromatic aberration correction; image super-resolution improves the image quality by increasing the image resolution. However, when traditional image enhancement and super-resolution methods improve the quality of panoramic images, problems such as artifacts are likely to be introduced.

[0059] To solve the above problems, embodiments of the present invention provide an image enhancement method, device, and computer-readable storage medium, which can improve the quality of images.

[0060] Embodiments of the present invention provide an image enhancement method, as Figure 1 shown, including:

[0061] Step 101, obtaining multiple groups of training data, each group of training data including a first image and a second image, wherein the image quality of the second image is better than that of the first image.

[0062] Here, both the first image and the second image are perspective views; or the first image is an equidistant cylindrical projection map and the second image is a perspective view. Of course, the first image and the second image can also be other types of images.

[0063] In the training data, the content captured by the first image and the second image in the same set of training data is the same. To obtain the training data, cameras with different imaging qualities can be used to capture the same shooting content in advance. For example, a camera with better imaging quality is used to capture content A to obtain a high-quality image, and then a camera with poorer imaging quality is used to capture content A to obtain a low-quality image. After that, the high-quality image and the low-quality image are matched to obtain the first image and the second image.

[0064] The parameters for measuring image quality include resolution, signal-to-noise ratio, and chromatic aberration. The image quality of the second image being better than that of the first image can be at least one of the following: the resolution of the second image is greater than that of the first image, the signal-to-noise ratio of the second image is higher than that of the first image, and the chromatic aberration of the second image is lower than that of the first image.

[0065] Step 102: Construct an image enhancement model including a first-level generative adversarial network and a second-level generative adversarial network, and use the multiple sets of training data to train the image enhancement model to obtain a trained image enhancement model.

[0066] Here, during the training process of the image enhancement model, using the second image as the target image (which can also be called the real image), the first-level generative adversarial network is trained with the enhanced low-frequency image generated based on the low-frequency features of the first image; using the second image as the target image (which can also be called the real image), the second-level generative adversarial network is trained with the enhanced image generated based on the fused image of the first image; the fused image is obtained by fusing the first image and the enhanced low-frequency image.

[0067] When training the image enhancement model, when the preset training end condition is reached, the process ends to obtain a trained image enhancement model. Specifically, the training end condition can be reaching the Nash equilibrium or the training process has converged.

[0068] Step 103: Input the third image to be enhanced into the image enhancement model, and output the fourth image after image enhancement.

[0069] Through the above steps, in the first-generation network of the first-level GAN in the embodiment of the present invention, the low-frequency features of the first image are used to generate an enhanced low-frequency image, and the second image is used as the real image. In the first discriminant network of the GAN, it is judged whether the enhanced low-frequency image is a real image. Since the low-frequency features of the first image are introduced and enhanced, the influence of noise can be directly reduced, and removing this part of the noise is beneficial to reducing the generation of artifacts, thereby improving the quality of the finally generated enhanced image.

[0070] For example, taking the first image as a panoramic image, the extreme regions of the panoramic image (such as the pole regions at the top or bottom edges) are stretched to the entire width of the image. The regions near the poles are stretched horizontally. The pole regions of the equirectangular image are severely distorted, making it very difficult to restore these regions. After adopting the above method, by extracting and enhancing the low-frequency feature part, the influence of noise can be directly reduced or removed. Since noise is the main factor causing artifacts (because high frequencies are usually encoded with fine details and noise, and low frequencies are usually encoded with global structures), the embodiments of the present invention can reduce the generation of artifacts.

[0071] Figure 2 A simplified structural schematic diagram of the image enhancement model of the embodiments of the present invention is given. For a more specific structure of the GAN, reference can be made to the descriptions of related prior arts, which will not be elaborated herein. The image enhancement model of the embodiments of the present invention may include two-level generative adversarial networks (GAN, Generative Adversarial Networks), that is, both the above-mentioned first-level generative adversarial network and the second-level generative adversarial network adopt GAN. Each level of GAN includes a generative network and a discriminative network.

[0072] The first-level generative adversarial network is an enhancement network, responsible for enhancing the low-frequency feature image to obtain an enhanced low-quality image without noise or with less noise. Among them, the first-level generative network generates an enhanced low-quality image (enhanced low-frequency image) based on the low-frequency features of the first image, and the first-level discriminative network determines whether the low-quality image (enhanced low-frequency image) is consistent with the second image. Considering that the second image is a high-quality image, before making a determination, the second image can be downsampled by a downsampling module, and then it is determined whether the low-quality image (enhanced low-frequency image) is consistent with the image obtained by downsampling the second image to obtain a determination result. The implementation methods of downsampling include bilinear interpolation, deconvolution, etc.

[0073] During the training process of the above image enhancement model, the embodiments of the present invention extract low-frequency features from the first image; then, based on the low-frequency features, image enhancement is performed to generate an enhanced low-frequency image of the first image; then, using the second image as the target image (which can also be called the reference image), the first-level generative adversarial network is trained with the enhanced low-frequency image of the first image until a preset training end condition is met. For example, taking the second image as the target image, by calculating and updating the loss function of the first-level generative adversarial network, the value of the loss function is continuously updated until the training end condition is met.

[0074] Here, as an implementation, the embodiments of the present invention may adopt an Octave Convolution (OctConv) module to extract low-frequency features from the first image. As Figure 3 shown, a schematic diagram of the Octave Convolution module is given. OctConv (Octave Convolution) is a plug-and-play structure that can save computational resource consumption while improving accuracy. Natural images can be decomposed into two parts: low spatial frequency and high spatial frequency. The output map of the convolutional layer can also be decomposed and grouped according to its spatial frequency. OctConv uses a coefficient α to factorize the feature map into X H and X L , which represent the high-frequency and low-frequency feature components of the feature map respectively. The multi-frequency feature representation method proposed by OctConv stores the smoothly varying low-frequency mapping in a low-resolution tensor to reduce spatial redundancy. Gaussian filtering used for X L changes its spatial resolution to half of the original, while X H does not perform any operation.

[0075] Figure 3 In, the input Cin: X ∈ R c×h×w , the input feature. Where h and w represent the spatial dimensions, and c represents the number of feature maps or channels.

[0076] X = {X H , X L};

[0077] X H ∈R (1-α)c×h×w is the high-frequency feature, including more details;

[0078] X L ∈R αc×h / 2×w / 2 is the low-frequency feature, which changes slowly in the spatial dimension.

[0079] Output Cout: X L , the low-frequency feature.

[0080] In the embodiments of the present invention, in order to obtain the low-frequency feature, α in can be set to 0, and α out can be set to 1.

[0081] For a more specific structure of OctConv, please refer to the relevant paper (Drop an Octave: Reducing Spatial Redundancy in Convolutional Neural Networks with Octave Convolution, ICCV 2019 arXiv:1904.05049 [cs.CV]), which will not be elaborated in this article.

[0082] The following provides a specific representation of the loss function Loss_G_1 of the first-level generative adversarial network. It should be noted that the following formula is only an example of a loss function that can be adopted in the embodiments of the present invention and is not used to limit the present invention:

[0083] Loss_G_1 = Lcobi 1 + λ 1 LG 1 + η 1 Lcolor 1 ;

[0084] Among them, Lcobi 1 represents the Contextual Bilateral Loss (CoBi-Loss) function for the enhanced low-frequency image and the second image. Introducing the Cobi-loss function can further reduce the generation of artifacts, especially in the extreme regions. Since the distortion in this region is very serious and the data is not easily aligned, and this loss function is insensitive to such data. For the detailed definition of CoBi-Loss, please refer to the description of Xuaner Zhang, et al. "Zoom to Learn, Learn to Zoom" arXiv:1905.05169v1 (2019) in the prior art.

[0085] LG 1 represents the adversarial loss function for the enhanced low-frequency image and the second image, that is, the adversarial loss of fidelity.

[0086] Lcolor 1It represents the color loss function for the enhanced low-frequency image and the second image, which can make the enhanced low-frequency image and the second image (high-quality image) have similar basic structures and colors. Here, in the embodiments of the present invention, the color loss function is used instead of the L1 loss function in the prior art. The L1 loss function is pixel-level. Considering that it is difficult to align the data (especially the distortion in the extreme regions is very serious), artifacts are easily generated. Therefore, the loss proposed in the embodiments of the present invention focuses more on the overall distribution of the data, which helps to reduce the generation of artifacts. As an example, the color loss function describes the data from three aspects: concentration trend (mean), separation trend (covariance), and distribution pattern (skewness), and then uses the L2 loss function to judge the similarity of the distribution.

[0087]

[0088] where n is the batch size, ∝ 1 ,∝ 2 ,∝ 3 are the corresponding parameters. Let C k =(R, G, B) T be a certain pixel in the image, then the mean, covariance, and skewness of the image are defined as:

[0089] E = ∑ k C k / N, Var = ∑ k (C k -E)(C k -E) T / N, N is the number of pixels in the image, and C represents the pixel value of a pixel in the image, which can be represented by the three channels of R, G, and B.

[0090] λ 1 and η 1 are both preset constants. For example, λ 1 = 5e-3; η 1 = 1e-2, etc.

[0091] The second - level generative adversarial network is a super - resolution (SR) network, and its input is a fused image (which can also be image features extracted from the fused image) obtained by fusing the original first image and the result of the first - level generative adversarial network (i.e., the enhanced low - frequency image). Specifically, it can be adding the pixels at the same positions of the first image and the enhanced low - frequency image, that is, by pixel addition, adding the pixels at the same position of the two images to obtain the fused image. By fusing the features of the above - mentioned images, the original image features can be introduced into the enhanced low - frequency image, thereby reducing artifacts and increasing the details of the texture. Among them, the second - level generative network generates an enhanced image based on the fused image, and the second - level discriminative network determines whether the enhanced image is consistent with the second image.

[0092] During the training process of the above - mentioned image enhancement model, the embodiment of the present invention fuses the first image and the enhanced low - frequency image to obtain a fused image, and generates an enhanced image of the first image based on the fused image; then, using the second image as the target image, the enhanced image of the first image is used to train the second - level generative adversarial network until a preset training end condition is met. For example, taking the second image as the target image, by calculating and updating the loss function of the second - level generative adversarial network, the value of the loss function is continuously updated until the preset training end condition is met.

[0093] The following provides a specific representation form of the loss function Loss_G_2 of the second - level generative adversarial network. It should be noted that the following formula is only an example of a loss function that can be adopted in the embodiment of the present invention and is not used to limit the present invention:

[0094] Loss_G_2 = η 2 Lcobi - hf+η 3 Lcobi 2 +λ 2 LG 2 +η 4 Lcolor 2 ;

[0095] Among them, Lcobi 2 represents the context - bilateral loss function for the enhanced image and the second image. LG 2 represents the adversarial loss function for the enhanced image and the second image. Lcolor 2 represents the color loss function for the enhanced image and the second image. η 2 、η 3 、λ 2 and η 4 are all preset constants.

[0096] The Lcobi-hf represents a contextual bilateral loss function for the high-frequency features of the enhanced image and the high-frequency features of the second image. A large number of background regions and texture details contained in the polar region are severely stretched. The above loss function only focuses on the high-frequency part of the image, thereby reducing the generation of artifacts. Since the edge information and artifacts of the image usually only exist in the high-frequency part, it is only necessary to compare the similarity between the high-frequency of the generated enhanced image and the high-frequency of the second image. Especially when not affected by the low-frequency background, artifacts can be better removed. For the problem of misalignment of training images, the embodiments of the present invention can draw on the implementation of the CoBi-loss function. For example, the CoBi-Loss function is an improvement of the Contextual loss. The CX loss function calculates the similarity of the distance between feature points; Cobi-loss imposes spatial constraints on this basis to obtain a better enhanced image. The Cobi-loss function pays more attention to VGG characteristics. A representation form of the Cobi-loss function in the prior art is as follows:

[0097]

[0098] Different from the Cobi-loss function that pays more attention to VGG characteristics, the embodiments of the present invention directly calculate the similarity of distances in the high-frequency image using RGB and spatial information. For example, 2x2 image blocks in the high-frequency features can be used as feature points, where, The similarity of 2x2 image blocks on the high-frequency features of the enhanced image and the high-frequency features of the second image can be used for replacement. The above loss function only focuses on the high-frequency part of the generated image, so it can directly reduce the generation of artifacts (especially in the extreme region, which contains a large number of background regions and the texture details are severely stretched).

[0099] As can be seen from the above, the embodiments of the present invention extract low-frequency features in the first-level generative adversarial network and enhance the low-frequency features; then add the original features of the image to the enhanced low-frequency features to reduce the influence of noise, and finally reduce the generation of artifacts and retain more texture details. In addition, the embodiments of the present invention introduce a high-frequency similarity loss, which only focuses on the high-frequency part of the generated image and can directly reduce the generation of artifacts. In addition, the embodiments of the present invention also replace the L1 loss in the existing network loss function with a color loss function. The color loss function pays more attention to the overall distribution of data and helps to reduce the generation of artifacts. In addition, the embodiments of the present invention also replace the perceptual loss in the existing network loss function with CoBi loss, which is insensitive to data alignment and also helps to reduce the generation of artifacts.

[0100] Based on the above image enhancement method, the embodiments of the present invention also provide an image enhancement device 40, as Figure 4As shown in the figure, it includes:

[0101] An acquisition module 41, configured to acquire multiple sets of training data, each set of training data including a first image and a second image, wherein the image quality of the second image is better than that of the first image;

[0102] A training module 42, configured to construct an image enhancement model including a first-level generative adversarial network and a second-level generative adversarial network, and use the multiple sets of training data to train the image enhancement model to obtain a trained image enhancement model. Among them, taking the second image as the target image, an enhanced low-frequency image generated based on the low-frequency features of the first image is used to train the first-level generative adversarial network; taking the second image as the target image, an enhanced image generated based on the fused image of the first image is used to train the second-level generative adversarial network; the fused image is obtained by fusing the first image and the enhanced low-frequency image;

[0103] An image processing module 43, which inputs a third image to be enhanced into the image enhancement model and outputs a fourth image with enhanced image quality.

[0104] Through the above modules, the image enhancement device 40 of the embodiment of the present invention can reduce the generation of artifacts while retaining more texture details and improving the quality of the image.

[0105] Optionally, the first-level generative adversarial network includes:

[0106] An octave convolution module, configured to extract low-frequency features from the first image;

[0107] A first generation network, configured to perform image enhancement based on the low-frequency features of the first image and generate an enhanced low-frequency image of the first image;

[0108] A first adversarial network, configured to determine whether the enhanced low-frequency image is consistent with the second image;

[0109] The training module is further configured to use the enhanced low-frequency image of the first image to train the first-level generative adversarial network with the second image as the target image until a preset training end condition is met.

[0110] As an implementation manner, the loss function Loss_G_1 of the first-level generative adversarial network is:

[0111] Loss_G_1 = Lcobi 1 + λ 1 LG 1 + η 1 Lcolor 1 ;

[0112] Among them, Lcobi 1 represents the context bilateral loss function for the enhanced low-frequency image and the second image; LG 1 represents the adversarial loss function for the enhanced low-frequency image and the second image; Lcolor 1 represents the color loss function for the enhanced low-frequency image and the second image; λ 1 and η 1 are both preset constants.

[0113] Optionally, the second-level generative adversarial network includes:

[0114] A fusion module for adding pixels at the same positions of the first image and the enhanced low-frequency image to obtain a fused image;

[0115] A second generation network for generating an enhanced image of the first image based on the fused image;

[0116] A second adversarial network for determining whether the enhanced image is consistent with the second image;

[0117] The training module is further configured to use the second image as the target image and train the second-level generative adversarial network with the enhanced image of the first image until a preset training end condition is met.

[0118] As an implementation, the loss function Loss_G_2 of the second-level generative adversarial network is:

[0119] Loss_G_2 = η 2 Lcobi - hf + η 3 Lcobi 2 + λ 2 LG 2 + η 4 Lcolor 2 ;

[0120] Among them, Lcobi 2 represents the context bilateral loss function for the enhanced image and the second image; Lcobi - hf represents the context bilateral loss function for the high-frequency features of the enhanced image and the high-frequency features of the second image; LG 2 represents the adversarial loss function for the enhanced image and the second image; Lcolor 2 represents the color loss function for the enhanced image and the second image; η 2 、η 3 、λ 2 and η 4 are both preset constants.

[0121] Optionally, the content captured by the first image and the second image in the same set of training data is the same.

[0122] Optionally, the first image is an equidistant cylindrical projection or a perspective view, and the second image is a perspective view.

[0123] Optionally, the image quality of the second image being better than that of the first image includes at least one of the following:

[0124] The resolution of the second image is greater than that of the first image;

[0125] The signal-to-noise ratio of the second image is higher than that of the first image;

[0126] The chromatic aberration of the second image is lower than that of the first image.

[0127] Please refer to Figure 5 , an embodiment of the present invention also provides a hardware structure block diagram of an image enhancement device, as Figure 5 shown. The image enhancement device 500 includes:

[0128] A processor 502; and

[0129] A memory 504, in which computer program instructions are stored,

[0130] wherein, when the computer program instructions are run by the processor, the processor 502 is caused to perform the following steps:

[0131] Obtain multiple sets of training data, each set of training data including a first image and a second image, wherein the image quality of the second image is better than that of the first image;

[0132] Construct an image enhancement model including a first-level generative adversarial network and a second-level generative adversarial network, and use the multiple sets of training data to train the image enhancement model to obtain a trained image enhancement model. Among them, using the second image as the target image, the enhanced low-frequency image generated based on the low-frequency features of the first image is used to train the first-level generative adversarial network; using the second image as the target image, the enhanced image generated based on the fused image of the first image is used to train the second-level generative adversarial network; the fused image is obtained by fusing the first image and the enhanced low-frequency image;

[0133] Input the third image to be enhanced into the image enhancement model, and output the fourth image after image enhancement.

[0134] Furthermore, as Figure 5As shown, the image enhancement device 500 may further include a network interface 501, an input device 503, a hard disk 505, and a display device 506.

[0135] The various interfaces and devices described above may be interconnected through a bus architecture. The bus architecture may include any number of interconnected buses and bridges. Specifically, one or more processors with computing capabilities represented by the processor 502, the processor may include a central processing unit (CPU) and / or a graphics processing unit (GPU), and various circuits of one or more memories represented by the memory 504 are connected together. The bus architecture may also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits. It can be understood that the bus architecture is used to implement the connection and communication between these components. In addition to the data bus, the bus architecture also includes a power bus, a control bus, and a status signal bus, which are well known in the art and will not be described in detail herein.

[0136] The network interface 501 may be connected to a network (such as the Internet, a local area network, etc.), receive data (such as training data) from the network, and may save the received data in the hard disk 505.

[0137] The input device 503 may receive various instructions input by an operator and send them to the processor 502 for execution. The input device 503 may include a keyboard or a pointing device (for example, a mouse, a trackball, a touchpad, or a touch screen, etc.).

[0138] The display device 506 may display the results obtained by the processor 502 executing instructions, such as displaying the progress of model training and the answer prediction results, etc.

[0139] The memory 504 is used to store the programs and data necessary for the operation of the operating system, as well as data such as intermediate results during the calculation of the processor 502.

[0140] It can be understood that the memory 504 in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. The memory 504 of the devices and methods described herein is intended to include, but is not limited to, these and any other suitable types of memories.

[0141] In some embodiments, the memory 504 stores the following elements, executable modules, or data structures, or subsets thereof, or extended sets thereof: an operating system 5041 and application programs 5042.

[0142] Among them, the operating system 5041 includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., and is used to implement various basic services and handle hardware-based tasks. The application programs 5042 include various application programs, such as a Browser, etc., and are used to implement various application services. The program for implementing the method of the embodiment of the present invention may be included in the application programs 5042.

[0143] The image enhancement method disclosed in the above embodiments of the present invention may be applied to the processor 502 or implemented by the processor 502. The processor 502 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above image enhancement method may be completed by the integrated logic circuit in the hardware of the processor 502 or instructions in the form of software. The above-mentioned processor 502 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and may implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention may be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 504, and the processor 502 reads the information in the memory 504 and combines its hardware to complete the steps of the above method.

[0144] It can be understood that these embodiments described herein may be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit may be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in this application, or a combination thereof.

[0145] For software implementation, the techniques described herein can be implemented by modules (e.g., procedures, functions, etc.) that execute the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented within the processor or external to the processor.

[0146] Specifically, when the computer program is executed by the processor 502, the following steps can also be implemented:

[0147] Extract low-frequency features from the first image;

[0148] Perform image enhancement based on the low-frequency features to generate an enhanced low-frequency image of the first image;

[0149] Use the enhanced low-frequency image of the first image to train the first-level generative adversarial network with the second image as the target image until a preset training end condition is met.

[0150] Specifically, the loss function Loss_G_1 of the first-level generative adversarial network is: Loss_G_1 = Lcobi 1 + λ 1 LG 1 + η 1 Lcolor 1 ;

[0151] Wherein, Lcobi 1 represents the context bilateral loss function for the enhanced low-frequency image and the second image; LG 1 represents the adversarial loss function for the enhanced low-frequency image and the second image; Lcolor 1 represents the color loss function for the enhanced low-frequency image and the second image; λ 1 and η 1 are both preset constants.

[0152] Specifically, when the computer program is executed by the processor 502, the following steps can also be implemented:

[0153] Add the pixels at the same positions of the first image and the enhanced low-frequency image, and generate an enhanced image of the first image based on the fused image;

[0154] Use the enhanced image of the first image to train the second-level generative adversarial network with the second image as the target image until a preset training end condition is met.

[0155] Specifically, the loss function Loss_G_2 of the second-level generative adversarial network is: Loss_G_2 = η 2 Lcobi-hf + η 3 Lcobi2 +λ 2 LG 2 +η 4 Lcolor 2 ;

[0156] Wherein, Lcobi 2 represents the context bilateral loss function for the enhanced image and the second image; Lcobi-hf represents the bilateral loss function for the high-frequency features of the enhanced image and the high-frequency features of the second image; LG 2 represents the adversarial loss function for the enhanced image and the second image; Lcolor 2 represents the color loss function for the enhanced image and the second image; η 2 η 3 λ 2 η 4 are all preset constants.

[0157] Optionally, the content captured by the first image and the second image in the same set of training data is the same.

[0158] Optionally, the first image is an equidistant cylindrical projection or a perspective view, and the second image is a perspective view.

[0159] Optionally, the image quality of the second image being better than that of the first image includes at least one of the following:

[0160] The resolution of the second image is greater than that of the first image;

[0161] The signal-to-noise ratio of the second image is higher than that of the first image;

[0162] The chromatic aberration of the second image is lower than that of the first image.

[0163] When this program is executed by a processor, it can implement all implementation manners in the above image enhancement method and achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0164] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0165] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0166] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0167] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present invention.

[0168] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0169] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the image enhancement method described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0170] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention and should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. An image enhancement method, characterized in that, comprising: obtaining multiple sets of training data, each set of training data including a first image and a second image, wherein the image quality of the second image is better than that of the first image; constructing an image enhancement model including a first - stage generative adversarial network and a second - stage generative adversarial network, training the image enhancement model using the multiple sets of training data to obtain a trained image enhancement model, wherein, taking the second image as the target image, training the first - stage generative adversarial network based on the enhanced low - frequency image of the first image output by the first - stage generative adversarial network; the enhanced low - frequency image of the first image is generated by using the first - stage generative adversarial network with the first image as the input; taking the second image as the target image, training the second - stage generative adversarial network based on the enhanced image output by the second - stage generative adversarial network; the enhanced image is generated by using the second - stage generative adversarial network with the fused image of the first image as the input; the fused image of the first image is obtained by fusing the first image and the enhanced low - frequency image; inputting a third image to be enhanced into the image enhancement model, and outputting a fourth image with enhanced image.

2. The image enhancement method according to claim 1, characterized in that, the training of the first - stage generative adversarial network by taking the second image as the target image and based on the enhanced low - frequency image of the first image output by the first - stage generative adversarial network includes: extracting low - frequency features from the first image; performing image enhancement based on the low - frequency features to generate the enhanced low - frequency image of the first image; taking the second image as the target image, and training the first - stage generative adversarial network using the enhanced low - frequency image of the first image until a preset training end condition is met.

3. The image enhancement method according to claim 2, characterized in that, The loss function Loss_G_1 of the first-level generative adversarial network is: Loss_G_1 = Lcobi 1 + λ 1 LG 1 + η 1 Lcolor 1 ; Among them, Lcobi 1 represents the context bilateral loss function for the enhanced low-frequency image and the second image; LG 1 represents the adversarial loss function for the enhanced low-frequency image and the second image; Lcolor 1 represents the color loss function for the enhanced low-frequency image and the second image; λ 1 and η 1 are both preset constants.

4. The image enhancement method according to any one of claims 1 to 3, characterized in that, the training of the second - stage generative adversarial network by taking the second image as the target image and based on the enhanced image output by the second - stage generative adversarial network includes: adding pixels at the same positions of the first image and the enhanced low - frequency image to obtain a fused image, and generating the enhanced image of the first image based on the fused image; taking the second image as the target image, and training the second - stage generative adversarial network using the enhanced image of the first image until a preset training end condition is met.

5. The image enhancement method according to claim 4, characterized in that, The loss function Loss_G_2 of the second - level generative adversarial network is: Loss_G_2 = η 2 Lcobi - hf+η 3 Lcobi 2 +λ 2 LG 2 +η 4 Lcolor 2 ; Among them, Lcobi 2 represents the context bilateral loss function for the enhanced image and the second image; Lcobi-hf represents the context bilateral loss function for the high-frequency features of the enhanced image and the high-frequency features of the second image; LG 2 represents the adversarial loss function for the enhanced image and the second image; Lcolor 2 represents the color loss function for the enhanced image and the second image; η 2 、η 3 、λ 2 and η 4 are all preset constants.

6. The image enhancement method according to claim 4, characterized in that, the content captured by the first image and the second image in the same set of training data is the same.

7. The image enhancement method according to claim 1, characterized in that, the first image is an equidistant cylindrical projection map or a perspective view, and the second image is a perspective view.

8. The image enhancement method according to claim 1, characterized in that, The image quality of the second image being better than that of the first image includes at least one of the following: The resolution of the second image is greater than that of the first image; The signal-to-noise ratio of the second image is higher than that of the first image; The chromatic aberration of the second image is lower than that of the first image.

9. An image enhancement device, characterized in that, it includes: An acquisition module, configured to acquire multiple sets of training data, each set of training data including a first image and a second image, wherein the image quality of the second image is better than that of the first image; A training module, configured to construct an image enhancement model including a first-level generative adversarial network and a second-level generative adversarial network, and use the multiple sets of training data to train the image enhancement model to obtain a trained image enhancement model, wherein, taking the second image as the target image, based on the enhanced low-frequency image of the first image output by the first-level generative adversarial network, training the first-level generative adversarial network; the enhanced low-frequency image of the first image is generated by using the first-level generative adversarial network with the first image as the input; taking the second image as the target image, based on the enhanced image output by the second-level generative adversarial network, training the second-level generative adversarial network; the enhanced image is generated by using the second-level generative adversarial network with the fused image of the first image as the input; the fused image of the first image is obtained by fusing the first image and the enhanced low-frequency image; An image processing module, which inputs the third image to be enhanced into the image enhancement model and outputs the fourth image after image enhancement.

10. The image enhancement device according to claim 9, characterized in that, The first-level generative adversarial network includes: An octave convolution module, configured to extract low-frequency features from the first image; A first generation network, configured to perform image enhancement based on the low-frequency features of the first image and generate the enhanced low-frequency image of the first image; A first adversarial network, configured to determine whether the enhanced low-frequency image is consistent with the second image; The training module is further configured to take the second image as the target image and use the enhanced low-frequency image of the first image to train the first-level generative adversarial network until a preset training end condition is met.

11. The image enhancement device according to claim 9 or 10, characterized in that, The second-level generative adversarial network includes: A fusion module, configured to add the pixels at the same positions of the first image and the enhanced low-frequency image to obtain a fused image; A second generation network, configured to generate the enhanced image of the first image based on the fused image; A second adversarial network, configured to determine whether the enhanced image is consistent with the second image; The training module is further configured to take the second image as the target image and use the enhanced image of the first image to train the second-level generative adversarial network until a preset training end condition is met.

12. An image enhancement device, including: A processor; and A memory, in which computer program instructions are stored, wherein, when the computer program instructions are run by the processor, the processor is caused to perform the following steps: Obtain multiple sets of training data, each set of training data including a first image and a second image, wherein the image quality of the second image is better than that of the first image; Construct an image enhancement model including a first-level generative adversarial network and a second-level generative adversarial network, and use the multiple sets of training data to train the image enhancement model to obtain a trained image enhancement model, wherein, taking the second image as the target image, based on the enhanced low-frequency image of the first image output by the first-level generative adversarial network, train the first-level generative adversarial network; the enhanced low-frequency image of the first image is generated by taking the first image as the input and using the first-level generative adversarial network; taking the second image as the target image, based on the enhanced image output by the second-level generative adversarial network, train the second-level generative adversarial network; the enhanced image is generated by taking the fused image of the first image as the input and using the second-level generative adversarial network; the fused image of the first image is obtained by fusing the first image and the enhanced low-frequency image; Input a third image to be enhanced into the image enhancement model, and output a fourth image after image enhancement.

13. A computer-readable storage medium, the computer-readable storage medium stores a computer program, characterized in that when the computer program is executed by a processor, it implements the steps of the image enhancement method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • High-frequency sensitive GAN network for LDCT image denoising

    CN110517198A

  • No-reference low-illumination image enhancement method and system based on generative adversarial network

    CN111798400A