Image enhancement method and device, model and training method thereof, and electronic equipment

By performing multi-level downsampling and upsampling of images, and combining pre-learned fusion parameters to fuse coded features and decoded features, the problem of large amount of calculation and poor effect of U-shaped network is solved, and efficient image enhancement and high-definition image generation are achieved.

CN120278915APending Publication Date: 2025-07-08SMARTER SILICON (SHANGHAI) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510423616.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-04
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing U-shaped network has a large amount of calculation during image enhancement, resulting in slower speed and difficult deployment, and the classic U-shaped network image enhancement effect is poor.

Method used

By performing downsampling and upsampling of the image at multiple levels, combining pre-learning fusion parameters to fuse the encoded features and decoded features, omitting the low-resolution feature search module, and using the encoding module, decoding module and fusion module of the enhanced network for image enhancement.

Benefits of technology

Improves image enhancement effect, reduces computational volume, improves processing speed, facilitates deployment, and generates high-definition images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278915A_ABST
    Figure CN120278915A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an image enhancement method and device, a model and a training method thereof, and electronic equipment, and the method comprises the steps: carrying out the down-sampling of a plurality of levels of a first image with a lower definition, and obtaining the coding features of a plurality of different sampling parameters; performing multi-level up-sampling on the coding features obtained by the down-sampling of the last level to obtain decoding features of a plurality of different sampling parameters; wherein the up-sampling of each non-first level comprises the following steps: fusing decoding features obtained by the up-sampling of the previous level and coding features of the same sampling parameter to obtain fused features, and carrying out up-sampling on the fused features; the fusion feature obtaining process of at least one non-first level comprises the following steps: fusing decoding features obtained by upsampling of a previous level and coding features of the same sampling parameter through a pre-learned fusion parameter; and obtaining a second image with higher definition based on the decoding feature obtained by the upsampling of the last level.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of an application with an application date of November 04, 2024, an application number of 202411563375.9, and an invention title of "Image Enhancement Method, Device, Model and Its Training Method, Electronic Device". Technical Field

[0002] This application relates to the field of image processing technology, and more specifically, to an image enhancement method, device, model and its training method, and electronic device. Background Art

[0003] Image enhancement is a method to improve the visual effect of images. Generally speaking, the main purpose of image enhancement is to improve the visual effect of images and enhance the clarity of images.

[0004] Currently, when enhancing a low-quality original image into a high-quality image, the classic U-shaped network is used in the image enhancement solution. However, the image enhancement effect of the classic U-shaped network is poor. In order to improve the image enhancement effect of the U-shaped network, a solution has been proposed to add a low-resolution feature search module between the encoding module and the decoding module of the U-shaped network to extract strong discriminative features required for enhancement, and then the decoding module decodes the strong discriminative features to obtain a high-quality image.

[0005] Although adding a low-resolution feature search module can improve the image enhancement effect, the scale of the low-resolution feature search module is relatively large, which significantly increases the computational complexity of the entire image enhancement network, resulting in a slow network speed and difficult network deployment. Summary of the Invention

[0006] The purpose of this application is to provide an image enhancement method, device, model and its training method, and electronic device, including the following technical solutions:

[0007] An image enhancement method, the method includes:

[0008] Performing downsampling on the first image at multiple levels to obtain encoded features with multiple different sampling parameters;

[0009] Performing upsampling on the encoded features obtained by the downsampling at the last level at multiple levels to obtain decoded features with multiple different sampling parameters; wherein, the upsampling of each non-first level includes: fusing the decoded features obtained by the upsampling of the previous level and the encoded features with the same sampling parameter to obtain a fused feature, and performing upsampling on the fused feature; the process of obtaining the fused feature at at least one non-first level includes: fusing the decoded features obtained by the upsampling of the previous level and the encoded features with the same sampling parameter through pre-learned fusion parameters;

[0010] The second image is obtained based on the decoded features obtained by upsampling the last level; the sharpness of the second image is higher than that of the first image.

[0011] In a possible implementation, the fusion of the decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameter by the pre-learned fusion parameters includes:

[0012] Based on the fusion parameters, the decoded features obtained by upsampling the previous level, and the encoded features with the same sampling parameter, a first weight of the encoded features with the same sampling parameter as the decoded features obtained by upsampling the previous level is calculated;

[0013] Based on the first weight, weighted fusion of the decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameter is performed.

[0014] In a possible implementation, the weighted fusion of the decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameter based on the first weight includes:

[0015] Based on the first weight, a second weight of the decoded features obtained by upsampling the previous level is obtained;

[0016] Based on the first weight and the second weight, weighted calculation of the decoded features obtained by upsampling the previous level and its encoded features with the same sampling parameter is performed.

[0017] In a possible implementation, the weighted fusion of the decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameter based on the first weight includes:

[0018] The decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameter are summed to obtain an initial fusion feature;

[0019] Based on the first weight, a second weight of the decoded features obtained by upsampling the previous level is obtained;

[0020] Based on the first weight and the second weight, weighted calculation of the initial fusion feature and the decoded features obtained by upsampling the previous level is performed.

[0021] In a possible implementation, the process of performing downsampling on the first image at multiple levels, performing upsampling on the encoded features obtained by downsampling the last level at multiple levels, and obtaining the second image based on the decoded features obtained by upsampling the last level includes:

[0022] Performing multiple levels of downsampling on the first image through the first encoding module of the enhanced network;

[0023] Performing multiple levels of upsampling on the encoded features obtained by downsampling the last level through the decoding module of the enhanced network, where the upsampling of each non-first level includes: obtaining the fused features obtained by the fusion module of the enhanced network by fusing the decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameter, and performing upsampling on the fused features; the process of obtaining the fused features in at least one non-first level includes: the fusion module fusing the decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameter through pre-learned fusion parameters;

[0024] Processing the decoded features obtained by upsampling the last level through the output module of the enhanced network to obtain the second image.

[0025] In a possible implementation manner, the enhanced network is trained as follows:

[0026] Performing unsupervised training on the initial network obtained by the second encoding module, the decoding module, and the output module through the high-quality images in the first image set to obtain a pre-trained network; the first image set includes multiple low-quality images and the corresponding high-quality images for each low-quality image; the second encoding module is used to perform multiple levels of downsampling on the input high-quality images to obtain encoded features with multiple different sampling parameters; based on the pre-trained network, performing supervised training on the first encoding module and the fusion module through the first image set to obtain the trained first encoding module and fusion module; the trained first encoding module and fusion module, and the decoding module and output module in the pre-trained network constitute the enhanced network.

[0027] An image enhancement model, including: a first encoding module, a decoding module, a fusion module, and an output module;

[0028] The first encoding module is used to perform multiple levels of downsampling on the first image to obtain encoded features with multiple different sampling parameters;

[0029] The decoding module is used to perform upsampling of multiple levels on the encoded features obtained by downsampling the last level, so as to obtain decoded features with multiple different sampling parameters; wherein, the upsampling of each non-first level includes: obtaining the fused features obtained by the fusion module by fusing the decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameter, and performing upsampling on the fused features; the process of obtaining the fused features in at least one non-first level includes: the fusion module fuses the decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameter through pre-learned fusion parameters;

[0030] The output module is used to obtain a second image based on the decoded features obtained by upsampling the last level; the clarity of the second image is higher than that of the first image.

[0031] In a possible implementation manner, when the fusion module fuses the decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameter through pre-learned fusion parameters, it is used to:

[0032] Based on the fusion parameters, the decoded features obtained by upsampling the previous level, and the encoded features with the same sampling parameter, calculate a first weight of the encoded features with the same sampling parameter as the decoded features obtained by upsampling the previous level;

[0033] Perform weighted fusion on the decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameter based on the first weight.

[0034] A model training method is used to train the image enhancement model described in any one of the foregoing items. The model training method includes:

[0035] Perform unsupervised training on the initial network obtained by the second encoding module, the decoding module, and the output module through the high-quality images in the first image set to obtain a pre-trained network; the first image set includes multiple low-quality images and the high-quality images corresponding to each low-quality image; the second encoding module is used to perform downsampling of multiple levels on the input high-quality images to obtain encoded features with multiple different sampling parameters;

[0036] Based on the pre-trained network, perform supervised training on the first encoding module and the fusion module through the first image set to obtain the trained first encoding module and fusion module; the trained first encoding module and fusion module, and the decoding module and output module in the pre-trained network constitute the image enhancement model.

[0037] In a possible implementation manner, based on the pre-trained network, the first encoding module is supervised and trained by the first image set, including:

[0038] For any low-quality image in the first image set, input the any low-quality image into the first encoding module to obtain encoding features of the any low-quality image with multiple different sampling parameters;

[0039] Input the high-quality image corresponding to the any low-quality image into the second encoding module in the pre-trained network to obtain encoding features of the high-quality image corresponding to the any low-quality image with multiple different sampling parameters;

[0040] Taking the first difference between the encoding features with the same sampling parameters of the any low-quality image and its corresponding high-quality image as smaller and smaller as the target, update the parameters of the first encoding module.

[0041] An image enhancement device, the device includes:

[0042] An encoding unit, configured to perform multi-level downsampling on a first image to obtain encoding features with multiple different sampling parameters;

[0043] A decoding unit, configured to perform multi-level upsampling on the encoding features obtained by the downsampling of the last level to obtain decoding features with multiple different sampling parameters; wherein, the upsampling of each non-first level includes: obtaining a fusion unit to fuse the decoding features obtained by the upsampling of the previous level and the encoding features with the same sampling parameter to obtain a fusion feature, and performing upsampling on the fusion feature; the process of obtaining the fusion feature of at least one non-first level includes: fusing the decoding features obtained by the upsampling of the previous level and the encoding features with the same sampling parameter through pre-learned fusion parameters;

[0044] An output unit, configured to obtain a second image based on the decoding features obtained by the upsampling of the last level; the clarity of the second image is higher than the clarity of the first image.

[0045] A model training device, configured to train the image enhancement model according to any one of the above, the model training device includes:

[0046] The first training unit is used to perform unsupervised training on the initial network obtained by the second encoding module, the decoding module, and the output module with high-quality images in the first image set to obtain a pre-trained network; the first image set includes multiple low-quality images and the high-quality images corresponding to each low-quality image; the second encoding module is used to perform downsampling at multiple levels on the input high-quality images to obtain encoded features with multiple different sampling parameters; the structure of the second encoding module is the same as or different from the structure of the first encoding module;

[0047] The second training unit is used to perform supervised training on the first encoding module and the fusion module with the first image set based on the pre-trained network to obtain a trained first encoding module and fusion module; the trained first encoding module and fusion module, and the decoding module and output module in the pre-trained network constitute the enhanced network.

[0048] A readable storage medium stores a computer program thereon. When the computer program is executed by a processor, it implements each step of the image enhancement method and / or the model training method described in any one of the above.

[0049] An electronic device includes:

[0050] A memory for storing a program;

[0051] A processor for calling and executing the program in the memory and implementing each step of the image enhancement method and / or the model training method described in any one of the above by executing the program.

[0052] It can be seen from the above solutions that an image enhancement method, device, model and its training method, electronic device and storage medium provided by the present application perform downsampling at multiple levels on a first image with lower clarity to obtain encoded features with multiple different sampling parameters; perform upsampling at multiple levels on the encoded features obtained by the downsampling of the last level to obtain decoded features with multiple different sampling parameters; wherein, each non-first-level upsampling includes: fusing the decoded features obtained by the upsampling of the previous level and the encoded features with the same sampling parameter to obtain a fused feature, and performing upsampling on the fused feature; the process of obtaining the fused feature in at least one non-first level includes: fusing the decoded features obtained by the upsampling of the previous level and the encoded features with the same sampling parameter through pre-learned fusion parameters; obtaining a second image with higher clarity based on the decoded features obtained by the upsampling of the last level. Description of the Drawings

[0053] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0054] Figure 1 It is a flowchart of an implementation of the image enhancement method provided by the embodiments of the present application;

[0055] Figure 2 It is a flowchart of an implementation of fusing the decoded features obtained by upsampling the previous level through pre-learned fusion parameters and the encoded features with the same sampling parameters provided by the embodiments of the present application;

[0056] Figure 3 It is a flowchart of an implementation of weighted fusion of the decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameters based on the first weight provided by the embodiments of the present application;

[0057] Figure 4 It is another flowchart of an implementation of weighted fusion of the decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameters based on the first weight provided by the embodiments of the present application;

[0058] Figure 5 It is a schematic structural diagram of an enhancement network provided by the embodiments of the present application;

[0059] Figure 6 It is another schematic structural diagram of an enhancement network provided by the embodiments of the present application;

[0060] Figure 7 It is yet another schematic structural diagram of an enhancement network provided by the embodiments of the present application;

[0061] Figure 8 It is a schematic structural diagram of an initial network provided by the embodiments of the present application;

[0062] Figure 9 It is another schematic structural diagram of an initial network provided by the embodiments of the present application;

[0063] Figure 10 It is a flowchart of an implementation of supervised training of the first encoding module through the first image set based on a pre-trained network provided by the embodiments of the present application;

[0064] Figure 11 It is a flowchart of an implementation of supervised training of the fusion module through the first image set based on a pre-trained network provided by the embodiments of the present application;

[0065] Figure 12 A schematic structural diagram of an image enhancement device provided by an embodiment of the present application;

[0066] Figure 13 A schematic structural diagram of a model training device provided by an embodiment of the present application;

[0067] Figure 14 A schematic structural diagram of an electronic device provided by an embodiment of the present application.

[0068] Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification, claims and the above-mentioned drawings are used to distinguish similar parts, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than that illustrated here. Detailed implementation manners

[0069] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0070] As mentioned above, in order to improve the image enhancement effect of the U-shaped network, one implementation is to embed a low-resolution feature search module after the encoding module and before the decoding module of the U-shaped network to extract strong discriminative features required for enhancement from the feature map with the smallest resolution output by the encoding module, and input the strong discriminative features into the decoding module for decoding processing. The low-resolution feature search module usually adopts a convolution module with a large kernel (convolution kernel), a non-local module, or a transformer module, and these modules often have a large amount of computation, and problems such as softmax and layernorm that are difficult to quantize or slow computing on the neural network processing unit (NPU) side will also be encountered during deployment.

[0071] In order to improve the processing speed of image enhancement and make the deployment more friendly, the solution of the present application is proposed.

[0072] The image enhancement method and device, image enhancement model, and model training method provided by the embodiments of the present application can be used in an electronic device, and the electronic device has a processor capable of processing images, such as a CPU, a GPU, etc.

[0073] The electronic device can be a terminal device or a server. The server can be a single server, a server cluster, or a cloud server, etc.

[0074] As Figure 1 shown, it is a flowchart of an implementation of the image enhancement method provided by an embodiment of the present application, which may include:

[0075] Step S101: Perform downsampling on the first image at multiple levels to obtain encoded features with multiple different sampling parameters.

[0076] The first image is a lower-quality image to be enhanced. The first image can be an RGB image, a grayscale image, a depth image, or an image in other formats. The present application does not make specific limitations on the format of the first image. The first image can be an image collected by an image acquisition device, or an image edited by an image editor, or an AI-generated image, etc.

[0077] The input of the first downsampling level is the first image. Starting from the second downsampling level, the input of each downsampling level is the encoded feature obtained from the previous downsampling level. That is, the input of the first downsampling level is the first image, and the input of the i-th (i = 2, 3,..., I; I is the total number of downsampling levels) downsampling level is the encoded feature output from the i-1-th downsampling level.

[0078] The sampling parameter of the encoded feature can refer to the resolution or size of the encoded feature. That is to say, perform downsampling on the first image at multiple levels to obtain encoded features with multiple different resolutions, or obtain encoded features with multiple different sizes.

[0079] The encoded feature obtained at each downsampling level is a feature map of the first image. Therefore, the downsampling at different levels obtains feature maps of different resolutions of the first image, or the downsampling at different levels obtains feature maps of different sizes of the first image. Among them, in the encoded features obtained from two adjacent downsampling levels, the resolution or size of the encoded feature obtained from the latter downsampling level is smaller than that of the encoded feature obtained from the previous downsampling level.

[0080] Optionally, any of the following downsampling methods can be used to perform downsampling on the first image at multiple levels: max pooling, convolution, average pooling, etc. The downsampling methods used in different downsampling levels can be the same or different.

[0081] Step S102: Perform upsampling on the encoded feature obtained from the downsampling of the last level at multiple levels to obtain decoded features with multiple different sampling parameters.

[0082] Among them, the input of the first upsampling level is the encoded feature output by the last downsampling level. That is to say, the object processed by the upsampling of the first level is the encoded feature obtained by the downsampling of the last level.

[0083] The upsampling of each non-first level includes: fusing the decoded feature obtained by the upsampling of the previous level and the encoded feature with the same sampling parameter to obtain a fused feature, and performing upsampling on the fused feature; the process of obtaining the fused feature for at least one non-first level includes: fusing the decoded feature obtained by the upsampling of the previous level and the encoded feature with the same sampling parameter through pre-learned fusion parameters.

[0084] Among the decoded features obtained by two adjacent upsampling levels, the resolution or size of the decoded feature obtained by the latter upsampling level is greater than that of the decoded feature obtained by the previous upsampling level.

[0085] In this application, the number of upsampling levels is the same as the number of downsampling levels. Among them, the sampling parameter of the encoded feature obtained by the jth (j = 1, 2, 3,..., I - 1) upsampling is the same as the sampling parameter of the encoded feature obtained by the (I - j)th downsampling level.

[0086] In this application, starting from the second upsampling level, instead of performing upsampling on the output of the previous upsampling level, first fuse the decoded feature output by the previous upsampling level with the encoded feature having the same sampling parameter to obtain a fused feature, and perform upsampling on the fused feature. When fusing the encoded feature and the decoded feature with the same sampling parameter, at least part of the fusion process corresponding to the upsampling level is obtained by fusing the encoded feature and the decoded feature with the same sampling parameter through pre-learned fusion parameters. That is to say, only part of the fusion process corresponding to the upsampling level can use the pre-learned fusion parameters, or the fusion processes corresponding to the I - 1 upsampling levels can all fuse the encoded feature and the decoded feature with the same sampling parameter through pre-learned fusion parameters.

[0087] In the case where only part of the fusion process corresponding to the upsampling level uses the pre-learned fusion parameters for fusion, the fusion processes corresponding to other upsampling levels can directly add the encoded feature and the decoded feature with the same sampling parameter.

[0088] Through the above multi-level upsampling, the decoded feature obtained by the last upsampling level has strong discriminative features required for image enhancement.

[0089] Optionally, any of the following upsampling methods can be used to perform multi-level upsampling on the encoded features obtained by downsampling the last level: inverse max pooling, transposed convolution, inverse mean pooling, etc. The upsampling methods used for different upsampling levels can be the same or different.

[0090] Step S103: Obtain a second image based on the decoded features obtained by upsampling the last level. The clarity of the second image is higher than that of the first image. The content in the second image is the same as the content in the first image.

[0091] Since the decoded features obtained by upsampling the last level have strong discriminative power required for image enhancement, it is ensured that the quality of the second image is higher than that of the first image.

[0092] The image enhancement method provided in this application no longer uses a low-resolution feature search module. Instead, after performing multi-level downsampling on a first image of relatively low quality to obtain encoded features with multiple different sampling parameters, it performs multi-level upsampling on the encoded features obtained by downsampling the last level to obtain decoded features with multiple different sampling parameters. Among them, the upsampling of each non-first level includes: fusing the decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameter to obtain a fused feature, and then performing upsampling on the fused feature. The process of obtaining the fused feature in at least one non-first level includes: fusing the decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameter through pre-learned fusion parameters, which can avoid being interfered by too many low-quality features when obtaining features of low-quality images, so that the decoded features obtained by upsampling the last level have strong discriminative power features required for enhancement. Moreover, the process of fusing the encoded features and decoded features with the same sampling parameter does not require complex calculations, so the computational complexity is small, ensuring the processing speed of image enhancement and being easy to deploy.

[0093] In an optional embodiment, a flowchart of an implementation of fusing the decoded features obtained by upsampling the previous level through pre-learned fusion parameters and the encoded features with the same sampling parameter is as Figure 2 shown, and may include:

[0094] Step S201: Calculate a first weight of the encoded features with the same sampling parameter as the decoded features obtained by upsampling the previous level based on the fusion parameters, the decoded features obtained by upsampling the previous level, and the encoded features with the same sampling parameter.

[0095] The first weight can be calculated in the following way:

[0096] Calculate the global feature score g: g = sigmoid(DownSample_r(e + d) * W);

[0097] Calculate the gated feature score s: s = sigmoid(UpSample_r(g)).

[0098] Among them, DownSample_r represents the downsampling operation with a window size of r. Here, maxPooling, convolution, averagepooling, etc. can be used.

[0099] UpSample_r represents the upsampling operation with a window size of r. Here, maxUnPooling, transposed convolution, averageunpooling, etc. can be used.

[0100] Sigmoid() is a threshold function used to map the variable in the parentheses to the range [0, 1].

[0101] e represents the encoded feature, and d represents the decoded feature; e and d have the same sampling parameters, that is, d represents the decoded feature obtained by upsampling the previous level, and e is the encoded feature with the same sampling parameters as d.

[0102] W is a pre-learned fusion parameter. It is a matrix with a dimension of h / r * w / r, where h is the height of the encoded feature e, and w is the width of the encoded feature e; the height and width of the decoded feature d are the same as those of the encoded feature e.

[0103] The gated feature score s is the first weight of the encoded feature e.

[0104] Step S202: Perform weighted fusion on the decoded feature obtained by upsampling the previous level and the encoded feature with the same sampling parameters based on the first weight.

[0105] After obtaining the first weight s of the encoded feature e, use the first weight s to perform weighted fusion on the decoded feature d obtained by upsampling the previous level and the encoded feature e with the same sampling parameters.

[0106] In an optional embodiment, a flowchart of an implementation of the above-mentioned weighted fusion of the decoded feature obtained by upsampling the previous level and the encoded feature with the same sampling parameters is as Figure 3 shown and may include:

[0107] Step S301: Obtain the second weight of the decoded feature obtained by upsampling the previous level based on the first weight.

[0108] Optionally, the second weight can be the difference between 1 and the first weight, that is, the second weight is 1 - s.

[0109] Step S302: Perform weighted calculation on the decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameters based on the first weight and the second weight to obtain fused features.

[0110] In this embodiment, the fused feature f is calculated as follows: f = s * e+(1 - s) * d.

[0111] In an alternative embodiment, another implementation flowchart of the above weighted fusion of the decoded features obtained by upsampling the previous level based on the first weight and the encoded features with the same sampling parameters is as Figure 4 shown, and may include:

[0112] Step S401: Sum the decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameters to obtain initial fused features.

[0113] That is, the initial fused feature is: e + d.

[0114] Step S402: Obtain the second weight of the decoded features obtained by upsampling the previous level based on the first weight.

[0115] Optionally, the second weight can be the difference between 1 and the first weight, that is, the second weight is 1 - s.

[0116] In this application, step S401 can be executed first, and then step S402, or step S402 can be executed first, and then step S401, or the two steps can be executed simultaneously. This application does not specifically limit the execution order of the two steps.

[0117] Step S403: Perform weighted calculation on the initial fused features and the decoded features obtained by upsampling the previous level based on the first weight and the second weight to obtain fused features.

[0118] In this embodiment, the fused feature f is calculated as follows: f = s *(e + d)+(1 - s) * d.

[0119] In an alternative embodiment, the process of performing multiple levels of downsampling on the first image, performing multiple levels of upsampling on the encoded features obtained by the last level of downsampling, and obtaining the second image based on the decoded features obtained by the last level of upsampling can be implemented by a pre-trained enhancement network (which can also be referred to as an image enhancement model). Specifically:

[0120] Perform multiple levels of downsampling on the first image through the first encoding module of the enhancement network.

[0121] The encoded features obtained by downsampling the last layer are upsampled by multiple levels through the decoding module of the enhanced network. Among them, the upsampling of each non-first level includes: obtaining the fusion features obtained by the fusion module of the enhanced network by fusing the decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameter, and upsampling the fusion features; the process of obtaining the fusion features at at least one non-first level includes: the fusion module fuses the decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameter through pre-learned fusion parameters.

[0122] The decoded features obtained by upsampling the last layer are processed by the output module of the enhanced network to obtain a second image.

[0123] As Figure 5 shown, a schematic structural diagram of an enhanced network provided by an embodiment of the present application may include:

[0124] An encoding module (for easy distinction and description, denoted as the first encoding module) 501, a decoding module 502, a fusion module 503, and an output module 504;

[0125] Among them, the first encoding module 501 is used to perform multi-level downsampling on the first image to obtain encoded features with multiple different sampling parameters.

[0126] Figure 5 In the example shown, the first encoding module 501 performs 4-level downsampling on the first image. In practical applications, the first encoding module 501 may have more levels of downsampling, or may have 3 or 2 levels of downsampling. The specific number of downsampling levels is not specifically limited.

[0127] The decoding module 502 is used to perform multi-level upsampling on the encoded features obtained by downsampling the last layer (such as Figure 5 the encoded feature 4) to obtain decoded features with multiple different sampling parameters; among them, the upsampling of each non-first level includes: obtaining the fusion features obtained by the fusion module 503 by fusing the decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameter, and upsampling the fusion features; the process of obtaining the fusion features at at least one non-first level includes: the fusion module 503 fuses the decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameter through pre-learned fusion parameters.

[0128] Figure 5In the illustrated example, the decoding module 502 performs upsampling in four levels on the encoded features obtained by downsampling the last level. In practical applications, the decoding module 502 may have more levels of upsampling, or may have three or two levels of upsampling. The specific number of upsampling levels is not specifically limited, as long as the number of upsampling levels is the same as the number of downsampling levels.

[0129] In the fusion module, the fusion parameters in different fusion gating modules may be the same or different, and the specific values are determined according to the training of the enhancement network.

[0130] The output module 504 is used to obtain a second image based on the decoded features obtained by upsampling the last level (such as Figure 5 the decoded feature 4 in). The clarity of the second image is higher than that of the first image.

[0131] Figure 5 In the illustrated example, the fusion module fuses the encoded features and decoded features with the same sampling parameters in three levels based on the fusion parameters. In other examples, it is possible to fuse only the encoded features and decoded features with the same sampling parameters in two levels, or only fuse the encoded features and decoded features with the same sampling parameters in one level based on the fusion parameters.

[0132] Such as Figure 6 shown, is another structural schematic diagram of the enhancement network provided by the embodiment of the present application. The structures of the first encoding module, decoding module, and output module in this example are the same as those of the first encoding module, decoding module, and output module in the Figure 5 illustrated example, the difference being that the structure of the fusion module is different. Figure 6 In the illustrated example, the fusion module 601 only fuses the encoded features and decoded features with the same sampling parameters in two levels based on the fusion parameters (only fuses the encoded feature 1 and the decoded feature 3 based on the fusion parameters learned by the fusion gating module 3, and fuses the encoded feature 3 and the decoded feature 1 based on the fusion parameters learned by the fusion gating module 1), and the encoded features and decoded features with the same sampling parameters in another level are not fused based on the learned fusion parameters, but are simply added and fused (that is, the encoded feature 2 and the decoded feature 2 are directly fused in an additive manner).

[0133] Such as Figure 7 shown, is yet another structural schematic diagram of the enhancement network provided by the embodiment of the present application. The structures of the first encoding module, decoding module, and output module in this example are the same as those of the first encoding module, decoding module, and output module in the Figure 5 illustrated example, the difference being that the structure of the fusion module is different. Figure 7In the illustrated example, the fusion module 701 only fuses the encoded features and decoded features with the same sampling parameters at one level based on the fusion parameters (only fuses the encoded feature 1 and the decoded feature 3 based on the fusion parameters learned by the fusion gating module 3), while the encoded features and decoded features with the same sampling parameters at the other two levels are not fused based on the learned fusion parameters, but are simply added (that is, the encoded feature 2 and the decoded feature 2 are directly fused by addition, and the encoded feature 3 and the decoded feature 1 are directly fused by addition).

[0134] In an optional embodiment, when the fusion module fuses the decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameters using the pre-learned fusion parameters, it is used for:

[0135] Based on the fusion parameters, the decoded features obtained by upsampling the previous level, and the encoded features with the same sampling parameters, calculate the first weight of the encoded features with the same sampling parameters as the decoded features obtained by upsampling the previous level. The specific implementation manner can refer to the foregoing embodiments and will not be elaborated here.

[0136] Perform weighted fusion on the decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameters based on the first weight. The specific implementation manner can refer to the foregoing embodiments and will not be elaborated here.

[0137] In the fusion module, each fusion gating module fuses the input encoded features and decoded features with the same sampling parameters based on the learned fusion parameters. For example, the fusion gating module 1 fuses the encoded feature 3 and the decoded feature 1 based on the learned first fusion parameter; the fusion gating module 2 fuses the encoded feature 2 and the decoded feature 2 based on the learned second fusion parameter; the fusion gating module 3 fuses the encoded feature 1 and the decoded feature 3 based on the learned third fusion parameter.

[0138] Figure 5 - 7 In, the process of each fusion module fusing the encoded features and decoded features with the same sampling parameters based on the learned fusion parameters can refer to the foregoing embodiments and will not be elaborated here.

[0139] In an optional embodiment, the enhanced network can be trained in the following manner:

[0140] Perform unsupervised training on the initial network obtained by the second encoding module, the decoding module 502, and the output module 504 using the high-quality images in the first image set to obtain a pre-trained network.

[0141] Among them, the first image set includes multiple low-quality images (with lower clarity) and the corresponding high-quality images (with higher clarity) for each low-quality image. The second encoding module is used to perform multi-level downsampling on the input high-quality image to obtain encoded features with multiple different sampling parameters. Among the encoded features with multiple different sampling parameters obtained by the second encoding module, there is at least one encoded feature that has the same sampling parameter as at least one encoded feature among the encoded features with multiple different sampling parameters obtained by the first encoding module performing multi-level downsampling on the input high-quality image.

[0142] As Figure 8 shown, it is a schematic structural diagram of an initial network provided by an embodiment of the present application. The initial network includes: a second encoding module 801, a decoding module 802, and an output module 803; among them, the structures of the second encoding module 801 and the first encoding module 501 may be the same or different, as long as the second encoding module 801 can perform multi-level downsampling on the input image and can obtain an encoded feature that has the same sampling parameter as at least one encoded feature among the encoded features output by the first encoding module 501. The structure of the decoding module 802 is the same as that of the decoding module 502, and the structure of the output module 803 is the same as that of the output module 504.

[0143] As Figure 9 shown, it is another schematic structural diagram of the initial network provided by an embodiment of the present application. In this example, the structures of the second encoding module 801 and the first encoding module 501 are the same, and both include 4 downsampling modules.

[0144] One implementation manner of performing unsupervised training on the initial network can be:

[0145] Input the high-quality image into the second encoding module 801.

[0146] The second encoding module 801 performs multi-level downsampling on the high-quality image to obtain encoded features with multiple different sampling parameters.

[0147] The decoding module 502 performs multi-level upsampling on the encoded feature obtained by the last-level downsampling of the second encoding module 801 to obtain decoded features with multiple different sampling parameters.

[0148] The output module 504 processes the decoded feature obtained by the last-level upsampling of the decoding module 502 to obtain a reconstructed high-quality image.

[0149] Update the parameters of the initial network based on the difference between the reconstructed high-quality image and the original high-quality image (i.e., the high-quality image input into the second encoding module 801).

[0150] Specifically, the parameters of the initial network can be updated with the goal that the difference between the reconstructed high-quality image and the original high-quality image becomes smaller and smaller (that is, compared with reconstructing the high-quality image using the initial network before parameter update, the reconstructed image obtained by reconstructing the high-quality image using the initial network after parameter update is closer to the original high-quality image).

[0151] The purpose of unsupervised training of the initial network using high-quality images is to make full use of the features of high-quality images to obtain a better decoding module, making the enhanced image more realistic.

[0152] After obtaining the pre-trained network, based on the pre-trained network, the first encoding module (the first encoding module here can be the second encoding module in the pre-trained network or not) and the fusion module are supervised-trained using the first image set to obtain the trained first encoding module and fusion module; the trained first encoding module and fusion module, together with the decoding module and output module in the pre-trained network, constitute the enhanced network, that is, the image enhancement model.

[0153] Optionally, if the structure of the second encoding module is the same as that of the first encoding module, then based on the pre-trained network, the fusion module and the second encoding module in the pre-trained network can be supervised-trained using the first image set to obtain the trained fusion module and first encoding module (that is, using the trained second encoding module as the first encoding module).

[0154] If the structure of the second encoding module is different from that of the first encoding module, then based on the pre-trained network, the first encoding module and the fusion module can be supervised-trained using the first image set to obtain the trained first encoding module and fusion module.

[0155] In an optional embodiment, a flowchart of an implementation of the above-mentioned supervised training of the first encoding module based on the pre-trained network using the first image set is as Figure 10 shown and may include:

[0156] Step S1001: For any low-quality image in the first image set, input the any low-quality image into the first encoding module to obtain the encoded features of the any low-quality image output by the first encoding module with multiple different sampling parameters.

[0157] The first encoding module here can be the second encoding module in the pre-trained network or a different encoding module from the second encoding module in the pre-trained network.

[0158] Step S1002: Input the high-quality image corresponding to any one of the low-quality images into the second encoding module in the pre-trained network, and obtain the encoded features of the high-quality image corresponding to any one of the low-quality images output by the second encoding module with multiple different sampling parameters.

[0159] When the first encoding module is the second encoding module in the pre-trained network, the sampling parameters of the encoded features output by the second encoding module and the first encoding module in this step are the same.

[0160] When the first encoding module is not the second encoding module in the pre-trained network, there are several possibilities as follows:

[0161] The sampling parameters of the encoded features output by the second encoding module and the first encoding module are the same.

[0162] Among the encoded features output by the second encoding module, some of the encoded features have the same sampling parameters as at least some of the encoded features output by the first encoding module.

[0163] The encoded features output by the second encoding module have the same sampling parameters as some of the encoded features output by the first encoding module.

[0164] Step S1003: Aim at making the first difference between the encoded features with the same sampling parameters of any one of the low-quality images and its corresponding high-quality image smaller and smaller, and update the parameters of the first encoding module.

[0165] That is to say, compared with encoding any one of the low-quality images using the first encoding module before parameter update, for the encoded features with multiple different sampling parameters obtained by performing multiple levels of downsampling on any one of the low-quality images using the first encoding module after parameter update, and the encoded features obtained by performing multiple levels of downsampling on the high-quality image corresponding to any one of the low-quality images using the second encoding module, the first difference between the encoded features with the same sampling parameters is smaller.

[0166] In this embodiment, the second encoding module in the pre-trained network is used as the teacher network, and the first encoding module is used as the student network to perform distillation training on the first encoding module, so that the features obtained by the student network are closer to the features obtained by the teacher network.

[0167] In an optional embodiment, for the first encoded feature with any sampling parameter of any one of the low-quality images and the second encoded feature with the same sampling parameter of the high-quality image corresponding to any one of the low-quality images (that is, the second encoded feature and the first encoded feature are encoded features with the same sampling parameter of different images), the first difference between the first encoded feature and the second encoded feature is calculated as follows:

[0168] Obtain the absolute error between the first encoded feature and the second encoded feature (denoted as Loss1), the sum of squared errors between the first classification result based on the first encoded feature and the second classification result based on the second encoded feature (denoted as Loss2), and the error at the maximum output of the first encoded feature and the second encoded feature (denoted as Loss3).

[0169] Optionally, the above three errors can be calculated as follows:

[0170] Loss1 = |f(teacher) – f(student)|1

[0171] Loss2 = |vgg(f(teacher)) - vgg(f(student))|2

[0172] Loss3 = |softmax(f(teacher) × f(teacher).T) - softmax(f(student) × f(student).T)|2

[0173] Where f(teacher) represents the second encoded feature, and f(student) represents the first encoded feature. vgg represents the pre-trained vgg network for image classification. vgg(f(teacher)) represents the second classification result obtained by the vgg network based on the second encoded feature, and vgg(f(student)) represents the first classification result obtained by the vgg network based on the first encoded feature. f(teacher).T represents the transpose matrix of the second encoded feature, and f(student).T represents the transpose matrix of the first encoded feature.

[0174] Weightedly sum the above absolute error, sum of squared errors, and error at the maximum output to obtain the first difference between the first encoded feature and the second encoded feature (denoted as Loss).

[0175] The first difference Loss can be calculated as follows:

[0176] Loss = λ1 × Loss1 + λ2 × Loss2 + λ3 × Loss3

[0177] Where λ1, λ2, and λ3 are the weights of the three errors, and these 3 weights are hyperparameters, that is, they are preset in advance.

[0178] Optionally, the weight of the absolute error is negatively correlated with the downsampling level for obtaining the first encoded feature. That is, the lower the downsampling level, the greater the weight of the corresponding absolute error; correspondingly, the larger the size (or resolution) of the first encoded feature, the greater the weight of the corresponding absolute error.

[0179] Optionally, the weight of the sum of squared errors is positively correlated with the downsampling level of the first encoded feature. That is, the higher the downsampling level, the greater the weight of the corresponding sum of squared errors; correspondingly, the smaller the size (or resolution) of the first encoded feature, the greater the weight of the corresponding sum of squared errors.

[0180] In an optional embodiment, after the first encoded feature is trained, a flowchart of an implementation of supervised training of the fusion module using a first image set based on a pre-trained network is as Figure 11 shown and may include:

[0181] Step S1101: Downsample any low-quality image through the trained first encoding module at multiple levels to obtain encoded features with multiple different sampling parameters.

[0182] The trained first encoding module refers to the first encoding module obtained through the aforementioned distillation training.

[0183] Step S1102: Upsample the encoded features obtained by downsampling at the last level through the trained decoding module at multiple levels to obtain decoded features with multiple sampling parameters. Among them, each non-first-level upsampling includes: obtaining the fused feature obtained by the fusion module by fusing the decoded feature obtained by upsampling the previous level and the encoded feature with the same sampling parameter, and upsampling the fused feature; the process of obtaining the fused feature at at least one non-first level includes: the fusion module fuses the decoded feature obtained by upsampling the previous level and the encoded feature with the same sampling parameter through the fusion parameter.

[0184] The process of the fusion module obtaining the fused feature can refer to the foregoing embodiments and will not be elaborated here.

[0185] The trained decoding module refers to the decoding module in the pre-trained network obtained by training the initial network as described above.

[0186] Step S1103: Obtain the enhanced image corresponding to any of the above low-quality images through the trained output module based on the decoded features obtained by upsampling at the last level.

[0187] The trained output module refers to the output module in the pre-trained network obtained by training the initial network as described above.

[0188] Step S1104: Take the second difference between the enhanced image corresponding to any of the above low-quality images and the high-quality image corresponding to the same low-quality image becoming smaller and smaller as the goal, and update the parameters of the fusion module.

[0189] The parameters of the fusion module include the above fusion parameters.

[0190] The fact that the second difference between the enhanced image corresponding to any one of the above low-quality images and the high-quality image corresponding to any one of the above low-quality images becomes smaller and smaller means that: compared with performing image enhancement on any one of the low-quality images by the fusion module before parameter update, the enhanced image obtained by performing image enhancement on any one of the low-quality images by the fusion module after parameter update is closer to the high-quality image corresponding to any one of the low-quality images.

[0191] When training the fusion module, the parameters of the first encoding module, the decoding module, and the output module are all frozen and unchanged.

[0192] Corresponding to the method embodiment, the present application further provides an image enhancement device. A schematic structural diagram of an image enhancement device provided by the present application is as Figure 12 shown, and may include:

[0193] an encoding unit 1201, a decoding unit 1202, a fusion unit 1203, and an output unit 1204;

[0194] The encoding unit 1201 is configured to perform downsampling on the first image at multiple levels to obtain encoded features with multiple different sampling parameters;

[0195] The decoding unit 1202 is configured to perform upsampling on the encoded features obtained by the downsampling at the last level at multiple levels to obtain decoded features with multiple different sampling parameters; wherein, the upsampling of each non-first level includes: obtaining a fused feature by the fusion unit 1203 fusing the decoded feature obtained by the upsampling of the previous level and the encoded feature with the same sampling parameter, and performing upsampling on the fused feature; the process of obtaining the fused feature in at least one non-first level includes: fusing the decoded feature obtained by the upsampling of the previous level and the encoded feature with the same sampling parameter through pre-learned fusion parameters;

[0196] The output unit 1204 is configured to obtain a second image based on the decoded features obtained by the upsampling at the last level; the clarity of the second image is higher than that of the first image.

[0197] The image enhancement device provided by the embodiment of the present application performs downsampling on a first image with a lower definition at multiple levels to obtain encoded features with multiple different sampling parameters; performs upsampling on the encoded features obtained by the downsampling of the last level at multiple levels to obtain decoded features with multiple different sampling parameters; wherein, the upsampling of each non-first level includes: fusing the decoded features obtained by the upsampling of the previous level and the encoded features with the same sampling parameter to obtain a fused feature, and performing upsampling on the fused feature; the process of obtaining the fused feature in at least one non-first level includes: fusing the decoded features obtained by the upsampling of the previous level and the encoded features with the same sampling parameter through pre-learned fusion parameters; obtaining a second image with a higher definition based on the decoded features obtained by the upsampling of the last level. By fusing the decoded features and the encoded features with the same sampling parameter through pre-learned fusion parameters and then performing upsampling, it is possible to obtain the features of a sufficiently low-quality image without being disturbed by too many low-quality features in the generation of a high-quality image, so that the decoded features obtained by the upsampling of the last level have strong discriminative features required for enhancement. Moreover, the feature fusion process does not require complex calculations. Therefore, the computational complexity of the image enhancement process is small, ensuring the processing speed of image enhancement and facilitating deployment.

[0198] In an alternative embodiment, when the fusion unit 1203 fuses the decoded features obtained by the upsampling of the previous level and the encoded features with the same sampling parameter through pre-learned fusion parameters, it is configured to:

[0199] Based on the fusion parameters, the decoded features obtained by the upsampling of the previous level, and the encoded features with the same sampling parameter, calculate a first weight of the encoded features with the same sampling parameter as the decoded features obtained by the upsampling of the previous level;

[0200] Perform weighted fusion on the decoded features obtained by the upsampling of the previous level and the encoded features with the same sampling parameter based on the first weight.

[0201] In an alternative embodiment, when the fusion unit 1203 performs weighted fusion on the decoded features obtained by the upsampling of the previous level and the encoded features with the same sampling parameter based on the first weight, it is configured to:

[0202] Obtain a second weight of the decoded features obtained by the upsampling of the previous level based on the first weight;

[0203] Perform weighted calculation on the decoded features obtained by the upsampling of the previous level and the encoded features with the same sampling parameter based on the first weight and the second weight.

[0204] In an optional embodiment, when the fusion unit 1203 performs weighted fusion on the decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameter based on the first weight, it is configured to:

[0205] Sum the decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameter to obtain initial fusion features;

[0206] Obtain a second weight of the decoded features obtained by upsampling the previous level based on the first weight;

[0207] Perform weighted calculation on the initial fusion features and the decoded features obtained by upsampling the previous level based on the first weight and the second weight.

[0208] In an optional embodiment, the image enhancement device performs multi-level downsampling on a first image through an enhancement network, performs multi-level upsampling on the encoded features obtained by downsampling the last level, and obtains a second image based on the decoded features obtained by upsampling the last level, specifically including:

[0209] Perform multi-level downsampling on the first image through a first encoding module of the enhancement network;

[0210] Perform multi-level upsampling on the encoded features obtained by downsampling the last level through a decoding module of the enhancement network, where each non-first-level upsampling includes: obtaining a fusion feature obtained by the fusion module of the enhancement network by fusing the decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameter, and performing upsampling on the fusion feature; the process of obtaining the fusion feature in at least one non-first level includes: the fusion module fusing the decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameter through pre-learned fusion parameters;

[0211] Process the decoded features obtained by upsampling the last level through an output module of the enhancement network to obtain the second image.

[0212] In an optional embodiment, the enhancement network is trained in the following manner:

[0213] Perform unsupervised training on an initial network obtained by a second encoding module, the decoding module, and the output module with high-quality images in a first image set to obtain a pre-trained network; the first image set includes multiple low-quality images and the high-quality images corresponding to each low-quality image; the second encoding module is configured to perform multi-level downsampling on the input high-quality images to obtain encoded features with multiple different sampling parameters;

[0214] Based on the pre-trained network, the first encoding module and the fusion module are supervised-trained by the first image set to obtain the trained first encoding module and fusion module; the trained first encoding module and fusion module, together with the decoding module and the output module in the pre-trained network, constitute the enhanced network.

[0215] The specific implementation can refer to the foregoing embodiments and will not be elaborated here.

[0216] Corresponding to the method embodiment, the present application also provides a model training device for training the foregoing image enhancement model. A schematic structural diagram of the model training device provided by the present application is as Figure 13 shown, and may include: a first training unit 1301 and a second training unit 1302;

[0217] Among them, the first training unit 1301 is used to unsupervised-train the initial network obtained by the second encoding module, the decoding module and the output module through the high-quality images in the first image set to obtain the pre-trained network; the first image set includes multiple low-quality images and the high-quality images corresponding to each low-quality image; the second encoding module is used to perform multi-level downsampling on the input high-quality images to obtain encoding features with multiple different sampling parameters; the structure of the second encoding module is the same as or different from the structure of the first encoding module;

[0218] The second training unit 1302 is used to supervised-train the first encoding module and the fusion module by the first image set based on the pre-trained network to obtain the trained first encoding module and fusion module; the trained first encoding module and fusion module, together with the decoding module and the output module in the pre-trained network, constitute the enhanced network.

[0219] In an optional embodiment, when the second training unit 1302 supervised-trains the first encoding module by the first image set based on the pre-trained network, it is used for:

[0220] For any low-quality image in the first image set, input the any low-quality image into the first encoding module to obtain encoding features of the any low-quality image with multiple different sampling parameters;

[0221] Input the high-quality image corresponding to the any low-quality image into the second encoding module in the pre-trained network to obtain encoding features of the high-quality image corresponding to the any low-quality image with multiple different sampling parameters;

[0222] Taking the goal that the first difference between the encoding features with the same sampling parameters of any of the low-quality images and their corresponding high-quality images becomes smaller and smaller, update the parameters of the first encoding module.

[0223] The specific implementation manner can refer to the foregoing embodiments and will not be elaborated here.

[0224] Corresponding to the method embodiment, the present application also provides an electronic device. A schematic structural diagram of the electronic device is as Figure 14 shown, and may include: at least one processor 1, at least one communication interface 2, at least one memory 3, and at least one communication bus 4.

[0225] In the embodiments of the present application, the number of the processor 1, the communication interface 2, the memory 3, and the communication bus 4 is at least one, and the processor 1, the communication interface 2, and the memory 3 complete mutual communication through the communication bus 4.

[0226] The processor 1 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application, etc.

[0227] The memory 3 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.

[0228] Among them, the memory 3 stores a program, and the processor 1 can call the program stored in the memory 3;

[0229] The program is used for: performing downsampling on the first image at multiple levels to obtain encoding features with multiple different sampling parameters; performing upsampling on the encoding features obtained by downsampling at the last level at multiple levels to obtain decoding features with multiple different sampling parameters; wherein, each upsampling of a non-first level includes: fusing the decoding features obtained by upsampling at the previous level and the encoding features with the same sampling parameter to obtain a fused feature, and performing upsampling on the fused feature; the process of obtaining the fused feature at at least one non-first level includes: fusing the decoding features obtained by upsampling at the previous level and the encoding features with the same sampling parameter through pre-learned fusion parameters; obtaining a second image based on the decoding features obtained by upsampling at the last level; the clarity of the second image is higher than that of the first image.

[0230] Alternatively, the program is used to: perform unsupervised training on the initial network obtained by the second encoding module, the decoding module, and the output module with high-quality images in the first image set to obtain a pre-trained network; the first image set includes multiple low-quality images and the high-quality images corresponding to each low-quality image; the second encoding module is used to perform downsampling at multiple levels on the input high-quality images to obtain encoded features with multiple different sampling parameters; based on the pre-trained network, perform supervised training on the first encoding module and the fusion module with the first image set to obtain the trained first encoding module and fusion module; the trained first encoding module and fusion module, and the decoding module and output module in the pre-trained network constitute the image enhancement model.

[0231] Optionally, the refinement function and the extension function of the program can refer to the above description.

[0232] The embodiment of the present application also provides a storage medium, which can store a program suitable for execution by a processor;

[0233] The program is used to: perform downsampling at multiple levels on the first image to obtain encoded features with multiple different sampling parameters; perform upsampling at multiple levels on the encoded features obtained by the downsampling of the last level to obtain decoded features with multiple different sampling parameters; wherein, the upsampling of each non-first level includes: fusing the decoded features obtained by the upsampling of the previous level and the encoded features with the same sampling parameter to obtain a fused feature, and performing upsampling on the fused feature; the process of obtaining the fused feature in at least one non-first level includes: fusing the decoded features obtained by the upsampling of the previous level and the encoded features with the same sampling parameter through pre-learned fusion parameters; obtaining a second image based on the decoded features obtained by the upsampling of the last level; the clarity of the second image is higher than that of the first image.

[0234] Alternatively, the program is used to: perform unsupervised training on the initial network obtained by the second encoding module, the decoding module, and the output module with high-quality images in the first image set to obtain a pre-trained network; the first image set includes multiple low-quality images and the high-quality images corresponding to each low-quality image; the second encoding module is used to perform downsampling at multiple levels on the input high-quality images to obtain encoded features with multiple different sampling parameters; based on the pre-trained network, perform supervised training on the first encoding module and the fusion module with the first image set to obtain the trained first encoding module and fusion module; the trained first encoding module and fusion module, and the decoding module and output module in the pre-trained network constitute the image enhancement model.

[0235] Optionally, the refinement function and the extension function of the program may refer to the above description.

[0236] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0237] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0238] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0239] In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0240] It should be understood that in the embodiments of this application, the dependent claims, each embodiment, and features can be combined with each other to solve the foregoing technical problems.

[0241] If the described function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or this part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0242] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An image enhancement method, the method comprising: Performing downsampling on a first image at multiple levels to obtain encoded features with multiple different sampling parameters; Performing upsampling on the encoded features obtained by downsampling at the last level at multiple levels to obtain decoded features with multiple different sampling parameters; wherein, the upsampling of each non-first level includes: calculating a first weight of the encoded features with the same sampling parameter as the decoded features obtained by upsampling at the previous level based on pre-learned fusion parameters, the decoded features obtained by upsampling at the previous level, and the encoded features with the same sampling parameter; performing weighted fusion on the decoded features obtained by upsampling at the previous level and the encoded features with the same sampling parameter based on the first weight to obtain a fused feature; performing upsampling on the fused feature; Obtaining a second image based on the decoded features obtained by upsampling at the last level; the sharpness of the second image is higher than that of the first image.

2. The method according to claim 1, wherein the performing weighted fusion on the decoded features obtained by upsampling at the previous level and the encoded features with the same sampling parameter based on the first weight includes: Obtaining a second weight of the decoded features obtained by upsampling at the previous level based on the first weight; Performing weighted calculation on the decoded features obtained by upsampling at the previous level and the encoded features with the same sampling parameter based on the first weight and the second weight.

3. The method according to claim 1, wherein the performing weighted fusion on the decoded features obtained by upsampling at the previous level and the encoded features with the same sampling parameter based on the first weight includes: Summing the decoded features obtained by upsampling at the previous level and the encoded features with the same sampling parameter to obtain an initial fused feature; Obtaining a second weight of the decoded features obtained by upsampling at the previous level based on the first weight; Performing weighted calculation on the initial fused feature and the decoded features obtained by upsampling at the previous level based on the first weight and the second weight.

4. The method according to claim 1, the process of performing downsampling on a first image at multiple levels, performing upsampling on the encoded features obtained by downsampling at the last level at multiple levels, and obtaining a second image based on the decoded features obtained by upsampling at the last level includes: Performing downsampling on the first image at multiple levels through a first encoding module of an enhancement network; Performing upsampling on the encoded features obtained by downsampling at the last level at multiple levels through a decoding module of the enhancement network, wherein the upsampling of each non-first level includes: obtaining a fused feature obtained by fusing the decoded features obtained by upsampling at the previous level and the encoded features with the same sampling parameter by a fusion module of the enhancement network, and performing upsampling on the fused feature; the process of obtaining a fused feature in at least one non-first level includes: the fusion module fusing the decoded features obtained by upsampling at the previous level and the encoded features with the same sampling parameter through pre-learned fusion parameters; The decoded features obtained by upsampling the last level through the output module of the enhancement network are processed to obtain the second image.

5. The method according to claim 4, wherein the enhancement network is trained as follows: The initial network obtained by the second encoding module, the decoding module, and the output module is unsupervised trained by high-quality images in the first image set to obtain a pre-trained network; the first image set includes multiple low-quality images and the high-quality images corresponding to each low-quality image; the second encoding module is configured to perform downsampling on the input high-quality image at multiple levels to obtain encoded features with multiple different sampling parameters; Based on the pre-trained network, the first encoding module and the fusion module are supervised trained by the first image set to obtain the trained first encoding module and fusion module; the trained first encoding module and fusion module, and the decoding module and output module in the pre-trained network constitute the enhancement network.

6. An image enhancement model, comprising: The first encoding module, the decoding module, the fusion module, and the output module; The first encoding module is configured to perform downsampling on the first image at multiple levels to obtain encoded features with multiple different sampling parameters; The decoding module is configured to perform upsampling on the encoded features obtained by downsampling the last level at multiple levels to obtain decoded features with multiple different sampling parameters; wherein, the upsampling of each non-first level includes: obtaining, by the fusion module, based on pre-learned fusion parameters, the decoded features obtained by upsampling the previous level, and the encoded features with the same sampling parameter, a first weight of the encoded features with the same sampling parameter as the decoded features obtained by upsampling the previous level; performing weighted fusion on the decoded features obtained by upsampling the previous level and the encoded features with the same sampling parameter based on the first weight to obtain a fused feature; performing upsampling on the fused feature; the output module is configured to obtain a second image based on the decoded features obtained by upsampling the last level; the clarity of the second image is higher than that of the first image.

7. A model training method for training the image enhancement model according to claim 6, the model training method comprising: The initial network obtained by the second encoding module, the decoding module, and the output module is unsupervised trained by high-quality images in the first image set to obtain a pre-trained network; the first image set includes multiple low-quality images and the high-quality images corresponding to each low-quality image; the second encoding module is configured to perform downsampling on the input high-quality image at multiple levels to obtain encoded features with multiple different sampling parameters; Based on the pre-trained network, the first encoding module and the fusion module are supervised trained by the first image set to obtain the trained first encoding module and fusion module; the trained first encoding module and fusion module, and the decoding module and output module in the pre-trained network constitute the image enhancement model.

8. The method according to claim 7, based on the pre-trained network, performing supervised training on the first encoding module through the first image set, includes: For any low-quality image in the first image set, inputting the any low-quality image into the first encoding module to obtain encoding features of the any low-quality image with multiple different sampling parameters; Inputting the high-quality image corresponding to the any low-quality image into the second encoding module in the pre-trained network to obtain encoding features of the high-quality image corresponding to the any low-quality image with multiple different sampling parameters; Taking the first difference between the encoding features with the same sampling parameter of the any low-quality image and its corresponding high-quality image as smaller and smaller as the goal, and updating the parameters of the first encoding module.

9. An image enhancement device, the device includes: An encoding unit, configured to perform downsampling on a first image at multiple levels to obtain encoding features with multiple different sampling parameters; A decoding unit, configured to perform upsampling on the encoding features obtained by the downsampling of the last level at multiple levels to obtain decoding features with multiple different sampling parameters; wherein, the upsampling of each non-first level includes: obtaining a first weight of the encoding features with the same sampling parameter as the decoding features obtained by the upsampling of the previous level by the fusion unit based on pre-learned fusion parameters, the decoding features obtained by the upsampling of the previous level, and the encoding features with the same sampling parameter; performing weighted fusion on the decoding features obtained by the upsampling of the previous level and the encoding features with the same sampling parameter based on the first weight to obtain a fusion feature; performing upsampling on the fusion feature; An output unit, configured to obtain a second image based on the decoding features obtained by the upsampling of the last level; the clarity of the second image is higher than that of the first image.

10. An electronic device, includes: A memory, configured to store a program; A processor, configured to call and execute the program in the memory, and implement each step of the image enhancement method as described in any one of claims 1-5 and / or implement the model training method as described in claim 7 or 8 by executing the program.