Image super-resolution model training method, device, equipment and storage medium

By processing and sampling RGB domain images, RAW domain image pairs are generated, and image super-score models are trained, which solves the problem that the prior art cannot super-score RAW domain images, and realizes efficient super-score of RAW domain images.

CN117788287BActive Publication Date: 2025-05-23AXERA SEMICON (SHANGHAI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311533683.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-16
Publication Date
2025-05-23
Estimated Expiration
2043-11-16

Smart Images

  • Figure CN117788287B_ABST
    Figure CN117788287B_ABST
Patent Text Reader

Abstract

The present application provides a training method, device, equipment and storage medium for an image super-resolution model, the method comprising: obtaining at least one first RGB domain image; processing each first RGB domain image to obtain a second RGB domain image corresponding to each first RGB domain image, each first RGB domain image and the corresponding second RGB domain image forming an initial image pair; sampling the first RGB domain image in each initial image pair to obtain a first RAW domain image corresponding to each first RGB domain image, and sampling the second RGB domain image in each initial image pair to obtain a second RAW domain image corresponding to each second RGB domain image, each first RAW domain image and the corresponding second RAW domain forming a RAW domain image pair. The technical solution of the present application can super-resolution the RAW domain image, thereby meeting the user's rich super-resolution image requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a training method, apparatus, device and storage medium for an image super-resolution model. Background Art

[0002] At present, many image generation works are based on the diffusion model. The diffusion model is divided into a forward denoising process and a reverse denoising process.

[0003] In the related art, super-resolution processing is performed on RGB domain images based on a diffusion model to obtain high-quality RGB domain images.

[0004] However, the above diffusion model only performs super-resolution processing on RGB domain images, so that users can only obtain RGB domain super-resolution images, which cannot meet the users' rich image needs. Summary of the invention

[0005] The embodiment of the present application provides a training method, apparatus, device and storage medium for an image super-resolution model, which can super-resolution RAW domain images, thereby meeting the user's rich super-resolution image requirements. The technical solution is as follows:

[0006] According to one aspect of an embodiment of the present application, a method for training an image super-resolution model is provided, the method comprising:

[0007] Acquire at least one first RGB domain image;

[0008] Processing each of the first RGB domain images to obtain a second RGB domain image corresponding to each of the first RGB domain images, each of the first RGB domain images and the corresponding second RGB domain image forming an initial image pair, and a resolution of the second RGB domain image in each of the initial image pairs is lower than a resolution of the first RGB domain image;

[0009] Sampling the first RGB domain image in each of the initial image pairs to obtain a first RAW domain image corresponding to each of the first RGB domain images, and sampling the second RGB domain image in each of the initial image pairs to obtain a second RAW domain image corresponding to each of the second RGB domain images, wherein each of the first RAW domain images and the corresponding second RAW domain image form a RAW domain image pair;

[0010] Based on multiple RAW domain image pairs, the image super-resolution model is trained.

[0011] In a possible implementation manner, processing each of the first RGB domain images includes:

[0012] For each of the first RGB domain images, obtain the value of each first pixel point of the first RGB domain image;

[0013] The value of each of the first pixel points is calculated based on a preset Gaussian kernel to obtain values ​​of a plurality of second pixel points in the second RGB domain image corresponding to the first RGB domain image.

[0014] In a possible implementation, sampling the first RGB domain image in each of the initial image pairs to obtain a first RAW domain image corresponding to each of the first RGB domain images, and sampling the second RGB domain image in each of the initial image pairs to obtain a second RAW domain image corresponding to each of the second RGB domain images, includes:

[0015] For each of the initial image pairs, a target value of each first pixel point of the first RGB domain image in the initial image pair is sampled to obtain the first RAW domain image in the RAW domain image pair corresponding to each of the initial image pairs, and a target value of each second pixel point of the second RGB domain image in the initial image pair is sampled to obtain the second RAW domain image in the RAW domain image pair corresponding to each of the initial image pairs;

[0016] Among them, the target value of the first pixel point is any one of the first color value, the second color value and the third color value corresponding to the first pixel point, and the target value of the second pixel point is any one of the fourth color value, the fifth color value and the sixth color value corresponding to the second pixel point.

[0017] In a possible implementation, the image super-resolution model is trained based on a plurality of RAW domain image pairs, including:

[0018] The image super-resolution model is iteratively trained based on each RAW domain image pair.

[0019] In a possible implementation, performing an iterative training on the image super-resolution model based on each RAW domain image pair includes:

[0020] In the i-th iteration process, an i-th initial noise image is obtained; the initial noise image is randomly generated;

[0021] performing noise processing on the second RAW domain image in the i-th RAW domain image pair based on the i-th initial noise image to obtain an i-th RAW domain noise image; i is an integer greater than 1 or equal to 1;

[0022] splicing the first RAW domain image and the i-th RAW domain noise image in the i-th RAW domain image pair to obtain an i-th spliced ​​image;

[0023] Based on the image super-resolution model, noise prediction is performed on the i-th spliced ​​image to obtain the i-th predicted noise image;

[0024] Calculate the i-th predicted noise image and the i-th initial noise image to obtain the i-th loss function value of the image super-resolution model;

[0025] The parameters of the image super-resolution model are updated based on the i-th loss function value.

[0026] In a possible implementation, the performing noise processing on the second RAW domain image in the i-th RAW domain image pair based on the i-th initial noise image to obtain the i-th RAW domain noise image includes:

[0027] Get the i-th noise parameter;

[0028] Based on the i-th noise parameter, obtaining a first weight of the i-th initial noise image and a second weight of the second RAW domain image in the i-th RAW domain image pair;

[0029] Calculating the i-th initial noise image and the first weight to obtain a first calculation result, and calculating the second RAW domain image in the i-th RAW domain image pair and the second weight to obtain a second calculation result;

[0030] The first calculation result and the second calculation result are summed to obtain the i-th RAW domain noise image.

[0031] In a possible implementation, the method further includes:

[0032] Obtain the RAW domain image to be super-resolved;

[0033] The RAW domain image is super-resolved based on the trained image super-resolution model to obtain the target RAW domain image.

[0034] According to a second aspect of an embodiment of the present application, a training device for an image super-resolution model is provided, the device comprising:

[0035] A first acquisition module, used to acquire at least one first RGB domain image;

[0036] a processing module, configured to process each of the first RGB domain images to obtain a second RGB domain image corresponding to each of the first RGB domain images, wherein each of the first RGB domain images and the corresponding second RGB domain image form an initial image pair, and a resolution of the second RGB domain image in each of the initial image pairs is lower than a resolution of the first RGB domain image;

[0037] a sampling module, configured to sample the first RGB domain image in each of the initial image pairs to obtain a first RAW domain image corresponding to each of the first RGB domain images, and to sample the second RGB domain image in each of the initial image pairs to obtain a second RAW domain image corresponding to each of the second RGB domain images, wherein each of the first RAW domain images and the corresponding second RAW domain image form a RAW domain image pair;

[0038] The training module is used to train the image super-resolution model based on a plurality of RAW domain image pairs.

[0039] In a possible implementation, the processing module includes:

[0040] a sampling unit, configured to sample, for each of the initial image pairs, a target value of each first pixel point of the first RGB domain image in the initial image pair to obtain the first RAW domain image in the RAW domain image pair corresponding to each of the initial image pairs, and to sample the target value of each second pixel point of the second RGB domain image in the initial image pair to obtain the second RAW domain image in the RAW domain image pair corresponding to each of the initial image pairs;

[0041] Among them, the target value of the first pixel point is any one of the first color value, the second color value and the third color value corresponding to the first pixel point, and the target value of the second pixel point is any one of the fourth color value, the fifth color value and the sixth color value corresponding to the second pixel point.

[0042] In a possible implementation, the training module includes:

[0043] The iteration unit is used to perform an iterative training on the image super-resolution model based on each RAW domain image pair.

[0044] In a possible implementation, the iteration unit includes:

[0045] An acquisition subunit is used to acquire an i-th initial noise image in an i-th iteration process; the initial noise image is randomly generated;

[0046] a processing subunit, configured to perform noise processing on the second RAW domain image in the i-th RAW domain image pair based on the i-th initial noise image to obtain an i-th RAW domain noise image; i is an integer greater than 1 or equal to 1;

[0047] a stitching subunit, configured to stitch the first RAW domain image and the i-th RAW domain noise image in the i-th RAW domain image pair to obtain an i-th stitched image;

[0048] A prediction subunit, used for performing noise prediction on the i-th spliced ​​image based on the image super-resolution model to obtain an i-th predicted noise image;

[0049] A calculation subunit, used to calculate the i-th predicted noise image and the i-th initial noise image to obtain the i-th loss function value of the image super-resolution model;

[0050] The updating subunit is used to update the parameters of the image super-resolution model based on the i-th loss function value.

[0051] In a possible implementation, the processing subunit includes:

[0052] A first acquisition subunit is used to acquire the i-th noise parameter;

[0053] a second acquisition subunit, configured to acquire, based on the i-th noise parameter, a first weight of the i-th initial noise image and a second weight of the second RAW domain image in the i-th RAW domain image pair;

[0054] a calculation subunit, configured to calculate the i-th initial noise image and the first weight to obtain a first calculation result, and to calculate the second RAW domain image in the i-th RAW domain image pair and the second weight to obtain a second calculation result;

[0055] The summing subunit is used to sum the first calculation result and the second calculation result to obtain the i-th RAW domain noise image.

[0056] In a possible implementation, the device further includes:

[0057] The second acquisition module is used to acquire the RAW domain image to be super-resolved;

[0058] The super-resolution module is used to super-resolution the RAW domain image based on a trained image super-resolution model to obtain a target RAW domain image.

[0059] According to a second aspect of an embodiment of the present application, a computer device is provided, comprising a processor and a memory, wherein the memory is used to store at least one program, and the at least one program is loaded by the processor and executes the training method of the image super-resolution model as described above.

[0060] According to a second aspect of an embodiment of the present application, a computer-readable storage medium is provided, in which at least one program is stored. The at least one program is loaded and executed by a processor to implement the training method of the image super-resolution model as described above.

[0061] In an embodiment of the present application, an embodiment of the present application provides a training method for an image super-resolution model, by processing each high-resolution first RGB domain image, a low-resolution second RGB domain image corresponding to each first RGB domain image is obtained, the corresponding first RGB domain image and the second RGB domain image constitute an RGB domain image pair, by sampling each RGB domain image pair, a RAW domain image pair corresponding to each RGB domain image pair is obtained, and an image super-resolution model is trained based on multiple RAW domain image pairs to obtain an image super-resolution model that super-resolutions the RAW domain images, thereby meeting the user's rich image needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0063] Figure 1 is a schematic diagram of an implementation environment provided according to an embodiment of the present application;

[0064] Figure 2 It is a flowchart of a training method for an image super-resolution model provided according to an embodiment of the present application;

[0065] Figure 3 This is a schematic diagram of an example flow chart of a training method for an image super-resolution model provided according to an embodiment of the present application;

[0066] Figure 4 is a schematic diagram of a RAW domain image provided according to an embodiment of the present application;

[0067] Figure 5 It is a structural schematic diagram of a training device for an image super-resolution model provided according to an embodiment of the present application;

[0068] Figure 6 is a schematic diagram of the structure of a terminal provided according to an embodiment of the present application;

[0069] Figure 7 It is a structural diagram of a server provided according to an embodiment of the present application. DETAILED DESCRIPTION

[0070] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.

[0071] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0072] In this application, the terms "first", "second", etc. are used to distinguish between identical or similar items with substantially the same effects and functions. It should be understood that there is no logical or temporal dependency between "first", "second", and "nth", nor is there a limitation on the quantity and execution order. It should also be understood that although the following description uses the terms first, second, etc. to describe various elements, these elements should not be limited by the terms.

[0073] These terms are only used to distinguish one element from another element. For example, without departing from the scope of various examples, a first action can be referred to as a second action, and similarly, a second action can also be referred to as a first action. Both the first action and the second action can be actions, and in some cases, can be separate and different actions.

[0074] Here, at least one means one or more than one, for example, at least one action can be one action, two actions, three actions, or any other action that is an integer greater than or equal to one. And multiple means two or more than two, for example, multiple actions can be two actions, three actions, or any other action that is an integer greater than or equal to two.

[0075] It should be noted that the data (including but not limited to training data and prediction data, such as user data, terminal side data, etc.) and signals involved in this application are authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions. For example, the training data involved in this application are all obtained with full authorization.

[0076] Figure 1 It is a schematic diagram of an implementation environment provided according to an embodiment of the present application, and the implementation environment may include a terminal 101 or a server 102.

[0077] In the terminal 101 or the server, there are components for taking pictures, such as an image signal processor (Image Signal Process, ISP), and an RGB domain image or a RAW domain image is obtained through the ISP.

[0078] The terminal 101 may be a smart phone with three-dimensional functions, a wearable device, a personal computer, a laptop computer, a tablet computer, a smart TV, a car terminal, etc.

[0079] The server 102 may be a single server, a server cluster consisting of multiple servers, or a cloud processing center.

[0080] The terminal 101 and the server 102 are respectively connected to a wired or wireless network.

[0081] In some embodiments, the wireless network or wired network uses standard communication technology and / or protocol. The network is usually the Internet, but it can also be any network, including but not limited to any combination of local area network (LAN), metropolitan area network (MAN), wide area network (WAN), mobile, wired or wireless network, private network or virtual private network. In some embodiments, the data exchanged through the network is represented by technology and / or format including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec) can also be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technology can also be used to replace or supplement the above data communication technology.

[0082] Figure 2 is a flow chart of a method for training an image super-resolution model according to an embodiment of the present application, such as Figure 2 As shown, in the embodiment of the present application, the application is described by taking the application on a terminal with an ISP as an example. The method comprises the following steps:

[0083] In step 201, the terminal obtains at least one first RGB domain image.

[0084] Among them, the RGB domain covers all color domains that the human eye can perceive. RGB represents the colors of three channels: red, green, and blue. Therefore, each first pixel in the first RGB image corresponds to three color channels. Therefore, the first RGB image is a three-dimensional image.

[0085] In some embodiments, the terminal obtains at least one first RGB domain image through an ISP.

[0086] In step 202, the terminal processes each of the first RGB domain images to obtain a second RGB domain image corresponding to each of the first RGB domain images. Each of the first RGB domain images and the corresponding second RGB domain image form an initial image pair, and the resolution of the second RGB domain image in each of the initial image pairs is lower than the resolution of the first RGB domain image.

[0087] The RAW domain is the original data output by the image sensor without any processing.

[0088] In some embodiments, the terminal performs degradation processing on each first RGB domain image to obtain a second RGB domain image corresponding to each first RGB domain image. Degrading each first RGB image is used to reduce the quality of each first RGB image, for example, reducing the resolution of each first RGB image to blur each first RGB image.

[0089] In step 203, the terminal samples the first RGB domain image in each of the initial image pairs to obtain a first RAW domain image corresponding to each of the first RGB domain images, and samples the second RGB domain image in each of the initial image pairs to obtain a second RAW domain image corresponding to each of the second RGB domain images. Each of the first RAW domain images and the corresponding second RAW domain form a RAW domain image pair.

[0090] In an embodiment of the present application, a terminal samples the first RGB domain image in each initial image pair to obtain a first RAW domain image corresponding to each first RGB domain image, thereby solving the problem that a high-resolution RAW domain image cannot be directly obtained based on a diffusion model.

[0091] In step 204, the terminal trains an image super-resolution model based on a plurality of RAW domain image pairs.

[0092] It should be noted that the above embodiments are described by taking application on a terminal as an example. In addition, the embodiments of the present application can also be applied on a server.

[0093] In an embodiment of the present application, by processing each high-resolution first RGB domain image, a low-resolution second RGB domain image corresponding to each first RGB domain image is obtained, and the corresponding first RGB domain image and the second RGB domain image form an RGB domain image pair. By sampling each RGB domain image pair, a RAW domain image pair corresponding to each RGB domain image pair is obtained, and an image super-resolution model is trained based on multiple RAW domain image pairs to obtain an image super-resolution model that super-resolutions the RAW domain images, thereby meeting the user's rich image requirements.

[0094] Figure 2 The embodiment shown is a brief process of the embodiment of the present application. Figure 3 The technical solution of this application is further explained. Figure 3 This is a schematic diagram of an example flow chart of a training method for an image super-resolution model provided in an embodiment of the present application. In the embodiment of the present application, the method is described by taking application to a terminal with an ISP as an example. The method comprises the following steps:

[0095] In step 301, the terminal obtains at least one first RGB domain image.

[0096] Step 301 is the same as step 201 and will not be described in detail.

[0097] In step 302, the terminal processes each first RGB domain image to obtain a second RGB domain image corresponding to each first RGB domain image. Each first RGB domain image and the corresponding second RGB domain image form an initial image pair. The resolution of the second RGB domain image in each initial image pair is lower than the resolution of the first RGB domain image.

[0098] In some embodiments, the above step 302 includes: the terminal obtains the value of each first pixel point of the first RGB domain image for each first RGB domain image; calculates the value of each first pixel point based on a preset Gaussian kernel to obtain the values ​​of multiple second pixel points in the second RGB domain image corresponding to the first RGB domain image.

[0099] In some embodiments, the terminal uses a bell curve of a Gaussian distribution as a weight distribution table of a Gaussian kernel to obtain a preset Gaussian kernel, and performs convolution calculation on the value of each first pixel point in the first RGB domain image through the preset Gaussian kernel to obtain the value of the second pixel point corresponding to each first pixel point, and then obtains a low-resolution second RGB domain image corresponding to the first RGB domain image, thereby achieving degradation of the first RGB domain image. For example, a preset Gaussian kernel is: 0.0947416 0.118318 0.0947416 0.118318 0.147761 0.118318 0.0947416 0.118318 0.0947416

[0103] Among them, 0.147761 in the above preset Gaussian kernel is the weight of each first pixel point, and 0.0947416 and 0.118318 are the weights of other first pixel points around each first pixel point.

[0104] In step 303, the terminal samples the first RGB domain image in each of the initial image pairs to obtain a first RAW domain image corresponding to each of the first RGB domain images, and samples the second RGB domain image in each of the initial image pairs to obtain a second RAW domain image corresponding to each of the second RGB domain images. Each of the first RAW domain images and the corresponding second RAW domain form a RAW domain image pair.

[0105] Among them, each pixel in the RAW domain image corresponds to a color channel, so the RAW domain image is a one-dimensional image.

[0106] In some embodiments, the step 303 includes: the terminal samples the target value of each first pixel point of the first RGB domain image in each of the initial image pairs to obtain the first RAW domain image in the RAW domain image pair corresponding to each of the initial image pairs, and samples the target value of each second pixel point of the second RGB domain image in the initial image pair to obtain the second RAW domain image in the RAW domain image pair corresponding to each of the initial image pairs;

[0107] Among them, the target value of the first pixel point is any one of the first color value, the second color value and the third color value corresponding to the first pixel point, and the target value of the second pixel point is any one of the fourth color value, the fifth color value and the sixth color value corresponding to the second pixel point.

[0108] In some embodiments, the color value may be an RGB value, and optionally, the first color value represents an R value corresponding to the RGB value, the second color value represents a G value corresponding to the RGB value, and the third color value represents a B value corresponding to the RGB value.

[0109] In some embodiments, a first pixel in the first RGB domain image corresponds one-to-one to a third pixel in the corresponding first RAW domain image, and a second pixel in the second RGB image corresponds one-to-one to a fourth pixel in the corresponding second RAW domain image.

[0110] In some embodiments, assuming that the first pixel point in the first RGB domain image is A, any one of the R value, B value, and G value corresponding to the target value of the first pixel point A is determined based on a preset mapping relationship as the value of the third pixel point corresponding to the first pixel point A, and then any one of the R value, B value, and G value corresponding to the first pixel point A is sampled to obtain the value of the pixel point corresponding to the first pixel point A. For example, assuming that the target value of the first pixel point A is determined based on a preset mapping relationship to be the corresponding R value, then the R value is sampled to obtain the value of the third pixel point corresponding to the first pixel point A (that is, the R value). Similarly, the value of the fourth pixel point corresponding to the first second pixel point C in the second RGB image is obtained.

[0111] In step 304, the terminal performs at least one iterative training on the image super-resolution model based on each of the RAW domain images.

[0112] In some embodiments, the above step 304 includes the following steps 304A to 304F:

[0113] In step 304A, the terminal obtains the i-th initial noise image during the i-th iteration; the initial noise image is randomly generated.

[0114] In step 304B, the terminal performs noise processing on the second RAW domain image in the i-th RAW domain image pair based on the i-th initial noise image to obtain an i-th RAW domain noise image; i is an integer greater than 1 or equal to 1.

[0115] In some embodiments, during the i-th iteration, an i-th initial noise image is randomly acquired, and the i-th initial noise image is superimposed M times (M is an integer greater than 1 or equal to 1) on the second RAW domain image to perform noise processing on the second RAW domain image, thereby obtaining the i-th RAW domain noise image.

[0116] In some embodiments, the above step 304B includes: the terminal obtains the i-th noise parameter; based on the i-th noise parameter, obtains the first weight of the i-th initial noise image and the second weight of the second RAW domain image in the i-th RAW domain image pair; calculates the i-th initial noise image and the first weight to obtain a first calculation result, and calculates the second RAW domain image in the i-th RAW domain image pair and the second weight to obtain a second calculation result; sums the first calculation result and the second calculation result to obtain the i-th RAW domain noise image.

[0117] In some embodiments, the noise parameter is positively correlated with M, that is, the terminal determines the i-th noise parameter based on M. When M is larger, the i-th noise parameter is also larger, and the i-th RAW domain noise image and the corresponding second RAW domain image in the i-th RAW domain image pair are closer. When M is smaller, the i-th noise parameter is also smaller, and the i-th RAW domain noise image and the pure noise image are closer. Optionally, (x, y) obeys the probability distribution p(x, y), x represents the first RAW domain image in the i-th RAW domain image pair, and y represents the second RAW domain image in the i-th RAW domain image pair. Through the above method, it is achieved that the i-th initial noise image is added to the second RAW domain image in the i-th RAW domain image pair at least once, thereby obtaining the i-th RAW domain noise image.

[0118] For example, based on the following formula (1), the i-th RAW domain noise image is obtained:

[0119]

[0120] In the above formula, z i represents the i-th RAW domain noise image, r represents the i-th noise parameter, represents the first weight, represents the first calculation result; represents the second weight, represents the second calculation result, y represents the second RAW domain image in the i-th RAW domain image pair, ∈ represents the i-th initial noise image, r obeys the probability distribution p(r), ∈ obeys the normal distribution N(0, 1).

[0121] In step 304C, the terminal splices the first RAW domain image in the i-th RAW domain image pair with the i-th RAW domain noise image to obtain the i-th spliced ​​image, wherein the splicing is channel superposition.

[0122] In some embodiments, the first RAW domain image in the ith RAW domain image pair is spliced ​​with the ith RAW domain noise image in the channel dimension to obtain the ith spliced ​​image. Optionally, the size of the first RAW domain image in the ith RAW domain image pair is 512*512*1, the size of the ith RAW domain noise image is 512*512*1, and the size of the ith spliced ​​image is 512*512*2, where 512*512 represents the size of the image, for example, the size of the first RAW domain image in the ith RAW domain image pair is 512*512*1, indicating that the first RAW domain image in the ith RAW domain image pair includes 512*512 pixels, and 1 represents one color channel.

[0123] In some embodiments, the image super-resolution model is a conditional diffusion model, and the first RAW domain image in the i-th stitched image is the condition for the i-th iterative training, so that the conditional diffusion model is trained for the i-th iterative training with the first RAW domain image as a guide.

[0124] In step 304D, the terminal performs noise prediction on the i-th spliced ​​image based on the image super-resolution model to obtain an i-th predicted noise image.

[0125] In some embodiments, the i-th spliced ​​image is input into an image super-resolution model, and noise prediction is performed on the i-th spliced ​​image based on the image super-resolution model to obtain an i-th predicted noise image. Optionally, the size of the i-th predicted noise image is 512*512*1.

[0126] In step 304E, the terminal calculates the i-th predicted noise image and the i-th initial noise image to obtain the i-th loss function value of the image super-resolution model.

[0127] In some embodiments, the terminal determines the difference between the i-th predicted noise image and the i-th initial noise image by calculating the i-th loss function value of the image super-resolution model, thereby determining the super-resolution effect of the image super-resolution model based on the difference.

[0128] In some embodiments, based on the following formula (2), the i-th loss function value of the image super-resolution model is obtained:

[0129]

[0130] In the above formula, L i represents the i-th loss function value, and fθ represents the image super-resolution model.

[0131] In step 304F, the terminal updates the parameters of the image super-resolution model based on the i-th loss function value.

[0132] In some embodiments, if the i-th loss function value meets the preset conditions, the training of the image super-resolution model is stopped; otherwise, if the i-th loss function value does not meet the preset conditions, the image super-resolution model is continued to be trained for the i+1th iterative time to continue updating the parameters of the image super-resolution model until the loss function value meets the preset conditions.

[0133] In step 305, the terminal obtains a RAW domain image to be super-resolved.

[0134] In some embodiments, the RAW domain image to be super-resolved is directly obtained based on ISP, optionally, as Figure 4 As shown, a GRBG RAW domain image to be super-resolved is directly obtained based on ISP.

[0135] In step 306, the terminal performs super-resolution on the RAW domain image based on the trained image super-resolution model to obtain a target RAW domain image.

[0136] In some embodiments, the terminal inputs the acquired RAW domain image to be super-resolved into a trained image super-resolution model, and super-resolves the RAW domain image based on the trained image super-resolution model to obtain a target RAW domain image with a higher resolution than the RAW domain image to be super-resolved.

[0137] It should be noted that the above embodiments are described by taking application on a terminal as an example. In addition, the embodiments of the present application can also be applied on a server.

[0138] In an embodiment of the present application, by processing each high-resolution first RGB domain image, a low-resolution second RGB domain image corresponding to each first RGB domain image is obtained, and the corresponding first RGB domain image and the second RGB domain image form an RGB domain image pair. By sampling each RGB domain image pair, a RAW domain image pair corresponding to each RGB domain image pair is obtained, and an image super-resolution model is trained based on multiple RAW domain image pairs to obtain an image super-resolution model that super-resolutions the RAW domain images, thereby meeting the user's rich image requirements.

[0139] Figure 5 is a structural schematic diagram of a training device 500 for an image super-resolution model provided according to an embodiment of the present application, the device comprising:

[0140] A first acquisition module 501 is used to acquire at least one first RGB domain image;

[0141] A processing module 502 is used to process each of the first RGB domain images to obtain a second RGB domain image corresponding to each of the first RGB domain images, each of the first RGB domain images and the corresponding second RGB domain image form an initial image pair, and a resolution of the second RGB domain image in each of the initial image pairs is lower than a resolution of the first RGB domain image;

[0142] a sampling module 503, configured to sample the first RGB domain image in each of the initial image pairs to obtain a first RAW domain image corresponding to each of the first RGB domain images, and to sample the second RGB domain image in each of the initial image pairs to obtain a second RAW domain image corresponding to each of the second RGB domain images, wherein each of the first RAW domain images and the corresponding second RAW domain image form a RAW domain image pair;

[0143] The training module 504 is used to train the image super-resolution model based on a plurality of the RAW domain image pairs.

[0144] In a possible implementation, the processing module 502 includes:

[0145] a sampling unit, configured to sample, for each of the initial image pairs, a target value of each first pixel point of the first RGB domain image in the initial image pair to obtain the first RAW domain image in the RAW domain image pair corresponding to each of the initial image pairs, and to sample the target value of each second pixel point of the second RGB domain image in the initial image pair to obtain the second RAW domain image in the RAW domain image pair corresponding to each of the initial image pairs;

[0146] Among them, the target value of the first pixel point is any one of the first color value, the second color value and the third color value corresponding to the first pixel point, and the target value of the second pixel point is any one of the fourth color value, the fifth color value and the sixth color value corresponding to the second pixel point.

[0147] In a possible implementation, the training module 504 includes:

[0148] The iteration unit is used to perform an iterative training on the image super-resolution model based on each RAW domain image pair.

[0149] In a possible implementation, the iteration unit includes:

[0150] An acquisition subunit is used to acquire an i-th initial noise image in an i-th iteration process; the initial noise image is randomly generated;

[0151] a processing subunit, configured to perform noise processing on the second RAW domain image in the i-th RAW domain image pair based on the i-th initial noise image to obtain an i-th RAW domain noise image; i is an integer greater than 1 or equal to 1;

[0152] a stitching subunit, configured to stitch the first RAW domain image and the i-th RAW domain noise image in the i-th RAW domain image pair to obtain an i-th stitched image;

[0153] A prediction subunit, used for performing noise prediction on the i-th spliced ​​image based on the image super-resolution model to obtain an i-th predicted noise image;

[0154] A calculation subunit, used to calculate the i-th predicted noise image and the i-th initial noise image to obtain the i-th loss function value of the image super-resolution model;

[0155] The updating subunit is used to update the parameters of the image super-resolution model based on the i-th loss function value.

[0156] In a possible implementation, the processing subunit includes:

[0157] A first acquisition subunit is used to acquire the i-th noise parameter;

[0158] a second acquisition subunit, configured to acquire, based on the i-th noise parameter, a first weight of the i-th initial noise image and a second weight of the second RAW domain image in the i-th RAW domain image pair;

[0159] a calculation subunit, configured to calculate the i-th initial noise image and the first weight to obtain a first calculation result, and to calculate the second RAW domain image in the i-th RAW domain image pair and the second weight to obtain a second calculation result;

[0160] The summing subunit is used to sum the first calculation result and the second calculation result to obtain the i-th RAW domain noise image.

[0161] In a possible implementation, the device further includes:

[0162] The second acquisition module is used to acquire the RAW domain image to be super-resolved;

[0163] The super-resolution module is used to super-resolution the RAW domain image based on a trained image super-resolution model to obtain a target RAW domain image.

[0164] It should be noted that: when executing the corresponding steps, the training device for the image super-resolution model provided in the above embodiment only uses the division of the above functional modules as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the training device for the image super-resolution model provided in the above embodiment and the training method embodiment of the image super-resolution model belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0165] In an embodiment of the present application, by processing each high-resolution first RGB domain image, a low-resolution second RGB domain image corresponding to each first RGB domain image is obtained, and the corresponding first RGB domain image and the second RGB domain image form an RGB domain image pair. By sampling each RGB domain image pair, a RAW domain image pair corresponding to each RGB domain image pair is obtained, and an image super-resolution model is trained based on multiple RAW domain image pairs to obtain an image super-resolution model that super-resolutions the RAW domain images, thereby meeting the user's rich image requirements.

[0166] An embodiment of the present application also provides a computer device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and the processor implements the above-mentioned image super-resolution model training method when executing the computer program.

[0167] Taking computer equipment as the terminal as an example, Figure 6 is a schematic diagram of the structure of a terminal provided in an embodiment of the present application, see Figure 6 The terminal 600 may be a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer or a desktop computer. The terminal 600 may also be called a user device, a portable terminal, a laptop terminal, a desktop terminal or other names.

[0168] Typically, the terminal 600 includes a processor 601 and a memory 602 .

[0169] The processor 601 may include one or more processing cores, such as a 4-core processor, a 6-core processor, etc. The processor 601 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 601 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 601 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 601 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0170] The memory 602 may include one or more computer-readable storage media, which may be non-transitory. The memory 602 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 602 is used to store at least one program code, which is used to be executed by the processor 601 to implement the process executed by the terminal in the training method of the image super-resolution model provided in the method embodiment of the present disclosure.

[0171] In some embodiments, the terminal 600 may further optionally include: a peripheral device interface 603 and at least one peripheral device. The processor 601, the memory 602 and the peripheral device interface 603 may be connected via a bus or a signal line. Each peripheral device may be connected to the peripheral device interface 603 via a bus, a signal line or a circuit board. Specifically, the peripheral device includes: at least one of a radio frequency circuit 604, a display screen 605, a camera assembly 606, an audio circuit 607 and a power supply 608.

[0172] The peripheral device interface 603 may be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 601 and the memory 602. In some embodiments, the processor 601, the memory 602, and the peripheral device interface 603 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 601, the memory 602, and the peripheral device interface 603 may be implemented on a separate chip or circuit board, which is not limited in the embodiments of the present disclosure.

[0173] The radio frequency circuit 604 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 604 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 604 converts the electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. In some embodiments, the radio frequency circuit 604 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and the like. The radio frequency circuit 604 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes, but is not limited to: a metropolitan area network, various generations of mobile communication networks (2G, 3G, 4G and 5G), a wireless local area network and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 604 may also include circuits related to NFC (Near Field Communication), which is not limited in the present disclosure.

[0174] The display screen 605 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos and any combination thereof. When the display screen 605 is a touch display screen, the display screen 605 also has the ability to collect touch signals on the surface or above the surface of the display screen 605. The touch signal can be input to the processor 601 as a control signal for processing. At this time, the display screen 605 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, the display screen 605 can be one, set on the front panel of the terminal 600; in other embodiments, the display screen 605 can be at least two, respectively set on different surfaces of the terminal 600 or in a folding design; in other embodiments, the display screen 605 can be a flexible display screen, set on the curved surface or folding surface of the terminal 600. Even, the display screen 605 can also be set to a non-rectangular irregular shape, that is, a special-shaped screen. The display screen 605 can be made of materials such as LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode) and the like.

[0175] The camera assembly 606 is used to capture images or videos. In some embodiments, the camera assembly 606 includes a front camera and a rear camera. Typically, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize the panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 606 may also include a flash. The flash may be a monochrome temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.

[0176] The audio circuit 607 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals and input them into the processor 601 for processing, or input them into the radio frequency circuit 604 to achieve voice communication. For the purpose of stereo acquisition or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 600. The microphone may also be an array microphone or an omnidirectional acquisition microphone. The speaker is used to convert the electrical signal from the processor 601 or the radio frequency circuit 604 into sound waves. The speaker may be a traditional film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for purposes such as ranging. In some embodiments, the audio circuit 607 may also include a headphone jack.

[0177] The power supply 608 is used to power various components in the terminal 600. The power supply 608 can be an alternating current, a direct current, a disposable battery, or a rechargeable battery. When the power supply 608 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology.

[0178] In some embodiments, the terminal 600 further includes one or more sensors 609 , including but not limited to: an acceleration sensor 610 , a gyroscope sensor 611 , a pressure sensor 66 , an optical sensor 613 , and a proximity sensor 614 .

[0179] The acceleration sensor 610 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established by the terminal 600. For example, the acceleration sensor 610 can be used to detect the components of gravity acceleration on the three coordinate axes. The processor 601 can control the display screen 605 to display the user page in a horizontal view or a vertical view according to the gravity acceleration signal collected by the acceleration sensor 610. The acceleration sensor 610 can also be used for collecting game or user motion data.

[0180] The gyro sensor 611 can detect the body direction and rotation angle of the terminal 600, and the gyro sensor 611 can cooperate with the acceleration sensor 610 to collect the user's three-dimensional actions on the terminal 600. The processor 601 can implement the following functions based on the data collected by the gyro sensor 611: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.

[0181] The pressure sensor 612 can be set in the side frame of the terminal 600 and / or the lower layer of the display screen 605. When the pressure sensor 612 is set in the side frame of the terminal 600, it can detect the user's holding signal of the terminal 600, and the processor 601 performs left and right hand recognition or shortcut operation according to the holding signal collected by the pressure sensor 612. When the pressure sensor 612 is set in the lower layer of the display screen 605, the processor 601 controls the operability controls on the UI page according to the user's pressure operation on the display screen 605. The operability controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.

[0182] The optical sensor 613 is used to collect the ambient light intensity. In one embodiment, the processor 601 can control the display brightness of the display screen 605 according to the ambient light intensity collected by the optical sensor 613. Specifically, when the ambient light intensity is high, the display brightness of the display screen 605 is increased; when the ambient light intensity is low, the display brightness of the display screen 605 is reduced. In another embodiment, the processor 601 can also dynamically adjust the shooting parameters of the camera assembly 606 according to the ambient light intensity collected by the optical sensor 613.

[0183] The proximity sensor 614, also called a distance sensor, is usually arranged on the front panel of the terminal 600. The proximity sensor 614 is used to collect the distance between the user and the front of the terminal 600. In one embodiment, when the proximity sensor 614 detects that the distance between the user and the front of the terminal 600 is gradually decreasing, the processor 601 controls the display screen 605 to switch from the screen-on state to the screen-off state; when the proximity sensor 614 detects that the distance between the user and the front of the terminal 600 is gradually increasing, the processor 601 controls the display screen 605 to switch from the screen-off state to the screen-on state.

[0184] Those skilled in the art will understand that Figure 6 The structure shown in the figure does not constitute a limitation on the terminal 600, and the terminal 600 may include more or less components than those shown in the figure, or combine some components, or adopt a different component arrangement.

[0185] Take the computer device as a server as an example. Figure 7It is a structural diagram of a server provided in an embodiment of the present application. The server 700 may have relatively large differences due to different configurations or performances, and may include one or more processors (Central Processing Units, CPU) 701 and one or more memories 702, wherein the one or more memories 702 store at least one computer program, and the at least one computer program is loaded and executed by the one or more processors 701 to implement the above-mentioned image super-resolution model training method. Of course, the server 700 may also have components such as a wired or wireless network interface, a keyboard, and an input and output interface for input and output. The server 700 may also include other components for implementing device functions, which will not be repeated here.

[0186] The embodiment of the present application also provides a computer-readable storage medium, the computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the above-mentioned image super-resolution model training method. Optionally, the computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, etc.

[0187] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.

[0188] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A training method for an image super-resolution model, It is characterized in that include: Acquire at least one first RGB domain image; Processing each of the first RGB domain images to obtain a second RGB domain image corresponding to each of the first RGB domain images, each of the first RGB domain images and the corresponding second RGB domain image forming an initial image pair, and a resolution of the second RGB domain image in each of the initial image pairs is lower than a resolution of the first RGB domain image; Sampling the first RGB domain image in each of the initial image pairs to obtain a first RAW domain image corresponding to each of the first RGB domain images, and sampling the second RGB domain image in each of the initial image pairs to obtain a second RAW domain image corresponding to each of the second RGB domain images, wherein each of the first RAW domain images and the corresponding second RAW domain image form a RAW domain image pair; Performing at least one iterative training on the image super-resolution model based on each of the RAW domain images; In the i-th iteration process, the i-th initial noise image is obtained; The initial noise image is randomly generated; Get the i-th noise parameter; Based on the i-th noise parameter, obtaining a first weight of the i-th initial noise image and a second weight of the second RAW domain image in the i-th RAW domain image pair; Calculating the i-th initial noise image and the first weight to obtain a first calculation result, and calculating the second RAW domain image in the i-th RAW domain image pair and the second weight to obtain a second calculation result; Summing the first calculation result and the second calculation result to obtain the i-th RAW domain noise image; In the channel dimension, stitching the first RAW domain image and the i-th RAW domain noise image in the i-th RAW domain image pair to obtain an i-th stitched image; Performing noise prediction on the i-th spliced ​​image based on the image super-resolution model to obtain an i-th predicted noise image; Calculating the i-th predicted noise image and the i-th initial noise image to obtain an i-th loss function value of the image super-resolution model; The parameters of the image super-resolution model are updated based on the i-th loss function value.

2. The method according to claim 1, It is characterized in that The processing of each of the first RGB domain images includes: For each of the first RGB domain images, obtaining a value of each first pixel point of the first RGB domain image; The value of each of the first pixel points is calculated based on a preset Gaussian kernel to obtain values ​​of multiple second pixel points in the second RGB domain image corresponding to the first RGB domain image.

3. The method according to claim 1, It is characterized in that The step of sampling the first RGB domain image in each of the initial image pairs to obtain a first RAW domain image corresponding to each of the first RGB domain images, and sampling the second RGB domain image in each of the initial image pairs to obtain a second RAW domain image corresponding to each of the second RGB domain images, comprises: For each of the initial image pairs, the target values ​​of each first pixel point of the first RGB domain image in the initial image pair are sampled to obtain the first RAW domain image in the RAW domain image pair corresponding to each of the initial image pairs, and the target values ​​of each second pixel point of the second RGB domain image in the initial image pair are sampled to obtain the second RAW domain image in the RAW domain image pair corresponding to each of the initial image pairs; wherein the target value of the first pixel point is any one of the first color value, the second color value and the third color value corresponding to the first pixel point, and the target value of the second pixel point is any one of the fourth color value, the fifth color value and the sixth color value corresponding to the second pixel point.

4. The method according to claim 1, It is characterized in that The method further comprises: Obtain the RAW domain image to be super-resolved; The RAW domain image is super-resolved based on a trained image super-resolution model to obtain a target RAW domain image.

5. A training device for an image super-resolution model, It is characterized in that include: An acquisition module, used to acquire at least one first RGB domain image; a processing module, configured to process each of the first RGB domain images to obtain a second RGB domain image corresponding to each of the first RGB domain images, wherein each of the first RGB domain images and the corresponding second RGB domain image form an initial image pair, and a resolution of the second RGB domain image in each of the initial image pairs is lower than a resolution of the first RGB domain image; a sampling module, configured to sample the first RGB domain image in each of the initial image pairs to obtain a first RAW domain image corresponding to each of the first RGB domain images, and to sample the second RGB domain image in each of the initial image pairs to obtain a second RAW domain image corresponding to each of the second RGB domain images, wherein each of the first RAW domain images and the corresponding second RAW domain image form a RAW domain image pair; A training module, configured to perform at least one iterative training on the image super-resolution model based on each of the RAW domain images; In the i-th iteration process, the i-th initial noise image is obtained; The initial noise image is randomly generated; Get the i-th noise parameter; Based on the i-th noise parameter, obtaining a first weight of the i-th initial noise image and a second weight of the second RAW domain image in the i-th RAW domain image pair; Calculating the i-th initial noise image and the first weight to obtain a first calculation result, and calculating the second RAW domain image in the i-th RAW domain image pair and the second weight to obtain a second calculation result; Summing the first calculation result and the second calculation result to obtain the i-th RAW domain noise image; In the channel dimension, stitching the first RAW domain image and the i-th RAW domain noise image in the i-th RAW domain image pair to obtain an i-th stitched image; Performing noise prediction on the i-th spliced ​​image based on the image super-resolution model to obtain an i-th predicted noise image; Calculating the i-th predicted noise image and the i-th initial noise image to obtain an i-th loss function value of the image super-resolution model; The parameters of the image super-resolution model are updated based on the i-th loss function value.

6. A computer device, It is characterized in that The computer device includes a processor and a memory, the memory is used to store at least one program, and the at least one program is loaded by the processor and executes the training method of the image super-resolution model according to any one of claims 1 to 4.

7. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores at least one program, and the at least one program is loaded and executed by the processor to implement the training method of the image super-resolution model according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Super-resolution image reconstruction method and electronic device

    CN110766610A

  • Image super-resolution method and system based on image noise prediction mechanism

    CN116342393A