A Super-Resolution Reconstruction Method for Camera Imaging Models
By super-resolution reconstruction of RAW image sequences in the camera imaging model and color correction, the problems of artifacts and visual quality in the existing RGB image reconstruction methods are solved, and higher quality image reconstruction is achieved.
Patent Information
- Application Number
- CN202210548154.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-18
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-05-18
AI Technical Summary
The existing super-resolution reconstruction methods are mainly based on RGB images, resulting in artifacts in the reconstruction results, affecting visual quality, and are not suitable for camera imaging models.
The super-resolution reconstruction method for the camera imaging model is used to super-resolution reconstruction of the RAW image sequence, and the reconstructed RAW image is converted into RGB images through the color correction model.
By combining more detailed information, unnecessary artifacts are eliminated, and reconstruction quality and visual senses are significantly improved.
Smart Images

Figure CN114897697B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing and reconstruction, and more specifically to a super-resolution reconstruction method for a camera imaging model. Background Art
[0002] Super-resolution reconstruction is an image processing technology that reconstructs a low-resolution image into a high-resolution image with rich texture details. Single-frame image super-resolution reconstruction means that the input image data is a single-frame image, and it is reconstructed into a single-frame high-resolution image. Currently, most image super-resolution reconstruction methods adopt deep learning methods, which are data-driven. For super-resolution reconstruction methods, existing methods generally process and study based on RGB images. However, a large amount of detailed information and image features are lost during the imaging process of RGB images themselves, resulting in artifacts in the reconstruction results and affecting the visual quality of the reconstruction. In the actual camera imaging process, the camera front-end analyzes photon information to first obtain a RAW image sequence, and then the camera Image Signal Processing (ISP) processes the fused RAW image in terms of color and brightness to generate an RGB image that conforms to the human eye's perception. Therefore, although RGB images are more in line with the human eye's perception, they are not suitable for super-resolution reconstruction.
[0003] Therefore, how to provide a super-resolution reconstruction method for a camera imaging model that can effectively improve the reconstruction quality and visual perception is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0004] In view of this, the present invention provides a super-resolution reconstruction method for a camera imaging model, which performs super-resolution reconstruction on the RAW image sequence and then post-processes the reconstructed RAW image to obtain an RGB image, which can effectively improve the reconstruction quality and visual perception.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] A super-resolution reconstruction method for a camera imaging model, comprising:
[0007] Construct an image super-resolution reconstruction model and a color correction model;
[0008] Obtain a RAW image sequence and determine a reference frame;
[0009] Based on the spatio-temporal correlation between multiple frames of RAW images obtained by the image super-resolution reconstruction model, and combining the spatio-temporal correlation between each frame of RAW images, perform super-resolution reconstruction on the reference frame to obtain a high-resolution reconstructed image;
[0010] Perform a transformation of the reconstructed image from the image linear domain to the color domain based on the color correction model.
[0011] Furthermore, the expression of the RAW image sequence is:
[0012]
[0013] where represents the set of RAW image sequences, represents the t-th frame in the image sequence, and there are N frames in total.
[0014] Furthermore, the image super-resolution reconstruction model includes a feature alignment module, a feature fusion module, and a feature reconstruction module;
[0015] The feature alignment module aligns the RAW image sequence in the feature space, aligns each sequence frame to the reference frame, and obtains a set of aligned feature maps;
[0016] The feature fusion module performs fusion and aggregation on the set of aligned feature maps to obtain a fused feature map;
[0017] The feature reconstruction module reconstructs the fused feature map to obtain a high-resolution reconstructed image.
[0018] Furthermore, the feature alignment module uses the following formula to align the RAW image sequence:
[0019]
[0020] where t represents the target frame number, F t represents the target frame, r represents the reference frame number, F r represents the reference frame; each target frame except the reference frame needs to be aligned with the reference frame; F t align represents the feature map after aligning the t-th target frame with the reference frame; represents the set of aligned feature maps, and t ∈ [1, N] represents from the 1st frame to the Nth frame.
[0021] Furthermore, the feature fusion module uses three-dimensional convolution to perform fusion and aggregation on the set of aligned feature maps, and its expression is:
[0022]
[0023] where ω t,k represents the weight of the convolution kernel; p represents the center position of the convolution kernel; p k represents K sampling points of the preset value of the convolution kernel, Δp [t,p],k represents the offset of the deformed convolution kernel; Ffusion Represents the fused feature map obtained through the feature fusion module.
[0024] Furthermore, the feature reconstruction module is composed of 16 cascaded residual blocks.
[0025] Furthermore, the color correction model adopts the U-Net network architecture, and the U-Net network architecture is trained using a hybrid color loss.
[0026] Furthermore, the calculation formula of the hybrid color loss is:
[0027] L color = L1 + 0.5·L lap
[0028] where L1 and L lap The calculation formulas are expressed as:
[0029] L1 = ||I HR - I SR ||1
[0030]
[0031] where L1 represents the pixel loss function; L lap represents the Laplacian loss, which calculates the pixel loss at different image scales to blur the details of the image and pay more attention to the color information of the image; φ j represents the j-th layer Laplacian pyramid; I HR represents the high-resolution image label, and I SR represents the reconstructed image.
[0032] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a super-resolution reconstruction method for a camera imaging model. Based on the imaging model of the camera, super-resolution reconstruction is performed on the RAW image sequence, and then the transformed RAW image after reconstruction is used to obtain the RGB color image. In the existing method of directly reconstructing the RGB image, since the RGB image itself has undergone complex nonlinear transformations, a lot of detail information has been lost in the RGB image itself, resulting in artifacts in the super-resolution result based on the RGB image. The present invention uses RAW images for super-resolution reconstruction, which can combine more detail information, eliminate unnecessary artifacts, and improve the reconstruction quality and visual perception. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to the provided drawings.
[0034] Figure 1 It is a flowchart of the super-resolution reconstruction method for a camera imaging model provided by the present invention;
[0035] Figure 2 It is a comparison diagram of image reconstruction using the super-resolution reconstruction method for a camera imaging model of the present invention and the prior art method;
[0036] Figure 3 It is a schematic structural diagram of the feature reconstruction module provided by the present invention. Detailed implementation manners
[0037] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0038] The embodiments of the present invention disclose a super-resolution reconstruction method for a camera imaging model, including the following steps:
[0039] Construct an image super-resolution reconstruction model and a color correction model;
[0040] Obtain a RAW image sequence and determine a reference frame;
[0041] Based on the spatio-temporal correlation between multiple frames of RAW images obtained by the image super-resolution reconstruction model, and combining the spatio-temporal correlation between each frame of RAW images, perform super-resolution reconstruction on the reference frame to obtain a high-resolution reconstructed image;
[0042] Based on the color correction model, perform a transformation on the reconstructed image from the image linear domain to the color domain.
[0043] In a specific embodiment, the image super-resolution reconstruction model includes a feature alignment module, a feature fusion module, and a feature reconstruction module;
[0044] The feature alignment module aligns the RAW image sequence in the feature space, aligns each sequence frame to the reference frame, and obtains a set of aligned feature maps;
[0045] The feature fusion module fuses and aggregates the aligned feature map set to obtain a fused feature map;
[0046] The feature reconstruction module reconstructs the fused feature map to obtain a high-resolution reconstructed image.
[0047] Specifically, for the RAW image sequence obtained by imaging, it consists of N similar images, that is
[0048]
[0049] where represents the RAW image sequence set, represents the t-th frame in the image sequence, and there are N frames in total.
[0050] As Figure 1 shown, the feature alignment module, the feature fusion module, and the feature reconstruction module together constitute an image super-resolution reconstruction model. The RAW image sequence is reconstructed in the image linear domain through the image super-resolution reconstruction model to obtain Then, through the color correction model, the reconstructed image is transformed from the image linear domain to the normal color domain I SR .
[0051] In a specific embodiment, the feature alignment module constructs a Pyramid Cascading and Deformable Convolution (PCD) using a Deformable Convolution Layer (DCN). Through the PCD, multiple frames of the RAW image sequence are aligned in the feature space, and each sequence frame is aligned to the reference frame image to obtain aggregated image features and eliminate unnecessary artifacts. The alignment function of the PCD module can be expressed by the following formula
[0052]
[0053] where t represents the target frame number, F t represents the target frame, r represents the reference frame number, and F r represents the reference frame; each target frame except the reference frame needs to be aligned with the reference frame; F t align represents the feature map after the t-th target frame is aligned with the reference frame; represents the aligned feature map set, and t ∈ [1, N] represents from the 1st frame to the Nth frame.
[0054] The feature fusion module is a three-dimensional deformable convolution fusion module (3D Deformable Convolution Fusion Module) designed based on the deformable convolution layer and capable of processing temporal information. Through the feature fusion module, the previously obtained set of aligned feature maps can be fused and aggregated, thereby enabling the extraction of more valuable image features. The specific implementation process is as follows: perform a convolution weighting operation on multiple frames of aligned feature maps in the time domain, where the feature maps closer to the reference frame are given greater weights. After the fusion of multiple frames of feature maps, they become one frame of feature map, and the high-resolution image is restored through subsequent feature reconstruction; it can be expressed by the following formula:
[0055]
[0056] Among them, ω t,k represents the weight of the convolution kernel; p represents the center position of the convolution kernel; p k represents K sampling points of the preset value of the convolution kernel, and Δp [t,p],k represents the offset of the deformable convolution kernel; F fusion represents the fused feature map obtained through the feature fusion module.
[0057] The feature reconstruction module generates an output high-resolution image from the fused feature map. 16 residual blocks are cascaded to form the feature reconstruction module. Finally, a upsampling layer is used to obtain a higher-resolution image. In addition, an extra residual connection is added to directly introduce low-frequency information into the feature reconstruction module. The specific structure of the feature reconstruction module is shown in Figure 3 , and the dimension of the feature map is kept the same for each layer. The image features undergo repeated convolution operations and gradually retain high-frequency information under the constraint of the loss function. The final image features are fed into the upsampling layer to achieve the restoration of the high-resolution image. In addition, the feature map output by the first residual block is additionally added to the last layer of feature map to introduce low-frequency information. The introduction of low-frequency information can make the reconstructed high-resolution image clearer.
[0058] In other embodiments, after the RAW image is super-resolution upsampled in the image linear domain in the present invention, it needs to undergo a color transformation to convert the image to the sRGB domain to meet the human eye's perception. The embodiments of the present invention construct a simple U-Net network to achieve the conversion from the image linear domain to the image sRGB domain.
[0059] Specifically, the embodiments of the present invention train the color correction model by performing backpropagation calculation of the gradient through pixel loss. Use pixel loss as the loss function of the color correction model to achieve super-resolution reconstruction in the image linear domain. The process can be expressed as follows:
[0060] In order to complete the transformation of the image from the linear domain to the color domain, this embodiment adopts mixed color loss to preserve the detailed texture information of the image as much as possible while ensuring color transformation. The mixed color loss is expressed as L color , by L1 and L lap The calculation formula is as follows:
[0061] L color =L1+0.5·L lap
[0062] Among them, L1 and L lap The calculation formula is expressed as:
[0063] L1=||I HR -I SR ||1
[0064]
[0065] Among them, L1 represents the pixel loss function; L lap represents the Laplace loss, which means calculating pixel loss at different image scales to blur the details of the image and focus more on the color information of the image; φ j represents the j-th level of the Laplacian pyramid; I HR Indicates the high-resolution image label, I SR Represents the reconstructed image.
[0066] To further demonstrate the effectiveness of the method of the present invention, the following experiments were performed:
[0067] The experiment uses the method of the present invention and the prior art models HighResNet and DeepJoint+RRDB for image reconstruction. The training image dataset used is the public image dataset DIV2K, and the verification images are the public image datasets BSD100, Urban100 and Manga109. The evaluation indicators are PSNR and SSIM. The higher the PSNR and SSIM, the higher the image quality. The experimental results are shown in Table 1 and Figure 2 shown.
[0068] Table 1 Comparison results
[0069]
[0070] Through Table 1 and Figure 2 The comparison results show that the method of the present invention has significant improvements in both performance indicators and visual effects.
[0071] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.
[0072] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A super-resolution reconstruction method for a camera imaging model, characterized in that, Including: Construct an image super-resolution reconstruction model and a color correction model; Obtain a RAW image sequence and determine a reference frame; Based on the image super-resolution reconstruction model, obtain the spatio-temporal correlation between multiple frames of RAW images, and combine the spatio-temporal correlation between each frame of RAW images to perform super-resolution reconstruction on the reference frame to obtain a high-resolution reconstructed image; Based on the color correction model, perform a transformation on the reconstructed image from the image linear domain to the color domain; The image super-resolution reconstruction model includes a feature alignment module, a feature fusion module, and a feature reconstruction module; The feature alignment module aligns the RAW image sequence in the feature space, aligns each sequence frame to the reference frame, and obtains a set of aligned feature maps; The feature alignment module is a cascaded pyramid alignment module constructed using deformable convolutional layers. It aligns the multi-frame RAW image sequence in the feature space through PCD, aligns each sequence frame to the reference frame image, and obtains aggregated image features to eliminate unnecessary artifacts; The feature fusion module fuses and converges the set of aligned feature maps to obtain a fused feature map; The feature reconstruction module reconstructs the fused feature map to obtain a high-resolution reconstructed image; The feature reconstruction module is composed of 16 cascaded residual blocks; The cascading method of the 16 residual blocks is as follows: The low-resolution fused features first pass through a convolutional layer and a residual block, then pass through 15 1*1 convolutional layers and residual blocks, and finally obtain a high-resolution image through an upsampling layer; among them, a feed-forward channel is added between adjacent convolutional layers, and finally the feature map output by the first residual block is additionally added to the last layer feature map to introduce low-frequency information.
2. The super-resolution reconstruction method for a camera imaging model according to claim 1, characterized in that, The expression of the RAW image sequence is: Among them, represents a set of RAW image sequences, represents the t-th frame in the image sequence, and there are N frames in total.
3. The super-resolution reconstruction method for a camera imaging model according to claim 1, characterized in that, The feature alignment module uses the following formula to align the RAW image sequence: Among them, t represents the target frame number, and F t represents the target frame, r represents the reference frame number, and F r represents the reference frame; each target frame except the reference frame needs to be aligned with the reference frame; F t align represents the feature map after the t-th target frame is aligned with the reference frame; represents the set of aligned feature maps, and t ∈ [1, N] represents from the 1st frame to the Nth frame.
4. The super-resolution reconstruction method for a camera imaging model according to claim 1, characterized in that, The feature fusion module uses three-dimensional convolution to fuse and converge the set of aligned feature maps, and its expression is: Among them, ω t,k represents the weight of the convolution kernel; p represents the center position of the convolution kernel; p k represents K sampling points of the preset value of the convolution kernel, and Δp [t,p],k represents the offset of the deformable convolution kernel; F fusion represents the fused feature map obtained through the feature fusion module.
5. The super-resolution reconstruction method for a camera imaging model according to claim 1, characterized in that, The color correction model adopts a U-Net network architecture, and it uses a hybrid color loss to train the U-Net network architecture.
6. The super-resolution reconstruction method for a camera imaging model according to claim 5, characterized in that, The calculation formula of the hybrid color loss is: L color = L1 + 0.5·L lap Among them, L1 and L lap are expressed by the calculation formula as: L1 = ||I HR -I SR ||1 Among them, L1 represents the pixel loss function; L lap represents the Laplacian loss, which calculates the pixel loss at different image scales to blur the details of the image and focus more on the color information of the image; φ j represents the j-th layer Laplacian pyramid; I HR represents the high-resolution image label, I SR represents the reconstructed image.