Image processing method and device, electronic equipment and computer readable storage medium
Patent Information
- Application Number
- CN202211499327.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-28
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2042-11-28
AI Technical Summary
[0003]然而,传统的图像处理方法,存在图像处理准确性不高的问题
[0025]上述图像处理方法、装置、电子设备、计算机可读存储介质和计算机程序产品,基于至少两个原始图像中参考帧的噪声信息,对从至少两个原始图像中提取的第一图像特征进行去噪处理和解马赛克处理,得到频率低于预设频率阈值的第二图像特征;对第二图像特征进行超分辨率处理,得到频率高于或等于预设频率阈值的第三图像特征;也就是说,通过去噪处理和解马赛克处理可以更准确地得到图像的低频色彩轮廓等信息,且通过超分辨率处理可以更准确地得到图像的高频纹理细节等信息,从而根据第二图像特征和第三图像特征,生成去噪、解马赛克和超分辨率的目标图像,提高了图像处理的准确性。
Smart Images

Figure CN118115375B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of imaging technology, and in particular to an image processing method, apparatus, electronic device, and computer-readable storage medium. Background Technology
[0002] With the development of imaging technology, multi-frame image fusion technology has emerged. In the process of multi-frame image fusion, ISP (Image Signal Processing) processing is usually required, including de-mosaicing, noise reduction, white balance, tone mapping, contrast enhancement, etc., so as to obtain the final image.
[0003] However, traditional image processing methods suffer from low accuracy. Summary of the Invention
[0004] This application provides an image processing method, apparatus, electronic device, computer-readable storage medium, and computer program product that can improve the accuracy of image processing.
[0005] Firstly, this application provides an image processing method. The method includes:
[0006] Based on the noise information of reference frames in at least two original images, the first image features extracted from the at least two original images are subjected to denoising and de-mosaic processing to obtain the second image features; the frequency of the second image features is lower than a preset frequency threshold.
[0007] The second image feature is subjected to super-resolution processing to obtain a third image feature; the frequency of the third image feature is higher than or equal to a preset frequency threshold.
[0008] A target image is generated based on the second image features and the third image features.
[0009] Secondly, this application also provides an image processing apparatus. The apparatus includes:
[0010] The first processing module is used to perform denoising and de-mosaic processing on the first image features extracted from the at least two original images based on the noise information of the reference frames in at least two original images to obtain the second image features; the frequency of the second image features is lower than a preset frequency threshold.
[0011] The second processing module is used to perform super-resolution processing on the second image feature to obtain a third image feature; the frequency of the third image feature is higher than or equal to a preset frequency threshold.
[0012] An image generation module is used to generate a target image based on the second image features and the third image features.
[0013] Thirdly, this application also provides an electronic device. The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:
[0014] Based on the noise information of reference frames in at least two original images, the first image features extracted from the at least two original images are subjected to denoising and de-mosaic processing to obtain the second image features; the frequency of the second image features is lower than a preset frequency threshold.
[0015] The second image feature is subjected to super-resolution processing to obtain a third image feature; the frequency of the third image feature is higher than or equal to a preset frequency threshold.
[0016] A target image is generated based on the second image features and the third image features.
[0017] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:
[0018] Based on the noise information of reference frames in at least two original images, the first image features extracted from the at least two original images are subjected to denoising and de-mosaic processing to obtain the second image features; the frequency of the second image features is lower than a preset frequency threshold.
[0019] The second image feature is subjected to super-resolution processing to obtain a third image feature; the frequency of the third image feature is higher than or equal to a preset frequency threshold.
[0020] A target image is generated based on the second image features and the third image features.
[0021] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:
[0022] Based on the noise information of reference frames in at least two original images, the first image features extracted from the at least two original images are subjected to denoising and de-mosaic processing to obtain the second image features; the frequency of the second image features is lower than a preset frequency threshold.
[0023] The second image feature is subjected to super-resolution processing to obtain a third image feature; the frequency of the third image feature is higher than or equal to a preset frequency threshold.
[0024] A target image is generated based on the second image features and the third image features.
[0025] The aforementioned image processing method, apparatus, electronic device, computer-readable storage medium, and computer program product, based on noise information from reference frames in at least two original images, perform denoising and de-mosaic processing on first image features extracted from at least two original images to obtain second image features with frequencies lower than a preset frequency threshold; and perform super-resolution processing on the second image features to obtain third image features with frequencies higher than or equal to the preset frequency threshold. In other words, denoising and de-mosaic processing can more accurately obtain low-frequency color contours and other information of the image, and super-resolution processing can more accurately obtain high-frequency texture details and other information of the image. Thus, based on the second and third image features, a denoised, de-mosaiced, and super-resolution target image is generated, improving the accuracy of image processing. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is a flowchart of an image processing method in one embodiment;
[0028] Figure 2 A flowchart of an image processing method in another embodiment;
[0029] Figure 3 A flowchart of an image processing method in another embodiment;
[0030] Figure 4 A flowchart of an image processing method in another embodiment;
[0031] Figure 5 A flowchart of an image processing method in another embodiment;
[0032] Figure 6 A flowchart of an image processing method in another embodiment;
[0033] Figure 7 A flowchart for generating multiple frames of images to be input to the network in another embodiment;
[0034] Figure 8 A flowchart of an image processing method in another embodiment;
[0035] Figure 9 This is a structural block diagram of an image processing device in one embodiment;
[0036] Figure 10 This is a diagram of the internal structure of an electronic device in one embodiment. Detailed Implementation
[0037] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0038] In one embodiment, such as Figure 1 As shown, an image processing method is provided. This embodiment illustrates the application of this method to an electronic device, which can be a terminal or a server. It is understood that this method can also be applied to a system including both a terminal and a server, and implemented through interaction between the terminal and the server. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, smart cars, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc. The server can be a standalone server or a server cluster consisting of multiple servers.
[0039] In this embodiment, the image processing method includes the following steps:
[0040] Step S102: Based on the noise information of the reference frame in at least two original images, the first image features extracted from at least two original images are subjected to denoising and de-mosaic processing to obtain the second image features; the frequency of the second image features is lower than a preset frequency threshold.
[0041] Understandably, the original image can be a RAW domain image, that is, raw, unprocessed image data.
[0042] A reference frame is an image used for reference during the registration process. Optionally, the reference frame can be the image with the highest sharpness among at least two original images, or the image with the highest brightness among at least two original images; there is no limitation on this. The noise information of the reference frame can be a noise map of the reference frame.
[0043] Denoising refers to the process of reducing noise in digital images. Demosaicing (also spelled de-mosaicing, demosaicking, or debayering) is a digital image processing algorithm that aims to reconstruct a full-color image from incomplete color samples output by a photosensitive imager covered with a color filter array (CFA). This method is also known as CFA interpolation or color reconstruction.
[0044] The preset frequency threshold can be set as needed. The frequency of the second image feature is lower than the preset frequency threshold, meaning the second image feature includes low-frequency image features. Low frequency refers to slowly changing colors or slowly changing grayscale, representing a continuously gradational area; this is the low-frequency part. For an image, the content within the edges is low-frequency, and this content within the edges contains most of the image information, that is, the general outline and contour of the image, which is approximate information about the image. For example, the second image feature may include contour, color, and other information.
[0045] Optionally, the electronic device acquires at least two raw images in the RAW domain using an image sensor; inputs the raw images in the at least two RAW domains into a joint demosaicing, denoising, and super-resolution network; and performs demosaicing, denoising, and super-resolution processing on the raw images in the at least two RAW domains in the joint demosaicing, denoising, and super-resolution network to obtain the target image.
[0046] Optionally, in the joint demosaicing, denoising, and super-resolution network, a first image feature is extracted from at least two original images, and a reference frame is determined from at least two original images; noise information of the reference frame is obtained; based on the noise information of the reference frame, the first image feature is denoised and demosaiced to obtain a second image feature.
[0047] Step S104: Perform super-resolution processing on the second image features to obtain the third image features; the frequency of the third image features is higher than or equal to a preset frequency threshold.
[0048] Super-resolution processing improves the resolution of an original image using hardware or software methods. It's the process of obtaining a high-resolution image from a series of low-resolution images. The super-resolution ratio can be set as needed. For example, super-resolution processing can be 2x super-resolution.
[0049] The frequency of the third image feature is higher than or equal to a preset frequency threshold, meaning the third image feature includes high-frequency image features. High frequency refers to rapidly changing frequencies. This is seen at image edges where there is a significant difference in grayscale between adjacent regions; and also in areas of sharp grayscale change where details appear. For example, the third image feature includes information such as edges and details.
[0050] Optionally, the electronic device inputs the second image features into the super-resolution imaging (SR) module, and performs super-resolution processing on the second image features to obtain the super-resolution processed third image features.
[0051] Step S106: Generate the target image based on the second image features and the third image features.
[0052] Optionally, the electronic device adds the second image features and the third image features to generate the target image.
[0053] Optionally, the electronic device multiplies the second and third image features by their respective weighting factors, and then adds them together to generate the target image. The weighting factors can be set as needed.
[0054] Optionally, the electronic device selects a first part of the features from the second image features, selects a second part of the features from the third image features, and adds the first part of the features and the second part of the features to generate the target image.
[0055] It should be noted that the method of generating the target image based on the second and third image features can be set as needed and is not limited here.
[0056] The aforementioned image processing method, based on noise information from reference frames in at least two original images, performs denoising and de-mosaic processing on first image features extracted from the at least two original images to obtain second image features with frequencies below a preset frequency threshold. Then, it performs super-resolution processing on the second image features to obtain third image features with frequencies above or equal to the preset frequency threshold. In other words, denoising and de-mosaic processing can more accurately obtain low-frequency color contours and other information of the image, while super-resolution processing can more accurately obtain high-frequency texture details and other information of the image. Therefore, based on the second and third image features, a denoised, detailed, realistic, and super-resolution target image is generated, improving the accuracy of image processing. Simultaneously, it can meet the requirements of electronic devices in terms of image processing speed, power consumption, computing power, and resolution.
[0057] Furthermore, by jointly constructing a demosaicing, denoising, and super-resolution network architecture, and outputting second image features of low-frequency components and third image features of high-frequency components, the utilization rate of each part in this architecture can be improved. This allows the super-resolution processing part to focus on restoring high-frequency texture details of the image, while the joint demosaicing and denoising part can focus on restoring low-frequency information. In addition, the high-low frequency separation strategy provides greater flexibility in terms of denoising intensity and super-resolution magnification, supporting customization for different application scenarios.
[0058] In one embodiment, based on noise information of reference frames in at least two original images, a first image feature extracted from at least two original images is subjected to denoising and de-mosaic processing to obtain a second image feature, comprising: denoising the first image feature extracted from at least two original images based on noise information of reference frames in at least two original images to obtain a denoised image feature; and de-mosaic processing the denoised image feature to obtain a second image feature.
[0059] It is understandable that the electronic device can perform noise denoising on the first image features extracted from at least two original images based on the noise information of the reference frames in at least two original images, thereby obtaining denoised image features. Then, the denoised image features can be de-mosaiced to accurately obtain denoised and de-mosaiced second image features.
[0060] In another embodiment, the electronic device may first perform de-mosaic processing on the first image features, and then perform denoising processing to obtain the second image features.
[0061] In one embodiment, based on the noise information of reference frames in at least two original images, the first image features extracted from at least two original images are denoised to obtain denoised image features, including: performing feature mapping on the noise information of reference frames in at least two original images and the first image features extracted from at least two original images to obtain mapped features; and performing denoising on the mapped features to obtain denoised image features.
[0062] Optionally, the mapped features are denoised to obtain denoised image features, including: sequentially downsampling and upsampling the mapped features to obtain upsampled features; the resolution of the upsampled features is the same as the resolution of the mapped features; and based on the upsampled features and the mapped features, denoised image features are obtained. The noise information can be a noise map.
[0063] Downsampling, also known as reduced sampling, is a multi-rate digital signal processing technique or a process of reducing the signal sampling rate, typically used to reduce data transmission rate or data size. Upsampling is also called upsampling or interpolating.
[0064] The electronic device uses a convolutional layer and an activation function layer to perform feature mapping between the noise map of the reference frame and the first image features, obtaining mapped features that the denoising module can process. These mapped features are then input into an UResNet module containing four 2x downsampling operations and four 2x upsampling operations. The mapped features are then downsampled and upsampled sequentially to obtain upsampled features. Finally, the upsampled features and the mapped features are concatenated along the channel dimension to obtain the denoised image features. The UResNet module is a UNet with residual modules as its basic processing units. The output after each upsampling operation is concatenated with the features of the same resolution before downsampling along the channel dimension before being input into the subsequent upsampling module. The number and scaling factor of downsampling and upsampling in the UResNet module can be set as needed and are not limited here.
[0065] In this embodiment, the electronic device performs feature mapping on the noise information of the reference frame in at least two original images and the first image features extracted from at least two original images to obtain mapped features. Then, the mapped features are denoised to obtain denoised image features, thereby improving the denoising effect of the image features.
[0066] In one embodiment, performing demosaic processing on the denoised image features to obtain second image features includes: upsampling and residual processing on the denoised image features to obtain second image features.
[0067] Optionally, the electronic device upsamples the denoised image features by a factor of 2, and then inputs the obtained image features into a network containing two deep residual modules for residual processing, thereby obtaining the second image features after joint demosaicing and denoising. The number of deep residual modules in the network can be set as needed.
[0068] Alternatively, the electronic device can process the second image features through two convolutional layers to obtain the first image component in the RGB domain.
[0069] In one embodiment, performing super-resolution processing on the second image features to obtain the third image features includes: performing residual processing on the second image features to obtain residual processed features; and upsampling the residual processed features to obtain the third image features.
[0070] Optionally, the electronic device inputs the second image features into a network containing a depth residual module, performs residual processing through the network containing the depth residual module to obtain residual processed features, and upsamples the residual processed features to obtain the third image features.
[0071] The number of deep residual modules in a network can be configured as needed. For example, to balance processing efficiency and processing effect, the network may include eight deep residual modules.
[0072] Optionally, the electronic device may use nearest neighbor interpolation to upsample the residual processing features to obtain the third image features.
[0073] Optionally, taking 2x super-resolution processing as an example, the electronic device performs residual processing on the second image features to obtain residual processed features; and performs 2x upsampling on the residual processed features to obtain the third image features.
[0074] For super-resolution magnification requirements greater than 1x and less than 2x, the electronic device upsamples the residual processing features by a preset magnification to obtain a third image feature; and upsamples the second image features by a preset magnification to obtain an upsampled second image feature. The preset magnification is greater than 1 and less than 2x.
[0075] For super-resolution magnification requirements greater than 2x, the electronic device connects a super-resolution processing branch with a preset magnification in parallel to obtain a third image feature with the preset magnification. At the same time, the second image feature with 2x (or other multiples of the preset magnification that are closest to the preset magnification, less than the preset magnification, or integer powers of 2) is multiplied by the corresponding upsampling factor to obtain the upsampled second image feature.
[0076] For example, for an 8x super-resolution magnification requirement, the electronic device increases the number of upsampling times to 3 during super-resolution processing and joint denoising and de-mosaic processing. That is, the first 2x upsampling results in 2x, the second 2x upsampling results in 4x, and the third 2x upsampling results in 8x.
[0077] Optionally, the electronic device can also process the third image features through two convolutional layers to obtain the second image component in the RGB domain; by adding the first image component in the RGB domain and the second image component in the RGB domain, the target image in the RGB domain can be obtained.
[0078] In this embodiment, residual processing is performed on the second image features to obtain residual processed features. Upsampling of the residual processed features can accurately obtain the super-resolution third image features.
[0079] In one embodiment, such as Figure 2As shown, another image processing method is also provided, including the following steps:
[0080] Step S202: Perform feature mapping on the noise information of the reference frame in at least two original images and the first image features extracted from at least two original images to obtain the mapped features.
[0081] Step S204: The mapped features are downsampled and upsampled sequentially to obtain upsampled features; the resolution of the upsampled features is the same as the resolution of the mapped features.
[0082] It should be noted that electronic devices can achieve noise reduction by sequentially downsampling and upsampling the mapped features.
[0083] Step S206: Based on the upsampling features and the mapping features, the denoised image features are obtained.
[0084] Step S208: Upsample and perform residual processing on the denoised image features to obtain second image features; the frequency of the second image features is lower than a preset frequency threshold.
[0085] Step S210: Perform residual processing on the second image features to obtain residual processed features; perform upsampling on the residual processed features at a preset magnification to obtain third image features; the frequency of the third image features is higher than or equal to a preset frequency threshold.
[0086] The preset magnification can be set as needed. For example, the preset magnification can be 2x.
[0087] Step S212: Upsample the second image features by a preset magnification to obtain the upsampled second image features.
[0088] It is understandable that the third image feature is based on the second image feature, and includes upsampling at a preset magnification. Therefore, the resolution of the third image feature is higher than that of the second image feature. In order to obtain a more accurate synthesized target image, the electronic device also upsamples the second image feature at a preset magnification, resulting in an upsampled second image feature; the resolution of the upsampled second image feature is the same as that of the third image feature.
[0089] Optionally, the electronic device upsamples the second image features by a preset ratio using bicubic interpolation to obtain the upsampled second image features.
[0090] Step S214: Generate the target image based on the features of the second and third images after upsampling.
[0091] In this embodiment, the electronic device performs residual processing on the second image features to obtain residual processed features, and then upsamples the residual processed features by a preset factor to obtain third image features. The second image features are also upsampled by a preset factor to obtain upsampled second image features. The resolution of the upsampled second image features and the third image features is consistent, so the target image can be generated more accurately based on the upsampled second image features and the third image features with consistent resolution.
[0092] In one embodiment, extracting a first image feature from at least two original images includes: extracting sub-image features from each of the at least two original images; and fusing the at least two sub-image features to obtain the first image feature.
[0093] Optionally, the electronic device extracts sub-image features of each of the at least two original images from multiple convolutional layers, and fuses the at least two sub-image features through convolutional layers or a self-attention mechanism to obtain the first image features.
[0094] In this embodiment, the electronic device extracts sub-image features of each of the at least two original images. It can utilize the complementary information between the at least two sub-image features to fuse the at least two sub-image features to obtain a first image feature with more image information.
[0095] In one embodiment, such as Figure 3 As shown, another image processing method is provided, which includes the following steps:
[0096] Step S302: Obtain the motion region mask for each non-reference frame in at least two original images, excluding the reference frame.
[0097] Optionally, the electronic device determines a reference frame and non-reference frames other than the reference frame from at least two original images; detects the motion region of each non-reference frame, and generates a motion region mask for each non-reference frame.
[0098] Step S304: Extract sub-image features of each original image based on the reference frame, non-reference frame and corresponding motion region mask.
[0099] Optionally, the electronic device does not need to perform motion detection on the reference frame, meaning the motion region mask of the reference frame is a completely black image. The electronic device concatenates the reference frame and its corresponding motion region mask along the channel dimension, and concatenates the non-reference frame and its corresponding motion region mask along the channel dimension, obtaining a concatenated reference frame and a concatenated non-reference frame. A feature extraction network extracts sub-image features of each original image from the concatenated reference frame and the concatenated non-reference frame. The feature extraction network may include multiple convolutional layers. For example, the original image in the RAW domain includes 4 channels. Concatenating the motion region mask and the original image in the RAW domain along the channel dimension generates a new original image in the RAW domain with 5 channels, including the original 4 channels and the motion region mask channel.
[0100] Alternatively, the electronic device can also input the reference frame and the concatenated non-reference frame into the feature extraction network for feature extraction.
[0101] It is understandable that the image and the motion region mask are concatenated in the channel dimension, that is, the motion region mask is an additional channel added to the image.
[0102] Step S306: Fuse at least two sub-image features to obtain the first image feature.
[0103] Step S308: Perform feature mapping on the noise information of the reference frame and the features of the first image in at least two original images to obtain the mapped features.
[0104] Step S310: The mapped features are downsampled and upsampled sequentially to obtain upsampled features; the resolution of the upsampled features is the same as the resolution of the mapped features.
[0105] Step S312: Based on the upsampling features and the mapping features, the denoised image features are obtained.
[0106] Step S314: Upsample and perform residual processing on the denoised image features to obtain second image features; the frequency of the second image features is lower than a preset frequency threshold.
[0107] Step S316: Perform residual processing on the second image features to obtain residual processed features.
[0108] Step S318: Upsample the residual processing features by a preset factor to obtain the third image features; the frequency of the third image features is higher than or equal to a preset frequency threshold.
[0109] Step S320: Upsample the second image features by a preset magnification to obtain the upsampled second image features.
[0110] Step S322: Generate the target image based on the features of the second and third images after upsampling.
[0111] In this embodiment, the electronic device can acquire the motion region mask of each non-reference frame in at least two original images, excluding the reference frame. Based on the reference frame, non-reference frames, and the corresponding motion region mask, the position of the motion region in each image can be determined. This allows for more accurate extraction of the required sub-image features for each original image, targeting both the motion region position and the non-motion region position, thereby improving the accuracy of image processing.
[0112] Furthermore, motion region masks are used to improve the noise reduction of moving regions. Understandably, joint denoising, de-mosaicing, and super-resolution networks tend to use multi-frame information for denoising because noise is generally zero-mean, and the more frames used, the stronger the denoising ability after averaging. However, the moving regions in each non-reference frame are the moving regions of the reference frame, meaning that the moving regions use information from one frame. Therefore, the joint denoising, de-mosaicing, and super-resolution networks implicitly improve the noise reduction of moving regions, avoiding stitching artifacts.
[0113] In one embodiment, the method for determining the noise information of the reference frame includes: determining the shot noise and readout noise corresponding to the reference frame based on the shooting parameters of the reference frame; generating a target noise map of the reference frame based on the shot noise and readout noise corresponding to the reference frame; the target noise map contains the noise information of the reference frame.
[0114] Shot noise is noise caused by non-uniform electron emission in active devices (such as vacuum tubes) in communication equipment. Read noise is electronic noise generated during the process of transferring charge from pixels out of the camera. It is a combination of all noise generated by system components when converting the charge of each pixel into a signal and then into a digital value, such as charge transfer noise, sense amplifier reset noise, analog-to-digital conversion quantization noise, and noise caused by crosstalk between the line transfer clock and the level register drive clock.
[0115] The shooting parameters of the reference frame may include ISO sensitivity, as well as information such as shutter speed or aperture value.
[0116] Optionally, the electronic device inputs the shooting parameters of the reference frame into the noise model, and outputs the shot noise and readout noise corresponding to the shooting parameters of the reference frame through the noise model. The noise model can be a Gaussian-Poisson noise model. It is understood that since different digital gain settings during shooting will affect the noise intensity, the noise model is obtained by calibrating the image sensor beforehand.
[0117] Optionally, based on the shot noise and readout noise corresponding to the reference frame, a target noise map of the reference frame is generated, including: multiplying each pixel in the reference frame by shot noise to obtain an intermediate noise map; and adding readout noise to the intermediate noise map to generate the target noise map of the reference frame.
[0118] The electronic device multiplies each pixel in the reference frame by the shot noise value to obtain an intermediate noise map; then, it adds the readout noise value to each pixel in the intermediate noise map to generate the target noise map of the reference frame. Each pixel value in the target noise map represents the noise information of the corresponding pixel in the reference frame.
[0119] In this embodiment, the electronic device determines the shot noise and readout noise corresponding to the reference frame based on the shooting parameters of the reference frame. Based on the shot noise and readout noise corresponding to the reference frame, the target noise map of the reference frame can be accurately generated, thereby obtaining the noise information of the reference frame from the target noise map.
[0120] In one embodiment, such as Figure 4 As shown, the electronic device acquires at least two raw images (Low Quality Images) and motion region masks for each raw image. Each raw image and its respective motion region mask are concatenated along the channel dimension to obtain new raw images. Sub-image features are extracted from each of the at least two new raw images. The at least two sub-image features are fused using a convolutional layer or a self-attention mechanism to obtain a first image feature. Based on the shooting parameters of a reference frame, the shot noise and readout noise corresponding to the reference frame are determined. Based on the shot noise and readout noise corresponding to the reference frame in the at least two raw images, a target noise map of the reference frame is generated. The target noise map and the first image feature are input into JDD (Joint Demosaic and...) In the Denoising (joint de-mosaicing) module, the JDD module performs denoising and de-mosaicing on the first image features based on the target noise map of the reference frame to obtain the second image features. The JDD module includes an UResNet network structure and 2x upsampling. The second image features are upsampled at a preset ratio to obtain upsampled second image features, which are low-frequency components. The second image features are input into the SR (Super-Resolution) module, which performs super-resolution processing on the second image features to obtain the third image features, which are high-frequency components. The super-resolution processing includes upsampling at a preset ratio. The upsampled second and third image features are added together to generate the target image.
[0121] In one embodiment, such as Figure 5 As shown, another image processing method is provided, which includes the following steps:
[0122] Step S502: Motion detection is performed on non-reference frames other than the reference frame in at least two original images to determine the motion region of each non-reference frame.
[0123] Optionally, the electronic device employs a motion detection algorithm to perform motion detection on non-reference frames in at least two original images, excluding the reference frame, to determine the motion region of each non-reference frame.
[0124] Optionally, the electronic device performs difference processing on the non-reference frames (excluding the reference frame) in at least two original images and the reference frame respectively to obtain a difference image corresponding to each non-reference frame; and determines the motion region of each non-reference frame from the difference image corresponding to each non-reference frame.
[0125] For each non-reference frame corresponding to the difference image, the electronic device can compare the pixel values in the difference image with a preset threshold, identify pixels with values greater than the preset threshold as moving pixels, and construct the motion region of the non-reference frame from these moving pixels. The preset threshold can be set as needed.
[0126] Optionally, in order to improve image processing efficiency, the electronic device can downsample at least two original images, perform motion detection on non-reference frames other than the reference frame in the at least two downsampled original images, and determine the motion region of each non-reference frame.
[0127] Step S504: Based on the image information of the reference frame, update the motion region of each non-reference frame to obtain the updated non-reference frame.
[0128] Optionally, the electronic device can update the motion region of each non-reference frame based on the image information of the motion region position corresponding to the non-reference frame in the reference frame, and obtain the updated non-reference frame.
[0129] In one alternative implementation, for each non-reference frame, the motion region of the non-reference frame is replaced with image information in the reference frame corresponding to the position of the motion region of the non-reference frame, to obtain an updated non-reference frame.
[0130] In another alternative implementation, for each non-reference frame, the image information of the motion region position in the reference frame corresponding to the non-reference frame is used to cover the motion region of the non-reference frame to obtain an updated non-reference frame.
[0131] Step S506: Using the reference frame and the updated non-reference frame as at least two new original images, based on the noise information of the reference frame in the at least two original images, the first image features extracted from the at least two original images are subjected to denoising and de-mosaic processing to obtain the second image features; the frequency of the second image features is lower than a preset frequency threshold.
[0132] The electronic device performs denoising and de-mosaic processing on the first image features extracted from the at least two new original images based on the noise information of the reference frames in the new at least two original images, to obtain the second image features.
[0133] Step S508: Perform super-resolution processing on the second image features to obtain the third image features; the frequency of the third image features is higher than or equal to a preset frequency threshold.
[0134] Step S510: Generate the target image based on the second image features and the third image features.
[0135] It is understandable that, since there is an interval between the shooting times of at least two original images, the subject will inevitably move on its own. When the movement is large, multi-frame registration will fail in the moving area, and problems such as motion ghosting will appear after multi-frame image fusion.
[0136] In this embodiment, the electronic device performs motion detection on non-reference frames other than the reference frame in at least two original images to determine the motion region of each non-reference frame; based on the image information of the reference frame, the motion region of each non-reference frame is updated to obtain the updated non-reference frame. This can avoid problems such as ghosting after image fusion and eliminate large displacements between multiple frames, thereby improving the accuracy of image processing.
[0137] In one embodiment, such as Figure 6 As shown, another image processing method is provided, which includes the following steps:
[0138] Step S602: Register the non-reference frames (excluding the reference frame) from at least two original images to the reference frame to obtain the registered non-reference frames.
[0139] Optionally, the electronic device registers the non-reference frame to the reference frame, which may include operations such as displacement and interpolation, as well as operations such as brightness alignment and image block alignment.
[0140] Optionally, registering the non-reference frames (excluding the reference frame) in at least two original images to the reference frame to obtain registered non-reference frames includes: determining the motion transformation relationship between the reference frame and each non-reference frame based on the reference frame and the non-reference frames in at least two original images; and registering the non-reference frames to the reference frame based on the motion transformation relationship corresponding to each non-reference frame to obtain registered non-reference frames.
[0141] The motion transformation relationship includes at least one of the affine transformation matrix and the optical flow of feature points.
[0142] Optionally, to improve image processing efficiency, the electronic device segments the reference frame and at least two non-reference frames from the original images (excluding the reference frame) into blocks, obtaining image blocks of the reference frame and image blocks of the non-reference frames from the original images (excluding the reference frame); corner detection is performed on each image block of the reference frame and each image block of the non-reference frames to obtain reference feature points of each image block of the reference frame and non-reference feature points of each image block of the non-reference frame; for each image block in each non-reference frame, the electronic device determines the motion transformation relationship between the non-reference feature points in the image block of the non-reference frame and the reference feature points of the corresponding image block in the reference frame, and registers the image block of the non-reference frame to the corresponding image block in the reference frame according to the motion transformation relationship; based on each registered image block in the non-reference frame, a registered non-reference frame is obtained.
[0143] Optionally, for each image block in each non-reference frame, the electronic device determines the affine transformation matrix or feature point optical flow between the non-reference feature points in the image block of the non-reference frame and the reference feature points of the corresponding image block in the reference frame, multiplies the image block of the non-reference frame by the affine transformation matrix or feature point optical flow, obtains the registered image block in the non-reference frame, and accurately registers it to the corresponding image block in the reference frame.
[0144] Optionally, the electronic device can perform smoothing filtering on adjacent image blocks, and then stitch the smoothed image blocks together to obtain a registered non-reference frame. In this registered non-reference frame, the transition area between adjacent image blocks is more natural.
[0145] Optionally, if there are overlapping pixels in the adjacent image blocks obtained by the electronic device during the slicing process, the overlapping pixels of the adjacent image blocks are fused to obtain a registered non-reference frame. In this registered non-reference frame, the transition area between adjacent image blocks is more natural.
[0146] Step S604: Motion detection is performed on non-reference frames other than the reference frame in at least two original images to determine the motion region of each non-reference frame.
[0147] Step S606: Based on the image information of the reference frame, update the motion region of each registered non-reference frame to obtain the updated non-reference frame.
[0148] For each non-reference frame, the electronic device replaces the motion region of the registered non-reference frame with the image information of the corresponding motion region position in the reference frame, thus obtaining the updated non-reference frame.
[0149] Step S608: Using the reference frame and the updated non-reference frame as at least two new original images, based on the noise information of the reference frame in the at least two original images, the first image features extracted from the at least two original images are subjected to denoising and de-mosaic processing to obtain the second image features; the frequency of the second image features is lower than a preset frequency threshold.
[0150] Step S610: Perform super-resolution processing on the second image features to obtain the third image features; the frequency of the third image features is higher than or equal to a preset frequency threshold.
[0151] Step S612: Generate the target image based on the second image features and the third image features.
[0152] It is understandable that during the process of an electronic device capturing at least two original images, there may be overall or local displacement between the original images. Specifically, this can include the displacement of the electronic device itself or the local displacement of the object being photographed, resulting in an overall image shift over a short period of time. Therefore, by registering the non-reference frames (excluding the reference frame) of the at least two original images to the reference frame, the electronic device can more accurately fuse multiple image features in subsequent processing, thereby improving the accuracy of image processing.
[0153] Electronic devices, with their high-precision and high-efficiency registration and alignment modules, can recover details that are difficult to obtain using traditional single-frame algorithms by leveraging sub-pixel-level complementary information across multiple frames. Furthermore, the effective utilization of multi-frame information can significantly improve image denoising capabilities. Moreover, for multi-frame input, the high-precision motion region detection module and the replacement of non-reference frame motion regions with reference frame motion regions can effectively avoid ghosting issues caused by the movement of the photographed object.
[0154] In one embodiment, determining a reference frame from at least two original images includes: determining the sharpness of each of the at least two original images; and determining a reference frame from the at least two original images based on the sharpness of each original image.
[0155] Optionally, the electronic device determines the original image with the highest sharpness from at least two original images as the reference frame.
[0156] Optionally, the electronic device determines the second-highest sharpness original image from at least two original images as the reference frame.
[0157] It should be noted that there is no limitation on the method of determining the reference frame from at least two original images.
[0158] Optionally, determining the sharpness of each of at least two original images includes: averaging the green channel in the original RAW image for each RAW domain original image to generate a grayscale image; extracting the difference of Gaussians operator from the grayscale image; and determining the sharpness of the original image based on the difference of Gaussians operator.
[0159] Optionally, for each RAW domain original image, the electronic device averages the two green channels in the RAW domain original image to generate a grayscale image; extracts the Difference of Gaussian (DoG) operator from the grayscale image; and determines the sharpness of the original image based on the Difference of Gaussian operator.
[0160] Optionally, the sharpness of the original image is determined based on the difference of Gaussians operator, including averaging the elements included in the difference of Gaussians operator to obtain the sharpness of the original image. The elements in the difference of Gaussians operator include the size and variance of the Gaussian kernel, which can be obtained through statistical analysis of the image data.
[0161] In the Gaussian difference operator, the size and variance of the Gaussian kernel are adjusted and fixed according to the data distribution. That is, the overall distribution of the size and variance of the Gaussian kernel is statistically analyzed, and a set of Gaussian kernel sizes and variances is determined based on the overall distribution of the parameters to make them conform to the overall data distribution of the parameters. The size and variance of the Gaussian kernel are then fixed.
[0162] Optionally, the electronic device can use the average value of the Gaussian difference operator as the sharpness of the original image, or it can use the average value of the Gaussian difference operator as the sharpness score of the original image, and determine the sharpness ranking of each original image according to the sharpness score. Here, the Gaussian difference operator is a matrix of image length * image width, and the sharpness score is obtained by averaging all the parameters of all elements of this matrix. The highest sharpness score corresponds to the highest sharpness of the original image.
[0163] It is understandable that, since the green channel in the original RAW image has a high sampling rate and the human eye is more sensitive to it, the green channel carries more information. Therefore, averaging the two green channels in the original RAW image can produce a grayscale image with more information, thus allowing for a more accurate determination of the sharpness of the original image.
[0164] Electronic devices can acquire raw images in the RAW domain, which can preserve the image signal received by the image sensor at the time of shooting to the greatest extent. This avoids the destruction of image detail, structural information, color and brightness information by traditional algorithms such as traditional demosaicing, noise reduction, and tone mapping. At the same time, it can avoid information loss caused by image dynamic range compression. Compared with traditional YUV or RGB domain images, it has greater tolerance and better visual effect.
[0165] In one embodiment, such as Figure 7 As shown, the electronic device captures images through the image sensor of the lens module and dumps at least two original images; the sharpness of each original image is calculated, sorted by sharpness, and the original images with low sharpness are removed to obtain the remaining original images; reference frames and non-reference frames are determined from the remaining original images; the reference frames and non-reference frames are segmented and feature points are extracted, and the optical flow or affine transformation matrix between the reference frames and non-reference frames is calculated; based on the optical flow or affine transformation matrix, the non-reference frames are registered to the reference frames; motion region detection is performed on each non-reference frame to determine the motion region mask of each non-reference frame; based on the motion region mask of each non-reference frame, the motion regions of the registered non-reference frames are replaced with pixels of the reference frames to obtain multi-frame images to be input into the network. Here, the lens module can be a telephoto module, and the network refers to a joint de-mosaicing, denoising, and super-resolution network.
[0166] In one embodiment, the method further includes: mapping a second image feature to the RGB domain and mapping a third image feature to the RGB domain; generating a target image based on the second image feature and the third image feature, including: adding the second image feature in the RGB domain and the third image feature in the RGB domain to generate the target image.
[0167] Optionally, the electronic device inputs the second image features into two convolutional layers, and maps the second image features to the RGB domain through the two convolutional layers to obtain the second image features in the RGB domain; inputs the third image features into two convolutional layers, and maps the third image features to the RGB domain through the two convolutional layers to obtain the third image features in the RGB domain; and adds the second image features in the RGB domain and the third image features in the RGB domain to generate the target image in the RGB domain.
[0168] In this embodiment, the electronic device maps both the second and third image features to the RGB domain, which can accurately generate the target image in the RGB domain.
[0169] In one embodiment, the method further includes capturing at least two raw images while locking the shooting parameters.
[0170] Optionally, shooting parameters include automatic exposure (AE), auto focus (AF), or automatic white balance (AWB). Shooting parameters also include ISO sensitivity, shutter speed, etc., but are not limited to these.
[0171] Optionally, the electronic device can continuously capture at least two raw images of the same shooting scene while locking the shooting parameters.
[0172] Understandably, electronic devices store raw images of each frame in a queue during continuous shooting, and when the user triggers the shutter, they retrieve at least two raw images from the register.
[0173] Optionally, to avoid large displacements caused by the movement of the electronic device or the object being photographed during shooting, the electronic device can control the shutter speed to be less than or equal to a preset shutter speed. The preset shutter speed can be set as needed, and it is also the maximum shutter speed of the electronic device during shooting.
[0174] In this embodiment, the electronic device captures at least two original images while locking the shooting parameters, which ensures that the captured at least two original images maintain consistency in image features such as brightness and color, thereby improving the accuracy of subsequent image processing.
[0175] In one embodiment, such as Figure 8 As shown, an electronic device captures and dumps multiple RAW images, the number of which can be N. Frames are selected from these RAW images, i.e., reference frames and non-reference frames are chosen. The electronic device calculates the sharpness of each RAW image, sorts them by sharpness, and selects the reference frame with the highest sharpness and other non-reference frames (total number less than N). The non-reference frames are registered to the reference frame to obtain a registered RAW image, which includes the reference frame and each registered non-reference frame. A motion region mask is calculated for each non-reference frame relative to the reference frame. Based on the motion region mask, the motion regions of the non-reference frames are replaced with the image information of the corresponding positions in the reference frame. Each image is then concatenated with the motion region mask and fed into a feature extraction network to extract features from each image, resulting in the first image feature. Based on a pre-calibrated sensor noise model, a target noise map of the reference frame is obtained, concatenated with the first image feature, and input into a joint denoising, de-mosaic, and super-resolution network to finally obtain a target image. This target image can then be input into a subsequent image processing engine for further processing.
[0176] In one embodiment, depending on the optimization requirements for image details, smearing effect, etc. in different scenarios, the electronic device can also support sharpening the reference frame and adjusting the noise map when the texture is adjusted, thereby controlling the denoising intensity and adding grayscale noise to the output of the JDD module.
[0177] Optionally, the electronic device sharpens the reference frame in the input Raw domain image, which can enable the joint denoising, demosaicing and super-resolution network to retain more weak textures during processing, reduce the smearing effect, and avoid artifacts such as black and white edges caused by sharpening in the RGB or YUV domains.
[0178] Optionally, the denoising strength of the joint denoising, demosaicing, and super-resolution network is influenced by the target noise map of the reference frame. By adjusting the target noise map input to the joint denoising, demosaicing, and super-resolution network during inference, the denoising strength of the network can be controlled, achieving a balance between a smooth, smeared appearance and noise.
[0179] Furthermore, the combined denoising, demosaicing, and super-resolution network can adjust the denoising intensity globally or locally on an image. For example, for images taken at night, the electronic device can control the combined denoising, demosaicing, and super-resolution network to improve the overall denoising intensity; for images taken during the day, the electronic device can control the combined denoising, demosaicing, and super-resolution network to identify dark areas of the image and improve the denoising intensity for those dark areas.
[0180] Optionally, due to the unique characteristics of human visual perception, granular grayscale noise can improve visual quality in textured areas. Electronic devices can decouple the JDD module from the SR module, enabling the folding of original image noise in the JDD module output and reducing the smearing effect.
[0181] In one embodiment, an image processing method is also provided, comprising the following steps:
[0182] Step A1: With the shooting parameters locked, capture at least two raw images.
[0183] Step A2: For each RAW domain original image, average the green channel in the RAW domain original image to generate a grayscale image; extract the difference of Gaussians operator from the grayscale image; average the elements included in the difference of Gaussians operator to obtain the sharpness of the original image.
[0184] Step A3: Determine a reference frame from at least two original images based on the sharpness of each original image.
[0185] Step A4: Based on the reference frame and at least two non-reference frames in the original images other than the reference frame, determine the motion transformation relationship between the reference frame and each non-reference frame; register the reference frame based on the motion transformation relationship corresponding to each non-reference frame to obtain the registered non-reference frame; the motion transformation relationship includes at least one of affine transformation matrix and feature point optical flow.
[0186] Step A5: Perform motion detection on non-reference frames (excluding the reference frame) in at least two original images to determine the motion region of each non-reference frame.
[0187] Step A6: For each non-reference frame, replace the motion region of the registered non-reference frame with the image information of the corresponding motion region position in the reference frame to obtain the updated non-reference frame.
[0188] Step A7: Using the reference frame and the updated non-reference frame as at least two new original images, obtain the motion region mask for each non-reference frame other than the reference frame in the at least two new original images; based on the reference frame, non-reference frame and the corresponding motion region mask, extract sub-image features for each original image.
[0189] Step A8: Fuse at least two sub-image features to obtain the first image feature.
[0190] Step A9: Based on the shooting parameters of the reference frame, determine the shot noise and readout noise corresponding to the reference frame; multiply each pixel in the reference frame by the shot noise to obtain an intermediate noise map; add the readout noise to the intermediate noise map to generate a target noise map of the reference frame; the target noise map contains the noise information of the reference frame.
[0191] Step A10: Perform feature mapping on the target noise map of the reference frame and the first image features to obtain the mapped features.
[0192] Step A11: The mapped features are downsampled and upsampled sequentially to obtain upsampled features; the resolution of the upsampled features is the same as that of the mapped features; based on the upsampled features and the mapped features, denoised image features are obtained; the denoised image features are upsampled and residual processed to obtain second image features; the frequency of the second image features is lower than a preset frequency threshold; the second image features are mapped to the RGB domain.
[0193] Step A12: Perform residual processing on the second image feature to obtain residual processed features; perform upsampling on the residual processed features at a preset ratio to obtain third image features; map the third image features to the RGB domain; the frequency of the third image feature is higher than or equal to a preset frequency threshold.
[0194] Step A13: Upsample the second image feature in the RGB domain by a preset ratio to obtain the upsampled second image feature in the RGB domain; add the upsampled second image feature in the RGB domain and the third image feature in the RGB domain to generate the target image.
[0195] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0196] Based on the same inventive concept, this application also provides an image processing apparatus for implementing the image processing method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more image processing apparatus embodiments provided below can be found in the limitations of the image processing method described above, and will not be repeated here.
[0197] In one embodiment, such as Figure 9 As shown, an image processing apparatus is provided, comprising: a first processing module 902, a second processing module 904, and an image generation module 906, wherein:
[0198] The first processing module 902 is used to perform denoising and de-mosaic processing on the first image features extracted from at least two original images based on the noise information of the reference frames in at least two original images to obtain the second image features; the frequency of the second image features is lower than a preset frequency threshold.
[0199] The second processing module 904 is used to perform super-resolution processing on the second image features to obtain the third image features; the frequency of the third image features is higher than or equal to a preset frequency threshold.
[0200] Image generation module 906 is used to generate a target image based on second image features and third image features.
[0201] The aforementioned image processing apparatus, based on noise information from reference frames in at least two original images, performs denoising and de-mosaic processing on first image features extracted from at least two original images to obtain second image features with frequencies lower than a preset frequency threshold; it then performs super-resolution processing on the second image features to obtain third image features with frequencies higher than or equal to the preset frequency threshold. In other words, denoising and de-mosaic processing can more accurately obtain information such as low-frequency color contours of the image, and super-resolution processing can more accurately obtain information such as high-frequency texture details of the image. Thus, based on the second and third image features, a denoised, de-mosaiced, and super-resolution target image is generated, improving the accuracy of image processing.
[0202] In one embodiment, the first processing module 902 is further configured to perform denoising processing on the first image features extracted from at least two original images based on the noise information of the reference frames in at least two original images to obtain denoised image features; and to perform de-mosaic processing on the denoised image features to obtain second image features.
[0203] In one embodiment, the first processing module 902 is further configured to perform feature mapping on the noise information of the reference frame in at least two original images and the first image features extracted from at least two original images to obtain mapped features; and to perform denoising processing on the mapped features to obtain denoised image features.
[0204] In one embodiment, the first processing module 902 is further configured to sequentially downsample and upsample the mapped features to obtain upsampled features; the resolution of the upsampled features is the same as the resolution of the mapped features; and based on the upsampled features and the mapped features, denoised image features are obtained.
[0205] In one embodiment, the first processing module 902 is further configured to upsample and perform residual processing on the denoised image features to obtain second image features.
[0206] In one embodiment, the second processing module 904 is further configured to perform residual processing on the second image features to obtain residual processed features; and to upsample the residual processed features to obtain third image features.
[0207] In one embodiment, the second processing module 904 is further configured to upsample the residual processing features by a preset factor to obtain a third image feature; the first processing module 902 is further configured to upsample the second image features by a preset factor to obtain an upsampled second image feature; and the image generation module 906 is further configured to generate a target image based on the upsampled second image feature and the third image feature.
[0208] In one embodiment, the first processing module 902 is further configured to extract sub-image features of each of the at least two original images; and fuse the at least two sub-image features to obtain a first image feature.
[0209] In one embodiment, the first processing module 902 is further configured to obtain the motion region mask of each non-reference frame in at least two original images, excluding the reference frame; and extract sub-image features of each original image based on the reference frame, the non-reference frame and the corresponding motion region mask.
[0210] In one embodiment, the first processing module 902 is further configured to determine the shot noise and readout noise corresponding to the reference frame based on the shooting parameters of the reference frame; generate a target noise map of the reference frame based on the shot noise and readout noise corresponding to the reference frame; the target noise map contains the noise information of the reference frame.
[0211] In one embodiment, the first processing module 902 is further configured to multiply each pixel in the reference frame by shot noise to obtain an intermediate noise map; and add readout noise to the intermediate noise map to generate a target noise map of the reference frame.
[0212] In one embodiment, the above-described apparatus further includes a motion detection module; the motion detection module is used to perform motion detection on non-reference frames other than the reference frame in at least two original images, determine the motion region of each non-reference frame; update the motion region of each non-reference frame based on the image information of the reference frame, obtain the updated non-reference frame, and use the reference frame and the updated non-reference frame as new at least two original images, wherein the first processing module 902 is used to perform denoising processing and de-mosaic processing on the first image features extracted from the at least two original images.
[0213] In one embodiment, the motion detection module is further configured to replace the motion region of each non-reference frame with image information in the reference frame corresponding to the position of the motion region of the non-reference frame, thereby obtaining an updated non-reference frame.
[0214] In one embodiment, the above apparatus further includes a registration module; the registration module is used to register at least two non-reference frames (excluding the reference frame) in the original images to the reference frame to obtain registered non-reference frames; the motion detection module is further used to update the motion region of each registered non-reference frame based on the image information of the reference frame to obtain updated non-reference frames.
[0215] In one embodiment, the registration module is further configured to determine the motion transformation relationship between the reference frame and each non-reference frame based on the reference frame and at least two non-reference frames in the original images other than the reference frame; and to register the reference frame based on the motion transformation relationship corresponding to each non-reference frame to obtain the registered non-reference frame.
[0216] In one embodiment, the upper motion transformation relation includes at least one of an affine transformation matrix and a feature point optical flow.
[0217] In one embodiment, the apparatus further includes a reference frame determination module; the reference frame determination module is used to determine the sharpness of each of the at least two original images; and to determine a reference frame from the at least two original images based on the sharpness of each original image.
[0218] In one embodiment, the reference frame determination module is further configured to, for each RAW domain original image, average the green channel in the RAW domain original image to generate a grayscale image; extract the Gaussian difference operator from the grayscale image; and determine the sharpness of the original image based on the Gaussian difference operator.
[0219] In one embodiment, the reference frame determination module is further configured to average the elements included in the Gaussian difference operator to obtain the sharpness of the original image.
[0220] In one embodiment, the first processing module 902 is further configured to map the second image features to the RGB domain, and the second processing module 904 is further configured to map the third image features to the RGB domain; the image generation module 906 is further configured to add the second image features in the RGB domain and the third image features in the RGB domain to generate a target image.
[0221] In one embodiment, the above-described apparatus is further configured as a shooting module; the shooting module is configured to capture at least two raw images while locking the shooting parameters.
[0222] Each module in the aforementioned image processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the electronic device in hardware form or independent of it, or stored in the memory of the electronic device in software form, so that the processor can call and execute the operations corresponding to each module.
[0223] In one embodiment, an electronic device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 10As shown, this electronic device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements an image processing method. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the electronic device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the electronic device, or external keyboards, touchpads, or mice, etc.
[0224] Those skilled in the art will understand that Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0225] This application also provides a computer-readable storage medium. One or more non-volatile computer-readable storage media containing computer-executable instructions, which, when executed by one or more processors, cause the processors to perform the steps of an image processing method.
[0226] This application also provides a computer program product containing instructions that, when run on a computer, cause the computer to perform an image processing method.
[0227] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0228] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0229] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0230] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. An image processing method, characterized in that, include: Based on the noise information of reference frames in at least two original images, the first image features extracted from the at least two original images are subjected to denoising and de-mosaic processing to obtain the second image features; the frequency of the second image features is lower than a preset frequency threshold. The second image feature is subjected to super-resolution processing to obtain a third image feature; the frequency of the third image feature is higher than or equal to a preset frequency threshold. A target image is generated based on the second image features and the third image features.
2. The method according to claim 1, characterized in that, The first image features extracted from the at least two original images are subjected to denoising and de-mosaic processing based on noise information from reference frames in at least two original images to obtain second image features, including: Based on the noise information of reference frames in at least two original images, the first image features extracted from the at least two original images are denoised to obtain denoised image features. The denoised image features are then subjected to de-mosaic processing to obtain the second image features.
3. The method according to claim 2, characterized in that, The first image features extracted from the at least two original images are denoised based on noise information from reference frames in at least two original images to obtain denoised image features, including: The noise information of the reference frame in at least two original images and the first image features extracted from the at least two original images are used to perform feature mapping to obtain the mapped features; The mapped features are then denoised to obtain denoised image features.
4. The method according to claim 3, characterized in that, The denoising process on the mapped features to obtain denoised image features includes: The mapped features are downsampled and upsampled sequentially to obtain upsampled features; the resolution of the upsampled features is the same as the resolution of the mapped features. Based on the upsampling features and the mapping features, the denoised image features are obtained.
5. The method according to claim 2, characterized in that, The step of performing de-mosaic processing on the denoised image features to obtain second image features includes: The denoised image features are upsampled and residual processed to obtain the second image features.
6. The method according to claim 1, characterized in that, The process of performing super-resolution processing on the second image features to obtain the third image features includes: The second image features are subjected to residual processing to obtain residual processed features; The residual processing features are upsampled to obtain the third image features.
7. The method according to claim 6, characterized in that, The upsampling of the residual processing features to obtain the third image features includes: The residual processing features are upsampled at a preset ratio to obtain the third image features; The method further includes: The second image feature is upsampled at the preset magnification to obtain the upsampled second image feature; The step of generating a target image based on the second image features and the third image features includes: The target image is generated based on the upsampled second image features and the third image features.
8. The method according to claim 1, characterized in that, Extracting first image features from the at least two original images includes: Extract sub-image features from each of the at least two original images; The first image feature is obtained by fusing at least two sub-image features.
9. The method according to claim 8, characterized in that, The extraction of sub-image features from each of the at least two original images includes: Obtain the motion region mask for each non-reference frame in the at least two original images, excluding the reference frame; Based on the reference frame, the non-reference frame, and the corresponding motion region mask, sub-image features of each original image are extracted.
10. The method according to claim 1, characterized in that, The method for determining the noise information of the reference frame includes: Based on the shooting parameters of the reference frame, determine the shot noise and readout noise corresponding to the reference frame; Based on the shot noise and readout noise corresponding to the reference frame, a target noise map of the reference frame is generated; the target noise map contains the noise information of the reference frame.
11. The method according to claim 10, characterized in that, The step of generating a target noise map of the reference frame based on the shot noise and readout noise corresponding to the reference frame includes: Each pixel in the reference frame is multiplied by the shot noise to obtain an intermediate noise map; The intermediate noise map is added to the readout noise to generate the target noise map of the reference frame.
12. The method according to claim 1, characterized in that, The method further includes: Motion detection is performed on non-reference frames other than the reference frame in the at least two original images to determine the motion region of each non-reference frame; Based on the image information of the reference frame, the motion region of each non-reference frame is updated to obtain the updated non-reference frame. The reference frame and the updated non-reference frame are then used as at least two new original images to perform the denoising and de-mosaic processing steps on the first image features extracted from the at least two original images.
13. The method according to claim 12, characterized in that, The step of updating the motion region of each non-reference frame based on the image information of the reference frame to obtain the updated non-reference frame includes: For each non-reference frame, the motion region of the non-reference frame is replaced with image information in the reference frame corresponding to the position of the motion region of the non-reference frame, to obtain the updated non-reference frame.
14. The method according to claim 12, characterized in that, The method further includes: The non-reference frames (excluding the reference frame) in the at least two original images are registered to the reference frame to obtain the registered non-reference frames. The step of updating the motion region of each non-reference frame based on the image information of the reference frame to obtain the updated non-reference frame includes: Based on the image information of the reference frame, the motion region of each registered non-reference frame is updated to obtain the updated non-reference frame.
15. The method according to claim 14, characterized in that, The step of registering the non-reference frames (excluding the reference frame) from the at least two original images to the reference frame to obtain the registered non-reference frames includes: Based on the reference frame and the non-reference frames in the at least two original images other than the reference frame, determine the motion transformation relationship between the reference frame and each non-reference frame; Based on the motion transformation relationship corresponding to each non-reference frame, the reference frame is registered to obtain the registered non-reference frame.
16. The method according to claim 15, characterized in that, The motion transformation relationship includes at least one of affine transformation matrix and feature point optical flow.
17. The method according to any one of claims 1 to 16, characterized in that, Determine the reference frame from at least two original images, including: Determine the sharpness of each of at least two original images; A reference frame is determined from at least two original images based on the sharpness of each original image.
18. The method according to claim 17, characterized in that, Determining the sharpness of each of the at least two original images includes: For each RAW domain original image, the green channel in the original RAW domain original image is averaged to generate a grayscale image; Extract the Gaussian difference operator from the grayscale image; The sharpness of the original image is determined based on the Gaussian difference operator.
19. The method according to claim 18, characterized in that, Determining the sharpness of the original image based on the Gaussian difference operator includes: The sharpness of the original image is obtained by averaging the elements included in the Gaussian difference operator.
20. The method according to any one of claims 1 to 16, characterized in that, The method further includes: Map the second image feature to the RGB domain, and map the third image feature to the RGB domain; The step of generating a target image based on the second image features and the third image features includes: The target image is generated by adding the second image feature in the RGB domain and the third image feature in the RGB domain.
21. The method according to any one of claims 1 to 16, characterized in that, The method further includes: With the shooting parameters locked, at least two raw images are captured.
22. An image processing apparatus, characterized in that, include: The first processing module is used to perform denoising and de-mosaic processing on the first image features extracted from the at least two original images based on the noise information of the reference frames in at least two original images to obtain the second image features; the frequency of the second image features is lower than a preset frequency threshold. The second processing module is used to perform super-resolution processing on the second image feature to obtain a third image feature; the frequency of the third image feature is higher than or equal to a preset frequency threshold. An image generation module is used to generate a target image based on the second image features and the third image features.
23. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the computer program is executed by the processor, the processor performs the steps of the image processing method as described in any one of claims 1 to 21.
24. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 21.
25. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 21.
Citation Information
Patent Citations
Image processing method and image processing system for super-resolution reconstruction
CN102651127A
Image enhancement method and device, electronic device and storage medium
CN109889800A