Data generation method and device, equipment, storage medium and product
By using scene metadata for degradation and noise processing in image processing, the target image pair matching the source data is generated, which solves the problems of restricted shooting scenes and incompetent ISP parameters in the prior art, and achieves efficient and robust training data generation and noise reduction improvement.
Patent Information
- Application Number
- CN202510125683.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-27
AI Technical Summary
In the prior art, shooting scenes are limited, resulting in the inability to successfully obtain training data, and the ISP parameters are not robust, affecting the effect of the noise reduction model.
By acquiring the source data image to be processed and metadata of at least one scene, degrading and noise processing are performed on the source data based on the metadata, a target image pair matching the source data image is generated, and the image content and scene performance are decoupled.
It saves labor costs required to obtain data, improves the robustness of training data, and improves the training speed and effect of noise reduction models.
Smart Images

Figure CN120047348A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technologies, and in particular, to a data generation method, apparatus, device, storage medium, and product. Background Art
[0002] With the rapid development of smartphones and short videos, people's demand for high-quality images and / or videos in daily life has become increasingly prominent. In the field of noise reduction of an Image Signal Processor (ISP), in order to pursue high-quality images and / or videos, noise reduction processing is implemented on images and / or videos based on a trained noise reduction model, and the accuracy of the noise reduction model is related to scene data. Summary of the Invention
[0003] Embodiments of the present application are expected to provide a data generation method, apparatus, device, storage medium, and product to solve the problem that training data cannot be successfully obtained due to limited shooting scenes in related technologies.
[0004] The technical solution of the present application is implemented as follows:
[0005] In a first aspect, an embodiment of the present application provides a data generation method, and the method includes:
[0006] Obtain a source data image to be processed and metadata of at least one scene;
[0007] Perform degradation processing and noise processing on the source data image to be processed based on the metadata of the at least one scene to obtain an original image pair under the at least one scene, where the original image pair includes a noisy original image and a corresponding noise-free original image;
[0008] Process the original image pair under the corresponding scene based on the metadata of the at least one scene to obtain at least one target image pair having the same data domain as the source data image to be processed.
[0009] In a second aspect, an embodiment of the present application provides a data generation apparatus, and the apparatus includes:
[0010] An obtaining module, configured to obtain a source data image to be processed and metadata of at least one scene;
[0011] A first processing module, configured to perform degradation processing and noise processing on the source data image to be processed based on the metadata of the at least one scene to obtain an original image pair under the at least one scene, where the original image pair includes a noisy original image and a corresponding noise-free original image;
[0012] A second processing module, configured to process the original image pair in the corresponding scenario based on the metadata of the at least one scenario, to obtain at least one target image pair having the same data domain as the source data image to be processed.
[0013] In a third aspect, an embodiment of the present application provides a photographing device, where the photographing device includes:
[0014] An image acquisition component, including a lens and a sensor, configured to acquire an image;
[0015] A processor, configured to execute some or all of the steps in the method described in the first aspect.
[0016] In a fourth aspect, an embodiment of the present application provides a storage medium, characterized in that the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement some or all of the steps in the method described in the first aspect.
[0017] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a processor, some or all of the steps in the method described in the first aspect are implemented.
[0018] A data generation method, apparatus, device, storage medium and product provided by an embodiment of the present application obtain a source data image to be processed and metadata of at least one scenario; perform degradation processing and noise processing on the source data image to be processed based on the metadata of the at least one scenario, to obtain an original image pair in at least one scenario, where the original image pair includes a noisy original image and a corresponding noise-free original image; process the original image pair in the corresponding scenario based on the metadata of the at least one scenario, to obtain at least one target image pair having the same data domain as the source data image to be processed. In this way, the present application decouples the process of obtaining the target image pair, and the source data image provides the image content, and the specific performance of the image content in different scenarios is obtained by performing image processing on the metadata in different scenarios. In this way, the labor cost required for obtaining data is saved, and the robustness of the training data is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is an optional flowchart of the data generation method provided by an embodiment of the present application;
[0020] Figure 2 is an optional flowchart block diagram of the data generation method provided by an embodiment of the present application;
[0021] Figure 3 is an optional hardware schematic diagram of the ISP provided by an embodiment of the present application;
[0022] Figure 4 An optional flowchart of the data generation method provided by the embodiment of the present application;
[0023] Figure 5 An optional flowchart of the data generation method provided by the embodiment of the present application;
[0024] Figure 6 An optional flowchart of the data generation method provided by the embodiment of the present application;
[0025] Figure 7 An optional structural diagram of the data generation device provided by the embodiment of the present application;
[0026] Figure 8 An optional structural diagram of the shooting device provided by the embodiment of the present application. Detailed implementation manners
[0027] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application.
[0028] It should be understood that the "embodiments of the present application" or "the foregoing embodiments" mentioned throughout the specification mean that specific features, structures, or characteristics related to the embodiments are included in at least one embodiment of the present application. Therefore, the "in the embodiments of the present application" or "in the foregoing embodiments" that appear throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner. In various embodiments of the present application, the magnitudes of the serial numbers of the foregoing processes do not mean the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages and disadvantages of the embodiments.
[0029] With the rapid development of smart phones and short videos, people's demand for high-quality images and / or videos in daily life has become increasingly prominent. In order to pursue high-quality images and / or videos, more and more manufacturers adopt a deep learning-based noise reduction model in the noise reduction scheme of the ISP. In order to train the noise reduction model, it is necessary to obtain training data including the image to be noise-reduced and the noise-reduced image, and the quality of the training data plays a decisive role in training the noise reduction model.
[0030] Through retrieval, it is found that the methods for obtaining training data in the related art have the following problems:
[0031] In a solution of the related art, a method of stacking frames is adopted. Specifically, in each scenario, the camera parameters are fixed to collect multiple frames of images, and the average of the collected multiple frames of images is taken to obtain a denoised image, and a frame is randomly selected from the collected multiple frames as the image to be denoised. However, the shooting scenario of this method is limited, and it is necessary to keep the scenario and the camera as stationary as possible during the shooting process. In an outdoor environment, the objects in the scenario are easily affected by natural factors (such as wind blowing, animal activities, etc.) or human factors (such as people walking, vehicle driving, etc.), and are extremely likely to change. This change of objects in the scenario will seriously interfere with the frame stacking effect and may ultimately lead to frame stacking failure; secondly, this method needs to shoot multiple frames of images in each scenario. For the diversity of scenarios, more than 100 scenarios may need to be shot, and the required time and labor costs are relatively high; finally, the ISP parameters before the target position on the ISP are not robust. For example, the stacked frames are the color mode (red green blue, RGB) images finally output by the ISP. Once some parameters in the ISP that have a greater impact on the noise form change, the stacked frames collected using the previous ISP parameters may need to face the problem of re-collection, otherwise it will lead to the mismatch between the image data used for training and the actual image data, thus affecting the effect of the denoising model.
[0032] In another solution of the related art, multiple first Raw images with different Bayer arrays in the same scenario are obtained, and the obtained multiple first Raw images are respectively subjected to fusion processing to obtain the corresponding number of second Raw images, and the corresponding first Raw images and second Raw images are constructed into sample data for machine learning to train a denoising model. First, for the diversity of scenarios, this method still needs to collect multiple Raw images with different Bayer arrays for each scenario, which is time-consuming and laborious; secondly, it is necessary to fuse multiple Raw images with different Bayer arrays collected in different scenarios. Although this method provides a solution to deal with the situation where the inconsistency of the same position in the picture of multiple images may lead to abnormal fusion results, there are still situations that this solution cannot handle, resulting in fusion failure and abnormal training data obtained, thus reducing the denoising performance of the denoising model.
[0033] In view of the above one or more problems, the exemplary embodiments of the present application first provide a data generation method for generating data pairs that can be used for training a denoising model, and training a denoising model deployed in different data domains of the ISP, so as to achieve efficient and high-quality image denoising processing.
[0034] Referring to Figure 1 , Figure 1 A data generation method provided by an embodiment of the present application is applied to a shooting device, and the method includes the following steps:
[0035] Step 101: Obtain the source data image to be processed and the metadata of at least one scenario.
[0036] In the embodiments of the present application, the source data image to be processed can be a data image that characterizes the image content and can provide basic information for subsequent image processing operations. Among them, the source data image to be processed is a data image with a noise value lower than a preset noise threshold. The source data image to be processed can be a single-frame image or a continuous multi-frame video image.
[0037] It can be understood that a source data image with a noise value lower than the preset noise threshold indicates that the image is relatively clear, with fewer interference factors, and can ensure the processing effect and accuracy in subsequent image processing, such as degradation simulation, application of noise reduction models, etc.
[0038] It can be understood that the image content of the source data image to be processed should cover as much as possible the content generated when using the noise reduction model. Because the noise reduction model will perform specific processing on the image and generate corresponding results during actual operation, and since the source data image contains the same or similar image content, the noise reduction model can better adapt to and process the data, thereby improving the noise reduction effect and optimizing the effectiveness of model training and application.
[0039] In practical applications, the source data image to be processed can be a data image with a noise value lower than the preset noise threshold captured by any image sensor or the target sensor of a shooting device; the source data image to be processed can also be an artificial intelligence (AI) data image with a noise value lower than the preset noise threshold created through software or work; the source data image to be processed can also be a publicly copyright-free data image set with a noise value lower than the preset noise threshold from the network. It should be noted that the number of source data images to be processed is at least greater than 2. In this way, the data set of source data images can contain rich and diverse data images, providing a wide range of data resources for research and processing.
[0040] In the embodiments of the present application, the source data image to be processed can be image data in any data domain, and any data domain includes but is not limited to image data in the raw (RAW) domain, image data in the RGB domain, and image data in the color encoding method (YUV) domain.
[0041] In some embodiments, the source data image to be processed can be a data image obtained after being processed by any data domain ISP processing module in the ISP. It should be noted that the source data image to be processed is associated with the location of the ISP where the noise reduction model to be trained is located.
[0042] Exemplarily, the source data image to be processed can be a data image obtained after being processed by one or more of the ISP processing modules in the RAW domain of the ISP, such as a Black Level correction (BLC) module, a Lens shading correction (LSC) module, a White Balance correction (WBC) module, or a Demosaic (DM) processing module.
[0043] Of course, the source data image to be processed can also be a data image obtained after being processed by one or more of the ISP processing modules in the RGB domain of the ISP, such as a Color Correction Matrix (CCM) module, a Global tone mapping (GTM) module / Local tone mapping (LTM), a GAMMA, or an RGB2YUV processing module.
[0044] Again, of course, the source data image to be processed can also be a data image obtained after being processed by one or more of the ISP processing modules in the YUV domain of the ISP, such as a Color Aberration Correction (CAC) module, an Edge Enhancement (EE) processing module, and a Hue & Saturation processing module.
[0045] In the embodiments of the present application, there can be included at least one scenario, and the scenario can be set based on at least one factor among a portrait, the surrounding environment, and time. It should be noted that the at least one scenario should preferably cover the scenarios required for the noise reduction model.
[0046] It can be understood that the portrait factor can distinguish whether it is a portrait scenario; the surrounding environment can distinguish whether it is an indoor scenario or an outdoor scenario, and can also distinguish landscape scenarios, architectural scenarios, animal scenarios, etc. Of course, it can also be divided into low-light scenarios, ordinary scenarios, strong wind scenarios, rain and fog scenarios, and strong light scenarios; the time factor can distinguish whether it is a daytime scenario or a nighttime scenario.
[0047] Exemplarily, the at least one scenario can include an indoor scenario, an outdoor scenario, a daytime scenario, a nighttime scenario, a portrait scenario, and a non-portrait scenario, and can also include an indoor daytime scenario, an outdoor daytime scenario, an outdoor nighttime scenario, an indoor nighttime scenario, a daytime portrait scenario, an indoor daytime portrait scenario, etc.
[0048] Based on the portrait, the surrounding environment, and the time, eight scenarios can be set, namely, indoor daytime portrait scenario, indoor daytime non-portrait scenario, indoor nighttime portrait scenario, indoor nighttime non-portrait scenario, outdoor daytime portrait scenario, outdoor daytime non-portrait scenario, outdoor nighttime portrait scenario, and outdoor nighttime non-portrait scenario. It should be noted that the above scenarios are only illustrative examples, and the present application does not specifically limit the scenarios.
[0049] In the embodiments of the present application, the metadata of at least one scenario can be understood as that one scenario corresponds to a set of metadata, and at least one scenario corresponds to at least one metadata. Among them, the metadata includes at least one image parameter for describing the scenario, and each image parameter in the at least one image parameter is provided with an image parameter value in this scenario, and the at least one image parameter can indicate the image effect presented by the final image.
[0050] In the embodiments of the present application, the image parameters can include one or more ISP parameters. For example, the ISP parameters can be BLC parameter, LSC parameter, WBC parameter, DM parameter, CCM parameter, GTM parameter, GAMMA parameter, RGB2YUV parameter, CAC parameter, EE parameter, and Hue&Saturation parameter. The image parameters can also include the brightness parameters for obtaining the RAW domain scene image (i.e., the unprocessed image) in this scenario. For example, the brightness parameters can be the average brightness parameter and / or the maximum brightness parameter of the RAW domain scene image. Of course, in the context of image processing, the metadata can also include information such as the camera model, shooting time, and exposure parameters (aperture, shutter speed, ISO). It should be noted that the image parameters included in the metadata are associated with the position of the ISP where the noise reduction model to be trained is located.
[0051] It should be noted that the types and quantities of image parameters in the metadata corresponding to different scenarios are not exactly the same, and / or the image parameter values of each image parameter are different. Exemplarily, different scenarios include Scenario A and Scenario B. In one case, the metadata corresponding to Scenario A includes image parameters such as LSC parameter, WBC parameter, CCM parameter, and GAMMA parameter, and the metadata corresponding to Scenario B includes image parameters such as LSC parameter, WBC parameter, CCM parameter, and GTM parameter. It can be seen that the types of image parameters in the metadata corresponding to Scenario A and the types of image parameters in the metadata corresponding to Scenario B are not exactly the same, and the image parameter values of the image parameters corresponding to different scenarios are also different. In the second case, the metadata corresponding to Scenario A includes image parameters such as LSC parameter, WBC parameter, CCM parameter, and GAMMA parameter, and the metadata corresponding to Scenario B includes image parameters such as LSC parameter, WBC parameter, CCM parameter, and GAMMA parameter. It can be seen that the types of image parameters in the metadata corresponding to Scenario A and the types of image parameters in the metadata corresponding to Scenario B are exactly the same. However, the image parameter values of the image parameters corresponding to different scenarios are completely different. It should be noted that the above is only an example, and the present application does not make specific limitations on this.
[0052] Step 102: Degrade and add noise to the source data image to be processed based on the metadata of at least one scenario, to obtain at least one pair of original images in at least one scenario, where the pair of original images includes a noisy original image and the corresponding noise-free original image.
[0053] In the embodiment of the present application, the noise-free original image may be a RAW image that has not been processed and has no noise added after degrading the source data image to be processed based on the metadata. The noisy original image may be a noisy RAW image obtained by adding noise to the noise-free original image.
[0054] In the embodiment of the present application, when degrading the source data image to be processed based on the metadata of at least one scenario, the source data image to be processed may be degraded respectively based on the image parameters in the metadata of at least one scenario. Here, the degradation process performed on the source data image to be processed includes performing inverse ISP processing corresponding to the image parameters in the metadata of at least one scenario on the source data image to be processed, including but not limited to inverse processing such as BLC, LSC, WBC, DM, CCM, GTM, GAMMA, RGB2YUV, CAC, EE, and Hue & Saturation.
[0055] In the embodiment of the present application, refer to Figure 2As shown, after the imaging device obtains the source data image to be processed and the metadata of at least one scene, based on the image parameters in the metadata of each scene, the source data image to be processed can be degraded by a data degenerate respectively, and then noise processing is performed to obtain a pair of original images including a noise-free original image and a corresponding noisy original image under at least one scene, and further obtain a pair of original images under at least one scene before image simulation in an input image signal processor (ISP).
[0056] Step 103: Process the pair of original images under the corresponding scene based on the metadata of at least one scene to obtain at least one pair of target images having the same data domain as the source data image to be processed.
[0057] In the embodiment of the present application, when processing the pair of original images under the corresponding scene based on the metadata of at least one scene, the pair of original images can be processed based on at least some of the image parameters in the metadata of at least one scene. Here, processing the pair of original images includes performing forward ISP processing corresponding to at least some of the image parameters on the pair of original images, including but not limited to BLC, LSC, WBC, DM, CCM, GTM, GAMMA, RGB2YUV, CAC, EE, and Hue&Saturation and other processing.
[0058] It can be understood that the pair of original images is RAW image data, and after image processing, the RAW image data will be processed into data that meets certain standards. For example, it is processed into standard image data that meets the RGB domain, or processed into standard image data that meets the YUV domain. The pair of target images is data that meets the aforementioned standards, for example, it is an RGB image or a YUV image.
[0059] In the embodiment of the present application, the pair of target images includes a target noisy image and a target denoised image. The pair of target images has the same data domain as the source data image to be processed, or the image attributes of the pair of target images are the same as those of the source data image to be processed; here, the same image attributes or data domain means that the pair of target images and the source data image to be processed both meet the standards of the same data domain.
[0060] In the embodiment of the present application, the pair of target images can be used as training data to train a denoising model with the same data domain as the source data image to be processed.
[0061] In some embodiments, the method further includes: obtaining a noise reduction model to be trained; wherein, the data domain of the noise reduction model in the ISP is the same as that of the target image pair; training the noise reduction model to be trained based on the target image pair to obtain a trained noise reduction model. In the embodiments of the present application, the generated training data, i.e., the target image pair, is used to train the noise reduction model to be trained whose data domain in the ISP is the same as that of the target image pair, so as to obtain a trained noise reduction model. In this way, the training speed and noise reduction effect of the noise reduction model are improved.
[0062] In some embodiments, the processing of the original image pair can be an image signal processor (ISP). The ISP can be an analog simulator corresponding to the image signal processor. The hardware analog simulator can use the same parameter combination as the hardware, and it can take effect without restarting. Moreover, the final image effect is equivalent to that generated by the hardware, so it does not affect the data quality. In addition, since no restart operation is required, the training duration can be reduced and the training efficiency can be improved during the training stage of the noise reduction model.
[0063] In an implementable scenario, a schematic diagram of an optional hardware of the image signal processor (ISP) is as Figure 3 shown. The ISP includes multiple ISP processing modules. Among them, the multiple ISP processing modules are distributed according to an implementable ISP pipeline. In sequence, they can be the BLC module, the LSC module, the WBC module, the DM module, the CCM module, the GTM module, the GAMMA module, the RGB2YUV module, the CAC module, the EE module, and the Hue&Saturation module.
[0064] Regarding the above BLC module, the black level refers to the signal level output by the image sensor when there is no light. The BLC module is used to adjust the black level, aiming to correct the output offset of the image sensor under the condition of no light. The BLC module can ensure that the black part of the image truly reaches black, avoiding the dark part of the image being grayish, which affects the contrast and color accuracy of the image, or the area that should be black becoming darker, resulting in the loss of dark part details of the image.
[0065] Regarding the above LSC module, due to the optical characteristics of the lens, when light passes through the lens and reaches the image sensor, there will be a situation where the center is bright and the edge is dark, that is, lens shading. The LSC module is used to compensate for this uneven illumination by enhancing the pixel values of the darker areas at the edges of the image, making the brightness of the entire picture more uniform.
[0066] For the above-mentioned WBC module, under different lighting conditions, white objects may appear color-shifted in the image. The WBC module is used to adjust the gains of the red, green, and blue channels of the image, so that white objects can be correctly displayed as white in any lighting environment, thereby restoring the true sense of the color of the entire image.
[0067] For the above-mentioned DM module, most image sensors use color filter arrays such as the Bayer array, and each pixel only records the information of one color (red, green, or blue). The DM module is used in the demosaicing process to interpolate and calculate the missing color components of each pixel through the color information of surrounding pixels, and restore each pixel to a complete color pixel, thereby obtaining a full-color image.
[0068] For the above-mentioned CCM module, different image sensors and optical systems have different response characteristics to colors. The CCM module is used to adjust the color space of the image by using a color correction matrix (CCM) to make the image colors conform to the RGB standard or the expected color style.
[0069] The above-mentioned GTM module is used to adjust the overall tone range of the image by controlling the intensity of tone mapping, so that the distribution of highlights, midtones, and shadows in the image is more in line with the visual perception of the human eye or specific display requirements.
[0070] For the above-mentioned GAMMA module, the human eye's perception of brightness is non-linear, while the signal output by the image sensor has a linear relationship with the actual brightness. The GAMMA module is used to perform a non-linear transformation on the image pixel values to make the brightness change of the image more in line with the visual perception of the human eye.
[0071] For the above-mentioned RGB2YUV module, RGB is a color space used to represent image colors, and YUV is also a color space, where Y represents luminance information, and U and V represent chrominance information. The RGB2YUV module can convert the image from the RGB data domain (also known as the color space) to the YUV data domain (or color space). Through the YUV data domain, the luminance and chrominance information can be better separated, which is convenient for compression and processing, thereby reducing the data volume while ensuring the image quality.
[0072] The above-mentioned CAC module is used to correct color abnormalities in the image, that is, color deviations caused by reasons such as sensor failures and abnormal lighting. By adjusting the color abnormality gain, the color can be restored to normal.
[0073] For the above-mentioned EE module, during the imaging process, due to factors such as the optical characteristics of the lens and sensor noise, the edges of the image may appear blurred. The EE module is used to detect the differences in brightness, color, etc. between the edge parts and other areas of the objects in the image, and process the edge pixels to highlight the edge information.
[0074] The above Hue & Saturation module is used to adjust the colors of an image, adjust the types and / or saturation of the colors in the image, and change the color style of the image, thereby enhancing or weakening the color expressiveness of the image.
[0075] It should be noted that Figure 3 This is only an example. In the actual operation process, the embodiments of the present application are not limited to Figure 3 the form of an ISP pipeline shown. Instead, it can be adapted to the ISP pipelines of different manufacturers. As long as an emulator for the corresponding ISP pipeline and the metadata required for simulation can be obtained, even if some modules cannot obtain the corresponding metadata, as long as the approximate value range of the input of its ISP module is roughly known, it can be roughly solved by adding corresponding random perturbations during the degradation process.
[0076] In the embodiments of the present application, continuing to refer to Figure 2 after the imaging device obtains at least one pair of original images in at least one scenario, it can perform data simulation processing in the direction indicated by the solid arrow in Figure 3 , perform forward ISP image simulation processing on the pair of original images corresponding to the metadata, obtain at least one pair of target noisy images and corresponding target noise-free images having the same data domain as the source data image to be processed, and determine the pair of target noisy images and corresponding target noise-free images as a pair of target images, so as to train a denoising model having the same data domain as the source data image to be processed.
[0077] A data generation method provided by the embodiments of the present application includes: obtaining a source data image to be processed and metadata of at least one scenario; performing degradation processing and noise processing on the source data image to be processed based on the metadata of at least one scenario to obtain at least one pair of original images in at least one scenario, where the pair of original images includes a noisy original image and a corresponding noise-free original image; processing the pair of original images in the corresponding scenario based on the metadata of at least one scenario to obtain at least one pair of target images having the same data domain as the source data image to be processed. In this way, the present application decouples the process of obtaining the pair of target images. The source data image provides the image content, and the specific performance of this image content in different scenarios is obtained by performing image processing on the metadata in different scenarios. In this way, the labor cost required to obtain data is saved, and the robustness of the training data is improved.
[0078] Referring to Figure 4 shown, an embodiment of the present application provides a data generation method, which is applied to an imaging device. The method includes the following steps:
[0079] Step 401, obtain a source data image to be processed.
[0080] Step 402: Collect a first quantity of first scene images through a target sensor in at least one scene.
[0081] In the embodiments of the present application, the target sensor is a specific device component for acquiring image data. The target sensor may be an image sensor in a photographing device that requires noise reduction processing. Exemplarily, in an intelligent movable terminal device, the target sensor may be a Complementary Metal Oxide Semiconductor Transistor (CMOS) image sensor; in a professional camera device, the target sensor may be a Charge Coupled Device (CCD) image sensor, etc. The target sensor is a key component for collecting images, and its performance, such as the number of pixels, sensitivity, color reproduction ability, etc., will directly affect the quality of the collected images.
[0082] In the embodiments of the present application, the first scene image is an original image obtained by converting the captured light source signal into a digital signal by the target sensor in different scenes, and the image data obtained after the original image is processed by ISP. The first scene image may be a scene image in the YUV or RGB domain, and the original image may be an image in the RAW domain. Therefore, the first scene image and the original image have different data domains or different image attributes.
[0083] In the embodiments of the present application, the first quantity is the quantity of first scene images that can be collected, and the first quantity may be 1. That is to say, in the case of multiple scenes, for each scene, the target sensor only needs to collect one first scene image, avoiding collecting multiple scene images for each scene, reducing time and lowering labor costs.
[0084] Step 403: Perform statistical analysis on the image parameters in the first scene images to obtain the image signal processing (ISP) parameters of at least one scene and the brightness parameters of the first scene images in the original domain.
[0085] Step 404: Use the ISP parameters and brightness parameters corresponding to at least one scene as the metadata of at least one scene.
[0086] In the embodiments of the present application, the image parameters may include one or more ISP parameters, and the image parameters may also include the brightness parameters for obtaining the RAW domain scene image (i.e., the unprocessed image) in this scene. Of course, in the context of image processing, the metadata may also include information such as the camera model, shooting time, exposure parameters (aperture, shutter speed, sensitivity), etc.
[0087] In the embodiments of the present application, after the imaging device acquires a first number of first scene images through a target sensor in at least one scene, the imaging device respectively performs statistical analysis on the image parameters in the first scene images to obtain one or more ISP parameters in at least one scene, and the brightness parameter of the first scene images in the RAW domain in this scene; further, the one or more ISP parameters and the corresponding brightness parameter are used as the metadata corresponding to this scene, and thus, the metadata of at least one scene is obtained.
[0088] Step 405: Based on the ISP parameters and the brightness parameters in the metadata of at least one scene, perform degradation processing on the source data image to be processed to obtain a noise-free raw image in at least one scene.
[0089] In the embodiments of the present application, continue to refer to Figure 3 , and based on the ISP parameters and the brightness parameters in the metadata of at least one scene, in the direction indicated by the dotted arrow in the ISP pipeline shown in Figure 3 , perform degradation processing on the source data image to be processed to obtain a noise-free raw image.
[0090] In an implementable scenario, continue to refer to Figure 2 and Figure 3 shown. Taking the image data obtained after the source data image to be processed passes through the CCM module in the RGB domain of the ISP as an example, at least one scene includes a first scene and a second scene. The metadata A of the first scene includes LSC parameters, WBC parameters, DM parameters, and CCM parameters, as well as brightness parameters. The metadata B of the second scene includes LSC parameters, WBC parameters, and CCM parameters, as well as brightness parameters.
[0091] Based on the LSC parameters, WBC parameters, DM parameters, and CCM parameters, as well as the brightness parameters, in the metadata 202 of the first scene, that is, metadata A, through the data degenerator 203, based on the ISP pipeline of the corresponding platform, the source data image 201 to be processed, whose current data domain is the CCM module in the RGB domain, is reversed along the direction of the ISP pipeline, passes through the CCM module corresponding to the CCM parameters, the DM module corresponding to the DM parameters, the WBC module corresponding to the WBC parameters, and the LSC module corresponding to the LSC parameters, and returns to the RAW image data in the RAW domain, and the RAW image data is processed based on the brightness parameters to obtain the noise-free RAW image corresponding to the first scene. It should be noted that since the BLC parameter is not included in the metadata, the BLC module and other ISP processing modules without ISP parameters in the data degenerator 203 can be set to the off state.
[0092] Similarly, based on the LSC parameters, WBC parameters, CCM parameters, and luminance parameters in the metadata 202 of the second scenario, i.e., metadata B, through the data degenerator 203, based on the ISP pipeline of the corresponding platform, the source data image 201 to be processed is reversed from the CCM module in the RGB domain where it currently resides along the direction of the ISP pipeline, passing through the CCM module corresponding to the CCM parameters, the WBC module corresponding to the WBC parameters, and the LSC module corresponding to the LSC parameters, and is backed off to the RAW image data in the RAW domain, and the RAW image data is processed based on the luminance parameters to obtain the noise-free RAW image corresponding to the second scenario. It should be noted that since the source data image to be processed is the image data obtained after being processed by the CCM module in the RGB domain of the ISP, the BLC module and other ISP processing modules without ISP parameters in the data degenerator 203 can be set to the off state.
[0093] It should be noted that most of the ISP processing modules in the ISP pipeline are reversible and can directly perform inverse operations based on the metadata. If there are ISP processing modules that are irreparable in principle, calculate the approximate value range of their influence based on the metadata and add perturbations to the data to be degraded.
[0094] Step 406: Determine the scene noise parameters of at least one scene of the target sensor based on the scene simulation gain of at least one scene.
[0095] Among them, the metadata further includes the scene simulation gain when the target sensor acquires the first scene image. The metadata of at least one scene includes the scene simulation gain of at least one scene.
[0096] In the embodiments of the present application, the simulation gain is the value of one or more parameters such as the exposure time, temperature, and light intensity of the image sensor when acquiring the scene image. In the same or different scenes, the simulation gain can be the same or different. In the same scene, the simulation gain of the image sensor is different; or in different scenes, the simulation gain of the image sensor can be the same or different.
[0097] In the embodiments of the present application, the scene simulation gain is the simulation gain set by the target sensor when acquiring the first scene image.
[0098] In the embodiments of the present application, the scene noise parameter is the RAW domain noise generated by the target sensor with the corresponding scene simulation gain set in at least one scene.
[0099] In the embodiments of the present application, based on the scene simulation gain of each scene in at least one scene, the scene noise parameter of the target sensor in this scene is determined, so as to obtain the scene noise parameters of the target sensor in at least one scene.
[0100] It can be understood that step 406 determines the scene noise parameters of at least one scene of the target sensor based on the scene simulation gain of at least one scene, which can be implemented through the following steps:
[0101] Step 461: Obtain at least two second scene images of the object to be photographed in the same scene collected by the target sensor at at least two calibration simulation gains.
[0102] In the embodiments of the present application, the second scene image is an image obtained by the target sensor collecting the object to be photographed in the same scene under the calibration simulation gain. Since the target sensor collects at least at two calibration simulation gains, correspondingly, at least two second scene images are obtained. It should be noted that multiple calibration simulation gains can be set according to actual needs.
[0103] Step 462: Based on the pixel values of each pixel point in at least two second scene images, calibrate the original domain noise of the target sensor to obtain the calibration noise parameters corresponding to the target sensor at at least two calibration simulation gains.
[0104] In the embodiments of the present application, an image is composed of numerous pixel points, and each pixel point has a corresponding pixel value. The pixel value of each pixel point in the second scene image reflects the intensity of the optical signal received by the pixel point, and after being subjected to photoelectric conversion and signal processing by the target sensor, it is stored in digital form. In the second scene images collected at different simulation gains, the pixel values of each pixel point are different due to the change of the simulation gain, and at the same time, they also include the influence of noise on the pixel values.
[0105] In the embodiments of the present application, the original domain noise can be understood as the noise introduced during the process of the target sensor converting the optical signal into an electrical signal and performing preliminary digital processing. This noise will be superimposed on the pixel values of the image, affecting the quality and accuracy of the image.
[0106] In the embodiments of the present application, the calibration of the original domain noise of the target sensor can be understood as establishing a mathematical relationship between the noise and the pixel values by analyzing and processing the pixel values of each pixel point in at least two second scene images, so as to determine the noise characteristics and parameters of the target sensor in the original domain.
[0107] In the embodiments of the present application, based on the pixel values of each pixel point in at least two second scene images, after calibrating the original domain noise of the target sensor, the parameters of the noise characteristics of the target sensor at different calibration simulation gains are obtained, that is, the calibration noise parameters. The calibration noise parameters can include one or more of the mean, variance, standard deviation, noise power spectral density, etc. of the noise, and can quantitatively reflect the noise level, noise distribution, and noise frequency characteristics of the target sensor at different simulation gains.
[0108] Step 463: Determine the calibration noise model of the target sensor based on at least two calibrated analog gains and the corresponding calibrated noise parameters;
[0109] Step 464: Input the scene analog gain of at least one scene into the calibration noise model to obtain the scene noise parameters of the target sensor in at least one scene.
[0110] In the embodiments of the present application, after obtaining the calibrated noise parameters corresponding to the target sensor at at least two calibrated analog gains, the calibration noise model of the target sensor is determined based on at least two calibrated analog gains and the corresponding calibrated noise parameters; further, the scene analog gains included in the metadata of each scene are respectively input into the calibration noise model of the target sensor to obtain the scene noise parameters corresponding to the scene analog gains of each scene, so as to obtain the scene noise parameters of at least one scene.
[0111] In an implementable scenario, a commonly used noise model is the Poisson-Gaussian noise model, as shown in formula (1):
[0112] I = Poisson(x)·g + Normal(0,y) (1)
[0113] E(I) = x·g (2)
[0114] σ(I) 2 = x·g 2 + y 2 = E(I)·g + y 2 (3)
[0115] Wherein, I is the pixel value of each pixel point in the second scene image, x is the light signal intensity received by each pixel point, Poisson(x) is the Poisson term, and Normal(0,y) is the Gaussian term.
[0116] It can be seen from formulas (1) to (3) that when the calibrated analog gain g is fixed, calculate the variance σ(I) of the pixels obtained after shooting the same pixel value on the photographed object in the second scene image 2 and the mean value E(I), and determine the slope of the curve formed by the variance σ(I) 2 and the mean value E(I) as the calibrated analog gain for calibrating the target sensor, so as to obtain the Poisson term part Poisson(x), and the Gaussian term Normal(0,y) can be obtained by setting the mean value to 0, that is, shooting a pure black image and performing calibration. In this way, a calibrated noise model is obtained, so as to obtain the scene noise parameters of the target sensor under the scene analog gains corresponding to each scene. In this way, it can be more adapted to the target sensor and improve the accuracy of training data.
[0117] Step 407: Add noise to the noiseless original image corresponding to at least one scenario according to the scenario noise parameters of the at least one scenario, to obtain the noisy original images of the at least one scenario;
[0118] Step 408: Use the noiseless original image and the noisy original image as the original image pairs for the at least one scenario.
[0119] In the embodiments of the present application, after obtaining the scenario noise parameters corresponding to at least one scenario, add noise to the noiseless original image corresponding to each scenario according to the scenario noise parameters of the at least one scenario, to obtain the noisy original images corresponding to each scenario, so as to obtain the noisy original images of the at least one scenario; further, use the noiseless original image and the noisy original image in the same scenario as the original image pair for this scenario, so as to obtain the original image pairs of the at least one scenario.
[0120] Step 409: Based on the ISP parameters in the metadata of at least one scenario, perform ISP processing on the original image pairs corresponding to the at least one scenario, to obtain at least one pair of target images having the same data domain as the source data image to be processed.
[0121] In the embodiments of the present application, continue to refer to Figure 3 , and based on the ISP parameters in the metadata of at least one scenario, in the direction indicated by the solid arrow in the ISP pipeline shown in Figure 3 , perform ISP processing on the original image pairs corresponding to the at least one scenario respectively, to obtain a pair of target images including a target noisy image and a target denoised image. Among them, the pair of target images has the same data domain as the image to be processed.
[0122] In a realizable scenario, continue to refer to Figure 2 and Figure 3 shown, and the above example is used for illustration. After obtaining the pair of RAW images corresponding to metadata A and the pair of RAW images corresponding to metadata B,
[0123] Based on the LSC parameter, WBC parameter, DM parameter, and CCM parameter in metadata 202 of the first scenario, that is, metadata A, through the ISP emulator 205, based on the ISP pipeline of the corresponding platform, for the pair of RAW images corresponding to the first scenario, along the direction of the ISP pipeline, perform forward ISP image simulation processing through the LSC module corresponding to the LSC parameter, the WBC module corresponding to the WBC parameter, the DM module corresponding to the DM parameter, and the CCM module corresponding to the CCM parameter, to obtain a pair of target images including a target noisy image and a target denoised image corresponding to the first scenario. It should be noted that since the BLC parameter is not included in the metadata, the BLC module in the ISP emulator 205 and other ISP processing modules without ISP parameters can be set to the off state.
[0124] Similarly, based on the metadata 202 of the second scenario, i.e., the LSC parameter, WBC parameter, and CCM parameter in metadata B, through the ISP emulator 205, based on the ISP pipeline of the corresponding platform, the RAW image pair corresponding to the second scenario is processed forward through the LSC module corresponding to the LSC parameter, the WBC module corresponding to the WBC parameter, and the CCM module corresponding to the CCM parameter along the direction of the ISP pipeline, so as to obtain the target image pair corresponding to metadata B, which includes the target noisy image and the target denoised image. It should be noted that since the source data image to be processed is the image data obtained after being processed by the CCM module in the RGB domain of the ISP, the BLC module and other ISP processing modules without ISP parameters in the ISP emulator 205 can be set to the off state.
[0125] In some embodiments, the imaging device includes an ISP. Before performing image processing on the RAW image pair corresponding to the metadata to obtain the target image pair in the same data domain as the image to be processed based on each metadata, the method further includes: when the parameter value of at least one image parameter in the metadata changes, performing alignment processing on the parameter value of the ISP processing module corresponding to the changed parameter value in the ISP.
[0126] It can be understood that when the parameter value of at least one image parameter in the metadata changes, after performing alignment processing on the parameter value of the ISP processing module corresponding to the changed parameter value in the ISP, then, based on the updated metadata, performing image processing on the RAW image pair corresponding to the changed metadata to obtain the target image pair in the same data domain as the image to be processed. In this way, when the parameters on the ISP pipeline are adjusted, it only needs to align the ISP with the actual ISP parameters and then re-simulate, without the need to re-collect any data.
[0127] It should be noted that the description of the same steps and the same content in this embodiment and other embodiments can refer to the description in other embodiments, and will not be repeated here.
[0128] Refer to Figure 5 As shown, an embodiment of the present application provides a data generation method, which is applied to an imaging device, and the method includes the following steps.
[0129] Step 501, obtain the source data image to be processed and the metadata of at least one scenario, where the source data image to be processed includes a source video image, and the metadata of at least one scenario is obtained by statistically analyzing the image parameters in each frame of the scenario video image of at least one scenario, and the number of frames in the source video image is greater than or equal to the number of frames in the metadata of each scenario.
[0130] In the embodiments of the present application, the source data image to be processed may be a source video image, and the metadata of the scene may include sub-metadata of multiple frames corresponding to the scene video image. The sub-metadata may include the image parameters of the frame images in the scene video image. The multiple frames of sub-metadata included in the metadata of the scene are arranged in the order of the image frames in the scene video image.
[0131] In the embodiments of the present application, the scene video image may be the original video image obtained by converting the light source signals captured by the target sensor in different scenes into digital signals, and the video image data obtained after the original video image is processed by ISP.
[0132] In the embodiments of the present application, since the image parameters of each frame image in the metadata of each scene are required, it is necessary to perform degradation processing and noise processing on each frame image in the source video image. In order to make full use of the image parameters of each frame image in the metadata of each scene, it is necessary to ensure that the number of frames in the source video image is greater than or equal to the number of frames in the metadata of each scene. In this way, when the number of frames in the source video image is greater than or equal to the number of frames in the metadata of each scene, the metadata of each scene can perform degradation processing and noise processing on some or all of the frame images in the source video image.
[0133] In the embodiments of the present application, the source video image to be processed and the metadata of at least one scene are obtained. Among them, the metadata of at least one scene can be obtained in the following way: by collecting the scene video images of at least one scene through the target sensor, statistically analyzing the image parameters of each frame image in the scene video image of each scene, obtaining the sub-metadata of each frame image in this scene, and arranging the sub-metadata of each frame image in this scene in the order of the image frames to obtain the metadata corresponding to the scene video image.
[0134] Step 502: Based on the metadata of at least one scene, perform degradation processing and noise processing on the source video image to obtain at least one pair of original video images in the scene, where the pair of original video images includes a noisy original video image and the corresponding noise-free original video image.
[0135] In the embodiments of the present application, since the number of frames of the source video image is greater than or equal to the number of frames of the metadata. In one case, the number of frames of the source video image is equal to the number of frames of the metadata, and the sub-metadata in the metadata corresponds one-to-one with all the image frames in the source video image, so that each sub-metadata in the metadata is used to perform degradation processing and noise processing on the corresponding image frames in the source video image in sequence. In the second case, the number of frames of the source video image is greater than the number of frames of the metadata, and a part of the source video image that is an integer multiple of the number of frames of the metadata can be intercepted from the source video image. For example, if the number of frames in the source video image is 32 frames and the number of frames in the metadata is 15 frames, any 15-frame part of the source video image with the same number of frames as the metadata can be intercepted from the source video image, and the sub-metadata in the metadata corresponds one-to-one with all the image frames in the source video image; it is also possible to intercept 30-frame part of the source video image that is twice the number of frames of the metadata from the source video image. Further, after performing degradation processing and noise processing on the corresponding image frames in the first 15-frame part of the source video image in sequence through each sub-metadata in the metadata, and then, based on each sub-metadata in the same metadata, perform degradation processing and noise processing on the corresponding image frames in the second 15-frame part of the source video image in sequence.
[0136] In the embodiments of the present application, for the metadata of at least one scene, when the metadata of each scene performs degradation processing and noise processing on the source video image, in accordance with the order of the sub-metadata in the metadata of each scene, the corresponding frame images in the source video image are sequentially subjected to degradation processing and noise processing to obtain the original video image pair under this scene, so as to obtain the original video image pairs under at least one scene, where the original video image pair includes a noisy original video image and the corresponding noise-free original video image.
[0137] Step 503: Process the original video image pairs under the corresponding scenes based on the metadata of at least one scene to obtain at least one pair of target video images having the same data domain as the source video image.
[0138] In the embodiments of the present application, the pair of target video images includes a target noisy video image and a target denoised video image.
[0139] In the embodiments of the present application, the pair of target video images is used to train a noise reduction model.
[0140] In the embodiments of the present application, for the metadata of at least one scene, when the metadata of each scene processes the original video image pairs under the corresponding scenes, in accordance with the order of the sub-metadata in the metadata of each scene, the corresponding frame images in the original video image pairs are sequentially subjected to simulation processing to obtain the pair of target video images under this scene, so as to obtain at least one pair of target video images having the same data domain as the source video image.
[0141] As can be seen from the above, in the embodiment of the present application, the acquisition of video data required for the noise reduction model to be trained is decoupled from the image content and the image shooting scene. The content of the target domain video is provided by the image content in the video image, and the specific noise performance of the target domain video image in different scenes is simulated by combining the metadata corresponding to each frame of the image in different video images with the ISP emulator for the video image degraded to the RAW domain. The source data without limited source can greatly expand the image content of the training data, and the degradation combined with the metadata and the simulation through the ISP emulator ensure the consistency between the training data distribution and the actual acquisition data distribution during actual use, solving the problem that video training data cannot be obtained by stacking frames in the prior art.
[0142] It should be noted that the descriptions of the same steps and the same content in this embodiment and other embodiments can refer to the descriptions in other embodiments and will not be repeated here.
[0143] Taking an implementation scenario as an example, the present application further describes a data generation method provided in the embodiment of the present application. Applied to a shooting device, refer to Figure 6 As shown, this method can be implemented in the following manner.
[0144] Step 601, determine the calibration noise model of the target sensor.
[0145] Here, first, determine the target sensor that needs to reduce noise and calibrate the RAW domain noise parameters of the target sensor. Here, the commonly used noise model is the Poisson-Gaussian noise model as shown in the above formula (1).
[0146] It can be seen from the above formulas (1) to (3) that when the analog gain g is fixed, the slope of the curve formed by the variance and the mean of the pixels obtained after shooting the same pixel value on the object to be photographed is the analog gain to be calibrated that we want, that is, the Poisson term part Poisson(x), and the Gaussian term Normal(0, y) can be obtained by setting the mean to 0, that is, shooting a pure black image and performing calibration. In this way, the calibration noise model can be obtained.
[0147] Step 602, collect the metadata of M application scenarios of the noise reduction model, and collect N source data images required for degradation.
[0148] Here, M pieces of metadata for the application scenarios of noise reduction models are collected, and N pieces of source data (also known as source data images) required for degradation are collected. Both M and N are integers greater than or equal to 2, and the numbers of M and N are obtained based on experience. Among them, the metadata is used for data degradation and simulation. Here, the metadata can include data describing the collected scenarios, such as the white balance parameters and lens shadow parameters of the collected scenarios for data degradation and simulation. The metadata can also include the brightness parameters describing the currently captured RAW images for data degradation, such as the average brightness and the maximum brightness of the RAW images captured in the current collection scenario. Of course, the metadata can also include the scene simulation gain of the image sensor in the collection scenario.
[0149] Here, the source data required for data degradation is not restricted in terms of its source. The source data can be captured by the target sensor or be an open copyright-free dataset on the network, as long as the image has little noise. It should be noted that the source data with unrestricted sources can greatly expand the image content of the training data. For each collection scenario, only one scene image needs to be captured by the target sensor, and the image parameters of this scene image are statistically analyzed to obtain the corresponding metadata for this scene image. The collected scenarios should cover as many scenarios as possible required by the noise reduction model, and the image content of the source data should also include as much content as possible obtained when the noise reduction model is used.
[0150] Step 603: Degrade the N source data images based on the calibrated noise model and the M pieces of metadata to generate P pairs of noisy RAW images and corresponding noiseless RAW images.
[0151] Here, based on the calibrated noise model and the M pieces of metadata, the N source data images are degraded to obtain P (P = M × N) noisy RAW images and noiseless RAW images. Exemplarily, a simplified ISP pipeline is as Figure 3 shown. Among them, the solid arrows indicate the direction of data flow, and degradation refers to the process of, based on the ISP pipeline of the corresponding platform and the metadata of each scenario, retreating the source data from the domain where the source data is located (such as the YUV domain or the RGB domain) back to the RAW domain against the direction of the ISP pipeline (the direction indicated by the dashed arrow). Most of the modules in the ISP pipeline are reversible and can directly perform inverse operations based on the metadata. If there are modules that are irreparable in principle, calculate the approximate range of their influence based on the metadata and add perturbations to the data to be degraded.
[0152] In the actual operation process, the embodiments of this application are not limited to Figure 3Rather than the ISP pipeline form shown, it can adapt to the ISP pipelines of different manufacturers. As long as the emulator of the corresponding ISP pipeline and the metadata required for simulation can be obtained, even if some modules cannot obtain the corresponding metadata, as long as the approximate value range affecting the module input is roughly known, it can be roughly solved by adding corresponding random perturbations during the degradation process.
[0153] Step 604: For the P pairs of noisy RAW images and noise-free RAW images, each pair of data is respectively simulated through the ISP emulator in combination with the corresponding metadata to obtain a target domain image pair, where the target domain image pair includes a noisy target domain image and a noise-free target domain image.
[0154] Here, the noisy target domain image corresponds to the above-mentioned target noisy image, and the noise-free target domain image corresponds to the above-mentioned target noise-free image. The target domain image pair is the data pair for training the denoising model.
[0155] In the embodiments of the present application, for the P pairs of noisy RAW images and noise-free RAW images that are degraded to the RAW domain, each pair of data is respectively simulated through the ISP emulator in combination with the corresponding metadata to obtain a target domain image pair. In this way, the degradation combined with metadata and the simulation through the ISP emulator ensure the consistency between the training data distribution and the actual acquisition data distribution during actual use.
[0156] It should be noted that when the parameters of the ISP pipeline are adjusted, it only needs to align the parameters of the ISP emulator with those of the actual ISP and then re-simulate, without the need to re-collect any data.
[0157] In some embodiments, this solution can be easily extended from the generation of image denoising data to the generation of video denoising data. The generation of video denoising data has always been a more difficult problem than the generation of image denoising data. Because the frame stacking scheme either cannot be simply extended to videos (fusing N frames every time the position moves), or cannot handle movement well (extending the single-frame fused image into a video).
[0158] The source data in the embodiments of the present application can be clean video frame data. The metadata collected in each scene is also obtained by collecting videos, so the previous steps can be reused. For each frame of source data, simulate in combination with the metadata, and then synthesize the simulated frames into video data.
[0159] It should be noted that the descriptions of the same steps and the same content in this embodiment and other embodiments can refer to the descriptions in other embodiments, and will not be repeated here.
[0160] As described above, in the embodiments of the present application, data pairs for training a denoising model are generated. According to the data domain on which the denoising model acts, the target domain generated in the embodiments of the present application can be the RAW domain, the RGB domain, the YUV domain, etc. Based on the ISP emulator, this solution splits the generation of training data for the target domain denoising model into two parts, that is, decouples the image content in the training images included in the training data required by the denoising model to be trained from the image shooting scene. Specifically, the source data provides the content of the target domain image pair, and the specific performance of this content in different scenes is simulated for the image content by combining the metadata in different scenes with the ISP emulator. In this way, through the decoupling method, the training data of the denoising model is not restricted by the scene, saving the labor cost required to obtain data, improving the robustness of the obtained data to the ISP pipeline parameters, and improving the speed and quality of data generation.
[0161] Referring to Figure 7 , Figure 7 FIG. 7 is a data generation device provided by an embodiment of the present application. The data generation device 7 includes:
[0162] An acquisition module 701, configured to acquire a source data image to be processed and metadata of at least one scene;
[0163] A first processing module 702, configured to perform degradation processing and noise processing on the source data image to be processed based on the metadata of at least one scene to obtain at least one pair of original images in at least one scene, where the pair of original images includes a noisy original image and a corresponding noise-free original image;
[0164] A second processing module 703, configured to process the pair of original images in the corresponding scene based on the metadata of at least one scene to obtain at least one pair of target images having the same data domain as the source data image to be processed.
[0165] In other embodiments of the present application, the acquisition module 701 is further configured to collect a first number of first scene images through a target sensor in at least one scene; a third processing module is configured to perform statistical analysis on the image parameters in the first scene images to obtain image signal processing (ISP) parameters of at least one scene and brightness parameters of the first scene images in the original domain; and use the ISP parameters and brightness parameters corresponding to at least one scene as the metadata of at least one scene.
[0166] In other embodiments of the present application, the metadata further includes the scene simulation gain when the target sensor captures the first scene image. The first processing module 702 is further configured to perform degradation processing on the source data image to be processed based on the ISP parameters and the brightness parameters in the metadata of at least one scene, so as to obtain a noise-free original image in at least one scene; determine the scene noise parameters of the target sensor in at least one scene based on the scene simulation gain of at least one scene; add noise to the noise-free original image in the corresponding scene according to the scene noise parameters of at least one scene, so as to obtain a noisy original image in at least one scene; and use the noise-free original image and the noisy original image as the original image pair in at least one scene.
[0167] In other embodiments of the present application, the acquisition module 701 is further configured to obtain at least two second scene images of the object to be photographed in the same scene captured by the target sensor under at least two calibration simulation gains; the first processing module 702 is further configured to perform original domain noise calibration on the target sensor based on the pixel values of each pixel point in at least two second scene images, so as to obtain the calibration noise parameters corresponding to the target sensor under at least two calibration simulation gains; determine the calibration noise model of the target sensor based on at least two calibration simulation gains and the corresponding calibration noise parameters; and input the scene simulation gain of at least one scene into the calibration noise model to obtain the scene noise parameters of the target sensor in at least one scene.
[0168] In other embodiments of the present application, the second processing module 703 is further configured to perform ISP processing on the original image pair in the corresponding scene based on the ISP parameters in the metadata of at least one scene, so as to obtain at least one target image pair having the same data domain as the source data image to be processed.
[0169] In other embodiments of the present application, the photographing device includes an ISP. The second processing module 703 is further configured to perform alignment processing on the parameter values of the ISP processing module corresponding to the changed parameter value in the ISP when the parameter value of at least one image parameter in the metadata changes.
[0170] In other embodiments of the present application, the source data image to be processed includes a source video image. The metadata of at least one scene is obtained by statistically analyzing the image parameters in each frame of image based on the scene video images of at least one scene; wherein the number of frames in the source video image is greater than or equal to the number of frames in the metadata of each scene.
[0171] In other embodiments of the present application, the acquisition module 701 is further configured to obtain a noise reduction model to be trained; wherein the data domain of the noise reduction model in the ISP is the same as the data domain of the target image pair; the training module is further configured to train the noise reduction model to be trained based on the target image pair to obtain a trained noise reduction model.
[0172] Reference Figure 8 , Figure 8 The present application embodiment provides a photographing device, and the photographing device includes:
[0173] An image acquisition component 801, including a lens and a sensor, for acquiring images;
[0174] A processor 802, configured to perform the following steps:
[0175] Obtain a source data image to be processed and metadata of at least one scene;
[0176] Perform degradation processing and noise processing on the source data image to be processed based on the metadata of the at least one scene to obtain an original image pair under the at least one scene, where the original image pair includes a noisy original image and a corresponding noise-free original image;
[0177] Process the original image pair under the corresponding scene based on the metadata of the at least one scene to obtain at least one target image pair having the same data domain as the source data image to be processed.
[0178] The present application embodiment provides an electronic device, and the electronic device includes a photographing device.
[0179] In practical applications, the electronic device includes but is not limited to various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices may be smart speakers, smart TVs, smart air conditioners, smart vehicle-mounted devices, etc. The portable wearable devices may be smart watches, smart bracelets, head-mounted devices, etc.
[0180] The present application embodiment provides a storage medium, and the storage medium stores one or more computer programs. The one or more computer programs can be executed by one or more processors to implement some or all of the steps in the above method. The storage medium can be transient or non-transient.
[0181] The present application embodiment provides a computer program, including computer-readable code. When the computer-readable code runs in an electronic device, the processor in the electronic device executes to implement some or all of the steps in the above method.
[0182] An embodiment of the present application provides a computer program product. The computer program product includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, some or all of the steps in the above method are implemented. The computer program product can be specifically implemented in a manner of hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium. In other embodiments, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0183] It should be noted here that the descriptions of the above embodiments tend to emphasize the differences between the embodiments, and their similarities can be referred to each other. The descriptions of the above device, storage medium, computer program, and computer program product embodiments are similar to the descriptions of the above method embodiments and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.
[0184] It should be noted here that the descriptions of the above storage medium and device embodiments are similar to the descriptions of the above method embodiments and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the embodiments of the storage medium and device of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.
[0185] The above processor can be at least one of an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Digital Signal Processing Device (DSPD), a Programmable Logic Device (PLD), a Field Programmable Gate Array (FPGA), a Central Processing Unit (CPU), a controller, a microcontroller, and a microprocessor. It can be understood that other electronic devices can also implement the functions of the above processor, and the embodiments of the present application do not make specific limitations.
[0186] The above computer storage medium / memory may be a read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; it may also be various terminals including one or any combination of the above memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0187] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the "in one embodiment" or "in an embodiment" that appears throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics may be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the magnitude of the sequence numbers of the above steps / processes does not mean the order of execution. The order of execution of each step / process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The sequence numbers of the embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0188] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including that element.
[0189] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the couplings, direct couplings, or communication connections between the various components shown or discussed can be through some interfaces. The indirect couplings or communication connections of devices or units can be electrical, mechanical, or other forms.
[0190] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0191] In addition, each functional unit in the embodiments of this application can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware, or in the form of hardware plus software functional units.
[0192] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments. The foregoing storage medium includes: various media such as removable storage devices, read-only memory (ROM), magnetic disks, or optical discs that can store program codes.
[0193] Alternatively, if the above-mentioned integrated units of this application are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a vehicle-mounted terminal (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods of the various embodiments of this application. The foregoing storage medium includes: various media such as removable storage devices, ROM, magnetic disks, or optical discs that can store program codes.
[0194] The above are only the implementation manners of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application.
Claims
1. A data generation method, characterized in that: The method comprises: Acquire a source data image to be processed and metadata of at least one scene; Performing degradation processing and noise processing on the source data image to be processed based on the metadata of the at least one scene to obtain an original image pair under the at least one scene, wherein the original image pair includes a noisy original image and a corresponding noise-free original image; The original image pair of the corresponding scene is processed based on the metadata of the at least one scene to obtain at least one target image pair having the same data domain as the source data image to be processed.
2. The method according to claim 1, characterized in that The obtaining metadata of at least one scene includes: Collecting a first number of first scene images by a target sensor in at least one of the scenes; Performing statistical analysis on image parameters in the first scene image to obtain image signal processing (ISP) parameters of the at least one scene and brightness parameters of the first scene image in the original domain; The ISP parameter and the brightness parameter corresponding to the at least one scene are used as metadata of the at least one scene.
3. The method according to claim 2, characterized in that The metadata also includes a scene simulation gain when the target sensor acquires the first scene image, and the metadata based on the at least one scene performs degradation processing and noise processing on the source data image to be processed to obtain an original image pair under the at least one scene, including: Based on the ISP parameter and the brightness parameter in the metadata of the at least one scene, performing degradation processing on the source data image to be processed to obtain a noise-free original image of the at least one scene; Determining a scene noise parameter of the target sensor in the at least one scene based on a scene simulation gain of the at least one scene; According to the scene noise parameter of the at least one scene, adding noise to the noise-free original image of the corresponding scene to obtain the noisy original image of the at least one scene; The noise-free original image and the noisy original image are used as an original image pair in the at least one scene.
4. The method according to claim 3, characterized in that The method of determining a scene noise parameter of the target sensor in the at least one scene based on the scene simulation gain of the at least one scene comprises: Obtain at least two second scene images of a photographed object in the same scene captured by the target sensor under at least two calibrated analog gains; Based on the pixel value of each pixel point in at least two second scene images, calibrate the original domain noise of the target sensor to obtain the calibration noise parameters corresponding to the target sensor under the at least two calibration analog gains; Determining a calibration noise model of the target sensor based on the at least two calibration analog gains and corresponding calibration noise parameters; The scene simulation gain of the at least one scene is input into the calibration noise model to obtain the scene noise parameter of the target sensor in the at least one scene.
5. The method according to any one of claims 2 to 4, characterized in that: The processing of the original image pair corresponding to the scene based on the metadata of the at least one scene to obtain at least one target image pair having the same data domain as the source data image to be processed includes: Based on the ISP parameters in the metadata of the at least one scene, ISP processing is performed on the original image pair in the corresponding scene to obtain at least one target image pair having the same data domain as the source data image to be processed.
6. The method according to any one of claims 1 to 4, characterized in that: The shooting device includes an ISP, and before the original image pair corresponding to the scene is processed based on the metadata of the at least one scene to obtain at least one target image pair having the same data domain as the source data image to be processed, the method further includes: When a parameter value of at least one image parameter in the metadata changes, an alignment process is performed on the parameter values of the ISP processing modules in the ISP corresponding to the changed parameter value.
7. The method according to any one of claims 1 to 4, characterized in that: The source data image to be processed includes a source video image, and the metadata of the at least one scene is obtained by statistically analyzing image parameters in each frame image based on the scene video image of the at least one scene; The number of frames in the source video image is greater than or equal to the number of frames in the metadata of each of the scenes.
8. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: Acquire a denoising model to be trained; wherein the data domain of the denoising model in the ISP is the same as the data domain of the target image pair; Based on the target image pair, the denoising model to be trained is trained to obtain a trained denoising model.
9. A data generating device, characterized in that: The method comprises: An acquisition module, used for acquiring a source data image to be processed and metadata of at least one scene; A first processing module, configured to perform degradation processing and noise processing on the source data image to be processed based on the metadata of the at least one scene, so as to obtain an original image pair under the at least one scene, wherein the original image pair includes a noisy original image and a corresponding noise-free original image; The second processing module is used to process the original image pair in the corresponding scene based on the metadata of the at least one scene to obtain at least one target image pair having the same data domain as the source data image to be processed.
10. A photographing device, characterized in that: The shooting device comprises: An image acquisition component, including a lens and a sensor, for acquiring images; A processor, configured to execute the data generation method according to any one of claims 1 to 8.
11. A storage medium, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the data generation method according to any one of claims 1 to 8.
12. A computer program product, comprising a computer program or instructions, wherein when the computer program or instructions are executed by a processor, the data generating method according to any one of claims 1 to 8 is implemented.