Image generation method and device, electronic equipment and storage medium
By applying Gaussian convolution blur and adding noise to RAW images, a noisy RAW image simulating point diffusion effect is generated, which solves the problem of noise distribution mismatch and improves the denoising effect of the image denoising model.
Patent Information
- Application Number
- CN202510985525.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies struggle to generate noisy RAW images with noise distributions that more closely resemble the actual noise distributions, resulting in poor denoising performance of image denoising models.
By applying Gaussian convolution blur and adding noise to RAW format images, noisy RAW images simulating point diffusion effects are generated and used to train image denoising models.
It improves the denoising effect of the image denoising model, making the noise distribution closer to the real noise distribution and improving the quality of image denoising.
Smart Images

Figure CN120876289A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image denoising technology, and in particular to an image generation method, apparatus, electronic device, and storage medium. Background Technology
[0002] In the field of image denoising technology, in order to obtain a well-trained image denoising model, noise can be added to the initial RAW image of the sample to obtain a noisy RAW image of the sample; the noisy RAW image of the sample is input into the image denoising model with an initial structure to obtain the image denoising result; based on the difference between the image denoising result and the initial RAW image of the sample, the model parameters of the image denoising model are adjusted until the model converges to obtain a well-trained image denoising model.
[0003] Compared to the initial RAW image, if the noise distribution in the noisy RAW image closely resembles the noise distribution in a real-world RAW image (which can be called the real noise distribution), then a trained image denoising model using both the initial and noisy RAW images can remove noise distributed according to this real noise distribution. Correspondingly, the trained image denoising model performs better on RAW images acquired from real-world scenes. Therefore, to improve image denoising performance, an image generation method is urgently needed to obtain noisy RAW images that more closely approximate the real noise distribution compared to the initial RAW image. Summary of the Invention
[0004] The purpose of this invention is to provide an image generation method, apparatus, electronic device, and storage medium to obtain a noisy RAW image whose noise distribution is closer to the true noise distribution than the initial RAW image of the sample. The specific technical solution is as follows:
[0005] In a first aspect of the present invention, an image generation method is provided, the method comprising:
[0006] Obtain the raw RAW format image to be processed;
[0007] Based on the interpolation algorithm, the image to be processed is converted into the standard red-green-blue sRGB format to obtain the image to be blurred;
[0008] For each color channel of the image to be blurred, Gaussian convolution blur is performed on the pixel value of each pixel position in the image under that color channel to obtain the blurred pixel value of each pixel position for that color channel.
[0009] For each pixel location, the blurred pixel value of the corresponding color channel to be used at that pixel location is obtained to obtain a blurred RAW image; wherein, the color channel to be used corresponding to a pixel location is: the color channel at that pixel location in the image to be processed;
[0010] Noise is added to the blurred RAW image to obtain a noisy RAW image.
[0011] Optionally, after adding noise to the blurred RAW image to obtain a noisy RAW image, the method further includes:
[0012] The noisy RAW image is input into the image denoising model with the initial structure to obtain the predicted denoised image;
[0013] Calculate the loss function value representing the difference between the predicted denoised image and the image to be processed;
[0014] The model parameters of the image denoising model with the initial structure are adjusted based on the loss function value until the preset convergence condition is met, thus obtaining the trained image denoising model.
[0015] Optionally, obtaining the raw RAW format image to be processed includes:
[0016] Acquire RAW format images captured by the image acquisition device under normal illumination conditions;
[0017] Based on the images acquired by the image acquisition device, the image to be processed is determined.
[0018] Optionally, based on the images acquired by the image acquisition device, determining the image to be processed includes:
[0019] The image acquired by the image acquisition device is subjected to preset processing to obtain the image to be processed; wherein, the preset processing includes at least one of the following: black level correction and pixel value normalization.
[0020] Optionally, the image acquisition device may use the minimum supported sensitivity when acquiring images.
[0021] And / or,
[0022] The brightness value of the image acquired by the image acquisition device falls within the range of neither overexposure nor underexposure.
[0023] Optionally, the interpolation algorithm is a linear interpolation algorithm.
[0024] In a second aspect of the present invention, an image generation apparatus is provided, the apparatus comprising:
[0025] The acquisition module is used to acquire the raw RAW format image to be processed;
[0026] The conversion module is used to convert the image to be processed into the standard red-green-blue sRGB format based on the interpolation algorithm to obtain the image to be blurred;
[0027] The blur module is used to perform Gaussian convolution blur on the pixel value of each pixel position in the image to be blurred for each color channel of the image to be blurred, so as to obtain the blurred pixel value of each pixel position for the color channel.
[0028] The sampling module is used to obtain the blurred pixel value of the corresponding color channel to be used at each pixel position, so as to obtain a blurred RAW image; wherein, the color channel to be used corresponding to a pixel position is: the color channel at the pixel position in the image to be processed;
[0029] The noise-adding module is used to add noise to the blurred RAW image to obtain a noisy RAW image.
[0030] Optionally, the device further includes:
[0031] The training module is used to add noise to the blurred RAW image after the noise-adding module performs the noise addition, to obtain a noisy RAW image, and then input the noisy RAW image into an image denoising model with an initial structure to obtain a predicted denoised image; calculate the loss function value representing the difference between the predicted denoised image and the image to be processed; and adjust the model parameters of the image denoising model with the initial structure based on the loss function value until a preset convergence condition is reached to obtain a trained image denoising model.
[0032] Optionally, the acquisition module is specifically used for:
[0033] Acquire RAW format images captured by the image acquisition device under normal illumination conditions;
[0034] Based on the images acquired by the image acquisition device, the image to be processed is determined.
[0035] Optionally, the acquisition module is specifically used for:
[0036] The image acquired by the image acquisition device is subjected to preset processing to obtain the image to be processed; wherein, the preset processing includes at least one of the following: black level correction and pixel value normalization.
[0037] Optionally, the image acquisition device may use the minimum supported sensitivity when acquiring images.
[0038] And / or, the brightness value of the image acquired by the image acquisition device belongs to the range that is neither overexposed nor underexposed.
[0039] Optionally, the interpolation algorithm is a linear interpolation algorithm.
[0040] In a third aspect of the present invention, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;
[0041] Memory, used to store computer programs;
[0042] When a processor executes a program stored in memory, it implements the image generation method steps described in any of the first aspects above.
[0043] In a fourth aspect of the present invention, a computer-readable storage medium is provided, wherein a computer program is stored therein, and when executed by a processor, the computer program implements the image generation method steps described in any of the first aspects above.
[0044] This invention also provides a computer program product containing instructions that, when run on a computer, cause the computer to execute any of the image generation methods described above.
[0045] The image generation method provided in this invention involves an electronic device acquiring a RAW format image to be processed; converting the image to be processed into sRGB format based on an interpolation algorithm to obtain a blurred image; performing Gaussian convolution blurring on the pixel values of each pixel position in the blurred image under that color channel for each color channel, obtaining blurred pixel values for each pixel position relative to that color channel; obtaining blurred pixel values for each pixel position relative to the corresponding color channel to be used, thus obtaining a blurred RAW image; the color channel to be used corresponding to a pixel position is the color channel at that pixel position in the image to be processed; and adding noise to the blurred RAW image to obtain a noisy RAW image.
[0046] Based on the above processing, for the RAW format image to be processed, the electronic device performs Gaussian convolution blurring on the image to obtain the blurred pixel values of each color channel. That is, the electronic device obtains the blurred pixel values of each color channel from the image to be processed, which are pixel values superimposed with dot diffusion effect. Therefore, based on the blurred pixel values of each color channel, the electronic device determines the blurred pixel value corresponding to each pixel position, obtaining a blurred RAW image. Noise is added to the blurred RAW image, that is, noise is added to the RAW image superimposed with dot diffusion effect. The resulting noisy RAW image, compared to the image to be processed, is an image with noise added after considering the dot diffusion effect. Since noise in RAW images acquired in real-world scenes often exhibits dot diffusion effect, adding noise after considering the dot diffusion effect makes the noise distribution in the resulting image closer to the real noise distribution. Accordingly, compared to the image to be processed, the noise distribution in the noisy RAW image obtained based on the image to be processed is closer to the real noise distribution. Subsequently, training an image denoising model based on noisy RAW images with noise distributions that are closer to the actual noise distribution can improve the image denoising performance of the trained model.
[0047] Of course, implementing any product or method of the present invention does not necessarily require achieving all of the advantages described above at the same time. Attached Figure Description
[0048] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these drawings.
[0049] Figure 1 This is a first flowchart of an image generation method provided in an embodiment of the present invention;
[0050] Figure 2 This is a second flowchart of the image generation method provided in an embodiment of the present invention;
[0051] Figure 3 This is a structural diagram of an image denoising model provided in an embodiment of the present invention;
[0052] Figure 4 A structural diagram of an image generation apparatus provided in an embodiment of the present invention;
[0053] Figure 5 This is a structural diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0054] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art based on the present invention are within the scope of protection of the present invention.
[0055] In real-world scenarios, under low ambient light conditions (also known as dark or low-light environments), images captured by image acquisition devices may contain significant noise. This can lead to loss of detail, low clarity, and ultimately, poor visual quality. Furthermore, processing images containing significant noise (low-quality images) can negatively impact the results and hinder downstream tasks. For instance, image analysis based on low-quality images can result in inaccurate analysis results. Therefore, image denoising (or noise reduction) of low-quality images is crucial.
[0056] In recent years, with the development of deep learning (i.e., neural network) technology, image denoising based on deep learning has become the mainstream image denoising method. A deep learning model (which can be called an image denoising model or noise reduction model) with an initial structure is trained using high-quality training data until it converges, resulting in a well-trained image denoising model. Using this trained model to denoise the image yields better results compared to traditional image denoising methods.
[0057] In the field of image denoising technology, image acquisition devices first acquire RAW format images (which can be simply referred to as RAW images). The pixel value of each pixel in a RAW image (which can be called RAW domain data) has a linear relationship with the incident light intensity. Performing ISP (Image Signal Processing) on the RAW image, that is, performing ISP on the pixel value of each pixel in the RAW image, can convert the RAW image into an sRGB (standard Red Green Blue) image; the pixel value of each pixel in the sRGB image can be called sRGB domain data. ISP can include various non-linear processing methods (such as white balance, gamma transform, etc.).
[0058] As can be seen, the RAW domain data is the data before ISP, while the sRGB domain data is the data after ISP.
[0059] Therefore, RAW domain data is unaffected by the nonlinear processing of ISP. Based on the principles of physical imaging, determining the noise parameters (which can be called noise calibration) of the noise included in RAW domain data is relatively easy. Consequently, the ease of accurately modeling RAW domain data also makes image denoising of RAW images easier.
[0060] Since sRGB domain data is obtained by ISP of RAW domain data, the noise included in sRGB domain data is obtained by ISP of the noise included in RAW domain data. As mentioned earlier, ISP includes various non-linear processing steps, which makes the noise included in sRGB domain data more complex, thus making noise calibration of sRGB domain data more difficult; correspondingly, accurate modeling of sRGB domain data is also more difficult, which in turn makes image denoising of sRGB images more difficult.
[0061] Therefore, in order to reduce the difficulty of image denoising, after obtaining the RAW image, the RAW image can be input into a trained image denoising model to obtain a denoised RAW image; then, ISP can be performed on the denoised RAW image to obtain an sRGB image, which is the denoised sRGB image.
[0062] To obtain a well-trained image denoising model, noise can be added to the initial RAW image to obtain a noisy RAW image. This noisy RAW image is then input into the initial structure of the image denoising model to obtain the denoised result. Based on the difference between the denoised result and the initial RAW image, the model parameters are adjusted until the model converges, resulting in a well-trained image denoising model. A noisy RAW image can also be simply referred to as a noisy image.
[0063] Compared to the initial RAW image, if the noise distribution in the noisy RAW image is close to the real noise distribution, then a trained image denoising model using both the initial and noisy RAW images can remove noise distributed as the real noise distribution. Correspondingly, the trained image denoising model performs well on RAW images acquired from real-world scenes. Therefore, the closeness of the noise distribution in the noisy RAW image to the real noise distribution—that is, the accuracy of noise data synthesis—is crucial to the overall performance of the image denoising model (i.e., the image denoising effect). Thus, to improve image denoising performance, an image generation method is urgently needed to obtain noisy RAW images with noise distributions closer to the real noise distribution than the initial RAW image.
[0064] Noise can include shot noise related to signal strength (i.e., the aforementioned ambient light intensity) and readout noise unrelated to signal strength. In related techniques, shot noise and / or readout noise are typically added to the initial RAW image of the sample. However, with the deepening of image denoising research, researchers have discovered that noise in RAW images acquired in real-world scenes exhibits characteristics similar to the "point spread effect" in optical systems. The point spread effect itself is a phenomenon in optical systems (which can be called spatial response), and the PSF (Point Spread Function) can be used as the aforementioned spatial response function to characterize the point spread effect. In an optical system, in an image of a point light source acquired by a camera, the pixel values of other pixels within the neighborhood of the point light source are affected by the light emitted by the point light source; this phenomenon is the point spread effect. In the field of image denoising, the influence of noise on the noise of a pixel by the noise of other pixels within its neighborhood is called the point spread effect of noise.
[0065] Specifically, experiments conducted by technicians revealed that for images acquired in low-light environments, the reduced number of photons in such environments leads to Poisson noise (i.e., the aforementioned shot noise) dominating the image due to the statistical randomness of photons reaching the sensor in the image acquisition device. Furthermore, processing the image using PSF (Pressure Sample Function) to amplify the impact of noise at each pixel on the noise of other pixels in its neighborhood (this can be termed amplifying the spatial correlation of noise) results in images with significant speckle or banded artifacts. This demonstrates that in RAW images acquired in real-world scenes, the noise at each pixel is indeed affected by the noise of other pixels in its neighborhood, indicating a point spread effect in the noise of RAW images acquired in real-world scenes.
[0066] Therefore, to obtain a noisy RAW image whose noise distribution is closer to the true noise distribution than the initial RAW image of the sample, this invention provides an image generation method applied to an electronic device. The electronic device can be a server; the server can process images acquired by an image acquisition device based on the image generation method provided by this invention to obtain a noisy RAW image. Alternatively, the electronic device can also be the aforementioned image acquisition device; for example, after acquiring an image, the image acquisition device can process its own acquired image based on the image generation method provided by this invention to obtain a noisy RAW image. The resulting noisy RAW image is an image with added noise considering the point spread effect, and the noise distribution in the noisy RAW image is closer to the true noise distribution. Correspondingly, compared to the image to be processed, the noise distribution in the noisy RAW image obtained based on the image to be processed is closer to the true noise distribution. Subsequently, training an image denoising model based on the noisy RAW image with a noise distribution closer to the true noise distribution can improve the image denoising effect of the trained image denoising model.
[0067] See Figure 1 , Figure 1 This is a first flowchart of an image generation method provided in an embodiment of the present invention. The method may include the following steps:
[0068] S101: Obtain the raw RAW format image to be processed.
[0069] S102: Based on the interpolation algorithm, the image to be processed is converted into the standard red-green-blue sRGB format to obtain the image to be blurred.
[0070] S103: For each color channel of the image to be blurred, perform Gaussian convolution blur on the pixel value of each pixel position in the image to be blurred in that color channel to obtain the blurred pixel value of each pixel position for that color channel.
[0071] S104: For each pixel location, obtain the blurred pixel value for the corresponding color channel to be used at that pixel location, and obtain a blurred RAW image.
[0072] Among them, the color channel to be used corresponding to a pixel position is: the color channel at that pixel position in the image to be processed.
[0073] S105: Add noise to a blurred RAW image to obtain a noisy RAW image.
[0074] Based on the image generation method provided in this invention, for a RAW format image to be processed, the electronic device performs Gaussian convolution blurring on the image to be processed to obtain blurred pixel values for each color channel. That is, the electronic device obtains blurred pixel values for each color channel after blurring processing based on the image to be processed, which are pixel values superimposed with dot diffusion effects. Therefore, the electronic device determines the blurred pixel value corresponding to each pixel position based on the blurred pixel values of each color channel, obtaining a blurred RAW image. Noise is added to the blurred RAW image, that is, noise is added to the RAW image superimposed with dot diffusion effects. The resulting noisy RAW image, compared to the image to be processed, is an image with noise added after considering the dot diffusion effect. Since noise in RAW images acquired in real-world scenes often exhibits dot diffusion effects, adding noise after considering the dot diffusion effect makes the noise distribution in the obtained image closer to the real noise distribution. Accordingly, compared to the image to be processed, the noise distribution in the noisy RAW image obtained based on the image to be processed is closer to the real noise distribution. Subsequently, training an image denoising model based on noisy RAW images with noise distributions that are closer to the actual noise distribution can improve the image denoising performance of the trained model.
[0075] For step S101, the image to be processed is a RAW format image; the pixel value of each pixel position in the RAW format image is RAW domain data. RAW domain data is the original data after the light signal recorded by the photosensitive element of the image acquisition device is converted into a digital signal. The pixel value of a pixel position includes the RAW domain data of a single color channel. The color channel corresponding to each pixel position in the image to be processed can be determined according to the Bayer color filter array in the sensor of the image acquisition device. For example, the color channel may include the R (Red) channel, the G (Green) channel, and the B (Blue) channel; the Bayer color filter array can be an RGGB array.
[0076] In step S102, for each pixel position in the image to be processed, the electronic device can determine the pixel positions within the neighborhood of that pixel position (which can be referred to as the interpolation pixel positions). Then, based on the pixel values of the interpolation pixel positions, according to the interpolation algorithm, the pixel values of that pixel position under each color channel included in the aforementioned Bayer color filter array are determined to convert the image to be processed into sRGB format. For example, based on the aforementioned example, when the Bayer color filter array is an RGGB array, the color channels can include the aforementioned R channel, G channel, and B channel. The resulting sRGB format image is the image to be blurred, and the pixel values of each pixel position in the image to be blurred include the pixel values under each of the aforementioned color channels. For example, the image to be processed can be converted into an sRGB format image to be blurred using Rawpy (Raw Python, a raw Python library).
[0077] For example, the interpolation algorithm can be nearest neighbor interpolation. For each color channel other than the color channel corresponding to the pixel position (which can be called the color channel to be interpolated), the electronic device can determine the position of the pixel to be interpolated that corresponds to the color channel to be interpolated and is closest to the pixel position. Then, the pixel value of the determined pixel position to be interpolated is used as the pixel value of the pixel position in the color channel to be interpolated.
[0078] For example, the interpolation algorithm can be bilinear interpolation. For each color channel to be interpolated, the electronic device can determine all the pixel positions corresponding to that color channel, and then use the statistical value of the pixel value at the determined pixel position as the pixel value for that color channel. The statistical value can be the average, median, etc.
[0079] For step S103, for each color channel of the image to be blurred, the electronic device can extract the pixel value of each pixel position in the image under that color channel. The image composed of the extracted pixel values of each pixel position under that color channel can be called the single-channel image corresponding to that color channel.
[0080] Furthermore, the electronic device performs Gaussian convolution blurring on the pixel values of each pixel position in the image to be blurred, under the corresponding color channel; that is, it performs channel-specific blurring on the image to be blurred. For example, the electronic device can substitute preset kernel_size (kernel size) and sigma (standard deviation) into the Gaussian function to generate a convolution kernel; then, it uses the generated convolution kernel to perform convolution processing on the single-channel image corresponding to the color channel, obtaining the blurred pixel value of each pixel position for that color channel. kernel_size and sigma can be set by technicians based on their work experience, or they can be nominal values marked in the relevant devices (such as the aforementioned image acquisition device); the electronic device can perform convolution processing based on convolution methods provided by Python libraries such as scipy (Scientific Python). The blurred pixel value of each pixel position for each color channel can be denoted as convolved_rgb (convolution_red-green-blue). Processing the blurred image based on Gaussian convolution blurring is equivalent to using PSF to process the blurred image.
[0081] In step S104, the electronic device can sample the convolved_rgb image to obtain a blurred RAW image. For example, the resolution of the image to be processed can be denoted as (H, W, 1), and the resolution of the image to be blurred can be denoted as (H, W, 3); where 1 indicates that each pixel position in the image to be processed includes pixel values under only one color channel; and 3 indicates that each pixel position in the image to be blurred includes pixel values under three color channels. The electronic device can initialize and generate an image with a resolution of (H, W, 1) (which can be called an empty image), and the empty image can be denoted as psf_raw (dot diffusion_raw); the pixel value at each pixel position in the empty image can be a preset value.
[0082] For each pixel location, the electronic device can determine the blurred pixel value for each color channel at that pixel location. Then, from the determined blurred pixel values, it determines the blurred pixel value for the corresponding color channel to be used at that pixel location, which is then used as the pixel value at that pixel location in the empty image. After determining the pixel values at each pixel location in the empty image, a blurred RAW image is obtained. A blurred RAW image can also be called a RAW image after simulating the PSF process.
[0083] For example, taking the Bayer color filter array as an RGGB array, the color channels corresponding to each pixel position in the image to be processed are arranged in the RGGB pattern. For each color channel of the image to be blurred, after the electronic device obtains the blurred pixel value of each pixel position for that color channel, it can merge the blurred pixel values of each pixel position for different color channels based on the channel merging method to obtain the merged result (which can be called the blurred image).
[0084] Furthermore, for each pixel location in the empty image, the electronic device can sample the blurred RAW image according to the following formula to determine the blurred pixel value for the corresponding color channel to be used at that pixel location:
[0085] psf_raw[i][j]=convolved_rgb[i, j, 0];
[0086] psf_raw[i][j+1]=convolved_rgb[i,j+1,1];
[0087] psf_raw[i+1][j]=convolved_rgb[i+1, j, 1];
[0088] psf_raw[i+1][j+1]=convolved_rgb[i+1,j+1,2];
[0089] Where i = 2 × k, k = 0, 1, 2, ..., H / 2; j = 2 × l, l = 0, 1, 2, ..., W / 2; psf_raw[x][y] represents the blurred pixel value of the pixel position (x, y) in the empty image for the corresponding color channel to be used; (H, W) represents the resolution of the empty image; convolved_rgb[x, y, z] represents the blurred pixel value of the pixel position (x, y) in the blurred image for the z color channel; z = 0 represents the R channel; z = 1 represents the G channel; z = 2 represents the B channel.
[0090] Regarding step S105, after acquiring the blurred RAW image, the electronic device can add noise to the blurred RAW image. For example, it can add at least one preset noise to the blurred RAW image according to preset noise parameters. The preset noise includes at least one of the following: Poisson noise and Gaussian noise, resulting in a noisy RAW image, which can be called a noise_raw image.
[0091] Because the image generation method provided by this invention simulates the PSF (Power-Sensitive Function) operation on the RAW format image to be processed, and then adds noise to the blurred RAW image based on the simulated PSF operation—that is, adding noise with a distribution close to the real noise distribution to the image to be processed—a noisy RAW image is obtained. Therefore, the less noise in the image to be processed, the closer the noise distribution in the obtained noisy RAW image is to the real noise distribution. Subsequently, the image denoising effect of the trained image denoising model based on the image to be processed and the noisy RAW image is better. Therefore, to further improve the image denoising effect of the obtained image denoising model, an image with even less noise can be obtained, and then a noisy RAW image can be obtained based on the image to be processed according to the image generation method provided in this embodiment of the invention.
[0092] In some embodiments, the aforementioned step S101 may include the following steps:
[0093] Step A1: Acquire a RAW format image captured by the image acquisition device under normal illumination conditions.
[0094] Step A2: Determine the image to be processed based on the image acquired by the image acquisition device.
[0095] In this embodiment of the invention, when the ambient light intensity is within the normal illuminance range, the probability of overexposure and underexposure in the images acquired by the image acquisition device is low. The lower limit of the normal illuminance range can be 100 LUX, and the upper limit can be several thousand LUX, such as 2000 LUX. When the ambient light intensity is less than the lower limit of the normal illuminance range, such as in a nighttime environment, the probability of underexposure in the images acquired by the image acquisition device is high; while when the ambient light intensity is greater than the lower limit of the normal illuminance range, the probability of overexposure in the images acquired by the image acquisition device is high.
[0096] Based on the physical properties of RAW domain data, when an image is neither overexposed nor underexposed (which can be termed normal exposure), there is a linear relationship between light intensity and the RAW domain data value. However, when an image is overexposed or underexposed (which can be termed abnormal exposure), the relationship between light intensity and the RAW domain data value becomes non-linear. In this case, the noise included in the RAW domain data is also more complex, causing interference from the abnormal exposure factor during the subsequent training process of the image denoising model based on the noisy RAW image and the image to be processed. This results in a poorly performing image denoising model. Therefore, acquiring images with light intensity within the normal illumination range can reduce the probability of abnormal exposure in the acquired images, thereby reducing the probability of the aforementioned problems.
[0097] In one approach, the electronic device can directly use the image acquired by the image acquisition device as the image to be processed.
[0098] In another approach, step A2 may include the following steps: performing preset processing on the image acquired by the image acquisition device to obtain an image to be processed. The preset processing includes at least one of the following: black level correction and pixel value normalization.
[0099] In this embodiment of the invention, the electronic device can perform at least one preset processing on the image acquired by the image acquisition device, and then use the processing result as the image to be processed.
[0100] Black level correction of an image can remove the interference of the current from the image acquisition device itself on pixel values, i.e., remove noise generated by the current, resulting in less noise in the processed image. Pixel value normalization of an image standardizes the image data, making subsequent processing based on standardized image data more efficient.
[0101] In some embodiments, the sensitivity of the image acquisition device when acquiring an image is the minimum supported sensitivity; and / or, the brightness value of the image acquired by the image acquisition device belongs to the range of neither overexposure nor underexposure.
[0102] In this embodiment of the invention, the ISO (sensitivity) of the image acquisition device when acquiring the image is the minimum supported sensitivity, which minimizes the noise present in the image due to the physical properties of the image acquisition device itself, thereby making the noise distribution in the subsequently obtained noisy RAW image closer to the real noise distribution.
[0103] In this embodiment of the invention, the brightness value of the image acquired by the image acquisition device falls within the range of neither overexposure nor underexposure. In practical scenarios, the image acquisition device may support multiple ISOs when acquiring images. For each ISO supported by the image acquisition device, technicians can determine, through testing, the exposure time range of the image acquisition device that ensures the brightness value of the acquired image falls within the range of neither overexposure nor underexposure when the ISO is set to that ISO, thus obtaining the selectable exposure time range corresponding to that ISO. The image brightness value falling within the range of neither overexposure nor underexposure indicates that the image exposure is normal.
[0104] Subsequently, if the technician sets the ISO of the image acquisition device to this ISO, the technician can select any duration from the selectable duration range corresponding to this ISO as the exposure duration for image acquisition. This allows the image acquisition device to acquire a properly exposed image. Determining the noisy RAW image based on the properly exposed image can reduce interference from abnormal image exposure during the training of the image denoising model based on the noisy RAW image and the image to be processed, thereby improving the image denoising effect of the trained model.
[0105] In practical scenarios, technicians can set the ISO of the image acquisition device to the minimum supported ISO, and the exposure time for image acquisition can be any duration within the selectable time range corresponding to the minimum ISO. Under normal illumination conditions, multiple RAW images are acquired under different environments using the image acquisition device. Each acquired RAW image can be used as an image to be processed, denoted as `clean_raw`. Subsequently, the image to be blurred based on this image can be denoted as `clean_rgb`. This further makes the noise distribution in the resulting noisy RAW image closer to the true noise distribution; and reduces the probability of abnormal exposure in the acquired images, thereby reducing the interference caused by abnormal image exposure during the training of the image denoising model based on the noisy RAW image and the image to be processed, and improving the image denoising effect of the trained image denoising model.
[0106] In some embodiments, the interpolation algorithm is a linear interpolation algorithm.
[0107] In this embodiment of the invention, the linear interpolation algorithm can be the aforementioned nearest neighbor interpolation method, or the bilinear interpolation method. By using a linear interpolation algorithm to convert the RAW format image to be processed into an sRGB format image to be blurred, the interference of nonlinear interpolation algorithms on subsequent processing can be avoided, making the noise distribution in the resulting noisy RAW image closer to the true noise distribution.
[0108] In some embodiments, Figure 1 Based on this, see Figure 2 Following step S105, the method may further include the following steps:
[0109] S106: Input the noisy RAW image into the image denoising model of the initial structure to obtain the predicted denoised image.
[0110] S107: Calculate the loss function value representing the difference between the predicted denoised image and the image to be processed.
[0111] S108: Adjust the model parameters of the initial structure image denoising model based on the loss function value until the preset convergence condition is reached to obtain the trained image denoising model.
[0112] In this embodiment of the invention, the electronic device can input a noisy RAW image into an image denoising model with an initial structure. The image denoising model then performs image denoising processing on the noisy RAW image to obtain a predicted denoised image. Since noise is added to the RAW image with superimposed point diffusion effect to simulate adding noise according to the real noise distribution in the image to be processed, the closer the predicted denoised image is to the image to be processed after removing noise from the noisy RAW image, the better the image denoising effect.
[0113] Therefore, electronic devices can adjust the model parameters of an image denoising model based on the difference between the predicted denoised image and the image to be processed. For example, the electronic device can calculate a loss function value representing the difference between the predicted denoised image and the image to be processed based on a preset loss function, and then adjust the model parameters of the image denoising model based on the calculated loss function value. The preset loss function can be a cross-entropy loss function, an L1 (absolute error loss) function, etc.
[0114] In real-world scenarios, electronic devices can acquire multiple images to be processed. For each image to be processed, the electronic device can generate a noisy RAW image corresponding to that image according to the image generation method in any of the aforementioned embodiments. This results in multiple noisy RAW images. In this embodiment, the electronic device calculates a loss function value based on an image to be processed and its corresponding noisy RAW image. Then, based on the calculated loss function value, it adjusts the model parameters of the initial structure of the image denoising model. It can then continue to adjust the model parameters of the image denoising model in the same manner as described above, based on other images to be processed and their corresponding noisy RAW images, until the model converges. For example, the convergence condition can be: the model parameters of the image denoising model have been adjusted based on a preset number of noisy RAW images; or, the convergence condition can also be: after inputting a noisy RAW image into the image denoising model, the loss function value between the predicted denoised image and the corresponding image to be processed is less than a preset threshold. After model convergence, a trained image denoising model is obtained. Subsequently, image denoising can be performed based on the trained image denoising model to obtain denoising results with better denoising effect; other processing based on the denoising results with better denoising effect, such as image analysis, can improve the accuracy of the obtained processing results.
[0115] See Figure 3 , Figure 3 This is a structural diagram of an image denoising model provided in an embodiment of the present invention. Figure 3 In the diagram shown, a rectangle represents a network layer; the length of the rectangle represents the size of the network layer.
[0116] Figure 3 The black-filled rectangles in the diagram represent a type of network layer (which can be called the first network layer). The convolution kernel size is 3×3, the stride is 1, and it includes the ReLU (activation function), which can be written as conv3×3_stride=1+relu.
[0117] Figure 3 The rectangle filled with black horizontal stripes represents a type of network layer (which can be called the second network layer). The convolution kernel size is 3×3, the stride is 2, and it includes the ReLU (activation function), which can be written as conv3×3_stride=2+relu.
[0118] Figure 3 The rectangle filled with black vertical stripes represents a type of network layer (which can be called the third network layer) with a stride of 2, which can be denoted as deconv_stride = 2.
[0119] Figure 3The white-filled rectangle in the middle represents a type of network layer (which can be called the fourth network layer), with a convolution kernel size of 3×3 and a stride of 1, which can be denoted as conv3×3_stride=1.
[0120] like Figure 3 As shown, the image denoising model consists of multiple network layers.
[0121] The first group of network layers includes multiple first network layers; the second group of network layers includes one second network layer and multiple first network layers; the third group of network layers includes one second network layer and multiple first network layers. These three groups of network layers are used to downsample the input data (i.e., the aforementioned noisy RAW image).
[0122] The fourth and fifth network layers each consist of one fourth network layer and multiple first network layers; the sixth network layer consists of one fourth network layer, multiple first network layers, and one third network layer. These three network layers are used to upsample the output data of the third network layer.
[0123] Furthermore, the image denoising model includes multiple `add` statements, each representing a residual connection. For example, the input of the first network layer and the output of the sixth network layer are residually concatenated, meaning they are stitched together according to their corresponding channels, and the residual connection result is used as the output. Similarly, the outputs of the first and fifth network layers are residually concatenated, and the result is used as the input of the sixth network layer. Likewise, the outputs of the second and fourth network layers are residually concatenated, and the result is used as the input of the fifth network layer. Residual connections can address the vanishing gradient problem and improve the model's convergence speed.
[0124] Based on the above processing, the image generation method provided by this invention, that is, using PSF to synthesize noise in the RAW domain, yields a noisy RAW image with a noise distribution that is closer to the real noise distribution; the image denoising model is trained based on the noisy RAW image, that is, the denoising model is fitted and trained, and the trained image denoising model has a better denoising effect. Using the trained image denoising model to denoise the RAW image, the denoised image has higher clarity, especially for image details such as hair strands and green plant textures.
[0125] Based on the same inventive concept as the image generation method described above, embodiments of the present invention also provide an image generation apparatus. See [link to previous document]. Figure 4 , Figure 4 A structural diagram of an image generation apparatus provided in an embodiment of the present invention; the apparatus includes:
[0126] The acquisition module 401 is used to acquire the raw RAW format image to be processed;
[0127] The conversion module 402 is used to convert the image to be processed into the standard red-green-blue sRGB format based on the interpolation algorithm to obtain the image to be blurred;
[0128] The blur module 403 is used to perform Gaussian convolution blur on the pixel value of each pixel position in the image to be blurred for each color channel of the image to be blurred, so as to obtain the blurred pixel value of each pixel position for the color channel.
[0129] The sampling module 404 is used to obtain the blurred pixel value of the corresponding color channel to be used for each pixel position, so as to obtain a blurred RAW image; wherein, the color channel to be used corresponding to a pixel position is: the color channel at the pixel position in the image to be processed.
[0130] The noise-adding module 405 is used to add noise to the blurred RAW image to obtain a noisy RAW image.
[0131] Optionally, the device further includes:
[0132] The training module is used to add noise to the blurred RAW image after the noise-adding module 405 performs the process, obtain a noisy RAW image, and then input the noisy RAW image into an image denoising model with an initial structure to obtain a predicted denoised image. The module calculates a loss function value representing the difference between the predicted denoised image and the image to be processed. Based on the loss function value, the module adjusts the model parameters of the image denoising model with the initial structure until a preset convergence condition is reached to obtain a trained image denoising model.
[0133] Optionally, the acquisition module 401 is specifically used for:
[0134] Acquire RAW format images captured by the image acquisition device under normal illumination conditions;
[0135] Based on the images acquired by the image acquisition device, the image to be processed is determined.
[0136] Optionally, the acquisition module 401 is specifically used for:
[0137] The image acquired by the image acquisition device is subjected to preset processing to obtain the image to be processed; wherein, the preset processing includes at least one of the following: black level correction and pixel value normalization.
[0138] Optionally, the image acquisition device may use the minimum supported sensitivity when acquiring images.
[0139] And / or, the brightness value of the image acquired by the image acquisition device belongs to the range that is neither overexposed nor underexposed.
[0140] Optionally, the interpolation algorithm is a linear interpolation algorithm.
[0141] This invention also provides an electronic device, such as... Figure 5 As shown, it includes a processor 501, a communication interface 502, a memory 503, and a communication bus 504, wherein the processor 501, the communication interface 502, and the memory 503 communicate with each other through the communication bus 504.
[0142] Memory 503 is used to store computer programs;
[0143] The processor 501, when executing the program stored in the memory 503, implements the steps of any of the above-described image generation methods.
[0144] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0145] The communication interface is used for communication between the aforementioned electronic devices and other devices.
[0146] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0147] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0148] In another embodiment of the present invention, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the steps of any of the above-described image generation methods.
[0149] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the image generation methods described above.
[0150] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).
[0151] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0152] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, electronic devices, computer storage media, and computer program products are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0153] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.
Claims
1. An image generation method, characterized in that, The method includes: Obtain the raw RAW format image to be processed; Based on the interpolation algorithm, the image to be processed is converted into the standard red-green-blue sRGB format to obtain the image to be blurred; For each color channel of the image to be blurred, Gaussian convolution blur is performed on the pixel value of each pixel position in the image under that color channel to obtain the blurred pixel value of each pixel position for that color channel. For each pixel location, the blurred pixel value of the corresponding color channel to be used at that pixel location is obtained to obtain a blurred RAW image; wherein, the color channel to be used corresponding to a pixel location is: the color channel at that pixel location in the image to be processed; Noise is added to the blurred RAW image to obtain a noisy RAW image.
2. The method according to claim 1, characterized in that, After adding noise to the blurred RAW image to obtain a noisy RAW image, the method further includes: The noisy RAW image is input into the image denoising model with the initial structure to obtain the predicted denoised image; Calculate the loss function value representing the difference between the predicted denoised image and the image to be processed; The model parameters of the image denoising model with the initial structure are adjusted based on the loss function value until the preset convergence condition is met, thus obtaining the trained image denoising model.
3. The method according to claim 1, characterized in that, The process of obtaining the original RAW format image to be processed includes: Acquire RAW format images captured by the image acquisition device under normal illumination conditions; Based on the images acquired by the image acquisition device, the image to be processed is determined.
4. The method according to claim 3, characterized in that, Based on the images acquired by the image acquisition device, the image to be processed is determined, including: The image acquired by the image acquisition device is subjected to preset processing to obtain the image to be processed; wherein, the preset processing includes at least one of the following: black level correction and pixel value normalization.
5. The method according to claim 3, characterized in that, The image acquisition device uses the minimum supported sensitivity when acquiring images; And / or, The brightness value of the image acquired by the image acquisition device falls within the range of neither overexposure nor underexposure.
6. The method according to any one of claims 1-5, characterized in that, The interpolation algorithm is a linear interpolation algorithm.
7. An image generation apparatus, characterized in that, The device includes: The acquisition module is used to acquire the raw RAW format image to be processed; The conversion module is used to convert the image to be processed into the standard red-green-blue sRGB format based on the interpolation algorithm to obtain the image to be blurred; The blur module is used to perform Gaussian convolution blur on the pixel value of each pixel position in the image to be blurred for each color channel of the image to be blurred, so as to obtain the blurred pixel value of each pixel position for the color channel. The sampling module is used to obtain the blurred pixel value of the corresponding color channel to be used at each pixel position, so as to obtain a blurred RAW image; wherein, the color channel to be used corresponding to a pixel position is: the color channel at the pixel position in the image to be processed; The noise-adding module is used to add noise to the blurred RAW image to obtain a noisy RAW image.
8. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the method of any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method described in any one of claims 1-6.
10. A computer program product, characterized in that, When the computer program product is run on a computer, the computer causes the computer to perform the method according to any one of claims 1-6.
Citation Information
Patent Citations
Image data conversion method and device
CN113744167A
Remote sensing image cloud detection method, system and device and storage medium
CN113792653A
Noise image generation method and device, neural network training method and device, equipment and medium
CN114399563A
RAW domain image denoising method and shooting device
CN116723413A
Image processing method, device and equipment, readable storage medium and program product
CN117726522A