Information processing device, learning device, and information processing method
The information processing apparatus addresses the challenge of maintaining high image quality with low-bit-depth neural networks by converting high-bit-depth images, estimating noise component maps, and deriving noise-removed images, achieving efficient and high-quality image processing.
Patent Information
- Application Number
- JP2024198334
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-28
- Filing Date
- 2024-11-13
- Publication Date
- 2025-06-09
AI Technical Summary
Existing neural networks (NNs) with low bit depth struggle to maintain high image quality due to coarse gradation and reduced output accuracy, especially when processing high-quality images like RAW images with 12 to 14 bits.
An information processing apparatus that converts an input image with a high bit depth into a low-bit-depth image, uses a neural network with a low bit depth to estimate a noise component map, and then derives a noise-removed image by subtracting the noise component map from the original image.
This approach enables the estimation of high-quality images using NNs with low bit depth, maintaining high noise removal performance and image quality, while reducing computational requirements.
Smart Images

Figure 2025086879000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to image processing technology for high image quality.
Background Art
[0002] In recent years, in high image quality processing for improving the image quality of images, various methods using neural networks (NNs) have been developed. Here, high image quality processing refers to image processing such as noise removal, aberration correction, and demosaicking. In methods using an NN, those with higher image processing performance tend to have a larger computational amount. Therefore, in order to enable processing on embedded devices, lightweight methods for reducing the computational amount while maintaining performance have been actively studied. In Non-Patent Document 1 and Non-Patent Document 2, methods for reducing the weight by quantizing the weights and feature amounts of the NN to a lower bit depth have been proposed.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Non-Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in quantization, when quantizing in a simple way such as thinning out values at equal intervals, the output accuracy deteriorates compared to before quantization. In the NN used for high image quality, when quantizing the weights and feature amounts of the NN to a low bit depth (for example, a bit depth lower than the bit depth of the image to be output), the gradation of the output from the NN becomes coarse and the output accuracy deteriorates. For example, when high-qualityizing a RAW image, since the bit depth of the RAW image is 12 to 14 bits, it is desirable to use an NN that originally has a bit depth of 12 to 14 bits or more. When using an NN in which the weights and feature amounts are quantized to a bit depth of 8 bits, the image output by the NN also has 8-bit gradation, which becomes coarser than the gradation of the image originally to be estimated. Therefore, when using an NN with a low bit depth, there is a problem that the high-qualityization performance deteriorates compared to an NN having a bit depth higher than the image to be output.
[0005] The present invention has been made in view of such problems, and an object thereof is to provide a technique that enables estimation of a high-quality image using an NN having a low bit depth.
Means for Solving the Problems
[0006] To solve the above problems, an information processing apparatus according to the present invention has the following configuration. That is, the information processing apparatus conversion means for converting an input image with a first bit depth into a low-bit-depth image with a second bit depth lower than the first bit depth; estimation means for estimating a noise component map in the input image from the low-bit-depth image using a neural network (NN) with a third bit depth lower than the first bit depth and equal to or higher than the second bit depth; derivation means for deriving a noise-removed image corresponding to the input image based on the input image and the noise component map; and includes.
Effects of the Invention
[0007] According to the present invention, it is possible to provide a technique for estimating a high-quality image using an NN having a low bit depth.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Modes for Carrying Out the Invention
[0009] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the invention according to the claims. Although a plurality of features are described in the embodiments, not all of these plurality of features are essential to the invention, and the plurality of features may be arbitrarily combined. Further, in the accompanying drawings, the same or similar configurations are denoted by the same reference numerals, and redundant descriptions are omitted.
[0010] (First Embodiment) As a first embodiment of the information processing apparatus according to the present invention, an information processing apparatus that performs high-quality processing using a neural network (NN) will be described below as an example.
[0011] <Overview> The present invention relates to a process of estimating a high-quality image from a low-quality image by machine learning. Regarding high-quality conversion from a low-quality image, there are, for example, noise removal (denoising) processing and aberration correction processing.
[0012] In the first embodiment, an inference process using an NN for noise removal and a learning method of the NN for noise removal will be described. The bit depth of the image to be processed is 14 bits, and the bit depth of the weights and intermediate feature amounts of the NN (hereinafter referred to as "bit depth of the NN") is 8 bits. However, the bit depth is not limited to these. Also, the type of the image to be processed may be a RAW image (for example, a mosaic image of a Bayer array) or an RGB image (a demosaicked image).
[0013] In the first embodiment, instead of directly estimating a high-quality image (denoised image) by the NN, a noise component is estimated by the NN. Then, the denoised image is derived by subtracting the estimated noise component from the noisy image. This is because, as will be described below, the variation range of the noise component is smaller than the range of values that the pixel values of the image can take, so the noise component can be expressed with relatively high accuracy even with an 8-bit depth.
[0014] FIG. 8 is a diagram for explaining the relationship between pixel values and noise components. Specifically, it is a diagram exemplarily showing the distribution (variation) of noise generated in an image (14-bit RAW image) captured by a certain image sensor. The horizontal axis represents pixel values, and the vertical axis represents noise. The curve shown in FIG. 8 indicates a curve corresponding to 2σ (σ is the standard deviation of the values that the noise can take) with respect to the pixel values. From FIG. 8, it can be seen that the variation of the noise increases as the pixel value increases (= the pixel with a larger number of bits).
[0015] Each point on the graph is generated by generating noise according to a normal distribution with σ corresponding to each pixel value and plotted. Here, even when the pixel value is the maximum value of 16383 (= 2 14 - 1), it can be seen that 2σ is about 512. This means that the noise component values are within about ±512 for about 90% of the pixels with a pixel value of 16383. That is, it shows that the noise component in the 14-bit RAW image can be sufficiently represented even with 10 bits (= 2 10 - 1 gradations).
[0016] Since the finally obtained denoised image is a 14-bit RAW image, when the 8-bit depth NN directly estimates the denoised image, it is necessary to convert 14 bits to 8 bits. On the other hand, when the 8-bit depth NN estimates the noise component, since 10 bits are converted to 8 bits, the error generated by quantization is smaller than when directly estimating the denoised image. Therefore, it can be expected that the denoised image obtained by estimating the noise component by the 8-bit depth NN and subtracting it from the noisy image will be a higher-quality image than the denoised image directly estimated by the 8-bit depth NN.
[0017] Also, although the noise has a larger σ as the pixel value increases, the ratio of the noise to the pixel value is larger as the pixel value is smaller. Therefore, it is important for image quality to accurately estimate the small absolute value noise that occurs in the region where the pixel value is small in the estimation of the noise component.
[0018] Hereinafter, first, the hardware configuration of the information processing apparatus will be described. After that, the functional configuration and operations in the inference process and the learning process will be described respectively.
[0019] <Hardware Configuration> FIG. 1 is a diagram showing the hardware configuration of the information processing apparatus according to the first embodiment. Note that the same information processing apparatus may be used for the inference process and the learning process, or different information processing apparatuses may be used.
[0020] The CPU 101 controls the entire device by executing the control program stored in the ROM 102. The RAM 103 temporarily stores various data from each component. Also, the RAM 103 functions as a work area for the CPU 101, and the control program is expanded in the RAM 103 in a state executable by the CPU 101. The storage unit 104 stores various data to be processed in this embodiment. For example, it stores images to be subjected to inference processing (noise removal processing), images used in learning processing, and various parameters. As the medium of the storage unit 104, an HDD, a flash memory, various optical media, etc. can be used.
[0021] <Functional configuration during inference processing> Figure 2(a) is a diagram showing the functional configuration of the information processing apparatus during inference. The information processing apparatus 1 includes a storage unit 201, an image acquisition unit 202, an image quantization unit 203, a difference estimation unit 204, and a high-quality image estimation unit 205. Each functional configuration unit will be briefly described.
[0022] The image acquisition unit 202 acquires an input image (an image with a bit depth of 14 bits) to be subjected to noise removal processing from the storage unit 201. Hereinafter, this image will be referred to as a "noisy image". The noisy image is an image in which a "noise component" is added to the original image. The noise component is caused by, for example, an imaging unit (such as an image sensor). Hereinafter, the original image will be referred to as a "clean image". As described above, in this inference processing, the noise component is estimated from the noisy image using the NN, and the clean image is derived by subtracting the noise component from the noisy image. Note that there may be cases where a clean image exactly the same as the original image cannot be derived, but for convenience, it is referred to as a "clean image".
[0023] The image quantization unit 203 performs quantization processing on the noisy image with a bit depth of 14 bits obtained from the image acquisition unit 202, and converts it into a noisy image (low bit-depth image) in which each pixel is represented by an unsigned 8-bit integer. In this embodiment, the same uniform quantization method as the bit-depth conversion layer 306 described later is used. However, the quantization method is not limited to this. For example, the image quantization unit 203 may include an NN with a bit depth of 14 bits or more. In that case, after inputting the 14-bit noisy image into the NN, quantization processing is performed on the output of the NN to obtain an 8-bit noisy image. Here, although the bit depth of the NN and the bit depth of the low bit-depth image are made to match (8 bits), they may be different. The bit depth of the NN may be lower than the bit depth of the input image and equal to or higher than the bit depth of the low bit-depth image.
[0024] The difference estimation unit 204 inputs the 8-bit noisy image 301 obtained from the image quantization unit 203 into an 8-bit NN, and estimates a difference map (noise component map) in which each pixel has 8-bit gradation but the value range is represented by a signed 10-bit integer.
[0025] The high-quality image estimation unit 205 derives a 14-bit clean image that is an image with noise removed. Specifically, it is derived by subtracting the noise component map, in which each pixel estimated by the difference estimation unit 204 has 8-bit gradation but the value range is represented by a signed 10-bit integer, from the 14-bit noisy image obtained from the image acquisition unit 202.
[0026] FIG. 3(a) is a diagram for explaining the structure of a difference estimation NN having an 8-bit bit depth. The intermediate layer exists from the first intermediate layer 302-1 to the nth intermediate layer 302-n, and finally the final layer 303 exists. The intermediate layer is an NN in which the weights are signed 8-bit integers and the output is an unsigned 8-bit integer. The final layer 303 is an NN in which the weights are signed 8-bit integers and the output has 8-bit gradation but the value range is a signed 10-bit integer.
[0027] The differential estimation NN composed of the intermediate layer and the final layer 303 takes the 8-bit noiseless image 301 without sign as the input to the first intermediate layer 302-1, and the final layer 303 outputs the differential estimation value 309 which is 8-bit grayscale but the value range is 10-bit integer with sign. Here, the differential estimation value refers to the estimated value 309 of the noise component map. Here, the number of intermediate layers can be any number.
[0028] The internal configurations of the first intermediate layer 302-1 to the nth layer 302-n are common, and the internal configuration within each intermediate layer will be described using the intermediate layer 302-1 as a representative example.
[0029] The intermediate layer 302-1 is composed of a convolutional layer 304-1, a ReLU layer 305-1, and a bit-depth conversion layer 306-1.
[0030] The convolutional layer 304-1 performs a convolution process which is a linear transformation with signed 8-bit integer weights. In the convolution process, since the weights with signed 8-bit integers (including biases) and the noiseless image 301 without sign are multiplied, the result of the operation becomes a signed 16-bit integer.
[0031] The ReLU layer 305-1 performs a ReLU (Rectified Linear Unit) process which is a non-linear transformation. Since ReLU outputs 0 for values of 0 or less, the intermediate features of the input signed 16-bit integer become unsigned 15-bit integers by ReLU.
[0032] The bit-depth conversion layer 306-1 performs a process of converting the unsigned 15-bit integer data obtained by the ReLU layer 305-1 into unsigned 8-bit integers. For the conversion of the bit depth, in this embodiment, a method of uniformly quantizing 15 bits to 8 bits is adopted, but a non-uniform quantization method represented by Non-Patent Document 2 may also be used. The details of the process in the bit-depth conversion layer 306-1 will be described later with reference to Fig. 3(b).
[0033] Next, the configuration of the final layer 303 will be described. The final layer 303 is composed of a convolutional layer 307 and a final bit-depth conversion layer 308.
[0034] Similar to the convolutional layer 304, the convolutional layer 307 performs a convolutional process with weights of 8-bit integers.
[0035] The final bit-depth conversion layer 308 converts a noise component map of signed 16-bit integers into a noise component map with a value range of signed 10-bit integers at 8-bit grayscale. The conversion from 16-bit grayscale to 8-bit grayscale uses a non-uniform quantization method represented by Non-Patent Document 2. Here, non-uniform quantization is a method of devising the tone representation when thinning out, finely representing the range effective for accuracy in the input, and reducing the quantization error, and an improvement in the accuracy of quantization NN can be expected. In the final bit-depth conversion layer 308, the non-uniform quantization method is devised to accurately quantize the noise components effective for image quality improvement. Details of the processing in the final bit-depth conversion layer 308 will be described later with reference to Fig. 3(c).
[0036] Note that the structure of the NN is not limited to that shown in Fig. 3(a), and a U-Net structure or the like may be used. The convolutional layers 304 and 307 and the ReLU layer 305 are not limited to these either, and other linear transformations and non-linear transformations can be used. Also, the type and number of layers of each layer in the intermediate layer 302 are not limited, and they do not have to be the same as the final layer 303. Furthermore, the bit-depth of the noisy image 301 may be greater than 8 bits.
[0037] <Operation during inference processing> Fig. 5 is a flowchart of the inference processing executed by the information processing apparatus. However, the information processing apparatus does not necessarily perform all the steps described in this flowchart.
[0038] In S501, the image acquisition unit 202 acquires a noisy image to be subjected to noise removal from the storage unit 201. Here, it is assumed that the noisy image is a RAW image and each pixel is an unsigned 14-bit integer.
[0039] In S502, the image quantization unit 203 converts the noise-free 14-bit integer noisy image acquired in S501 into a noise-free 8-bit integer noisy image 301.
[0040] In S503, the difference estimation unit 204 obtains an estimated value of a noise component map that is 8-bit grayscale but whose value range is represented by a signed 10-bit integer from the noise-free 8-bit integer noisy image 301 obtained in S502.
[0041] More specifically, the difference estimation unit 204 inputs the noise-free 8-bit integer noisy image 301 obtained in S502 to the difference estimation NN in Fig. 3(a) and sequentially performs the processing of each intermediate layer 302 and the final layer 303. As a result, a difference estimation value 309 (noise component map) that is 8-bit grayscale and whose value range is a signed 10-bit integer is output. Here, the case where the biases of the convolutional layers 304 and 307 are "0" and the weights are represented by "signed 8-bit integers" will be described.
[0042] At this time, as a result of the convolution operation between the signed 8-bit integer weights of the convolutional layers 304 and 307 and the noise-free 8-bit integer noisy image 301 or the intermediate features, the obtained output is an intermediate feature of a signed 16-bit integer. When the Relu layer 305 is applied to the output of the convolutional layer 304, negative values are converted to "0" and positive values are output as they are, so the obtained output is represented by an unsigned 15-bit integer. The bit depth conversion layer 306 converts the unsigned 15-bit integer obtained by the ReLU layer 305 into an unsigned 8-bit integer. In the final bit depth conversion layer 308, the signed 16-bit integer obtained by the convolutional layer 307 is converted into an integer with a value range of a signed 10-bit integer in 8-bit grayscale.
[0043] Fig. 3(b) is a diagram for explaining the processing in the bit depth conversion layer 306. This processing is to convert an unsigned 15-bit integer input into an unsigned 8-bit integer.
[0044] In S311, the 15-bit unsigned integer output by the ReLU layer 305 is normalized. Specifically, the processing as shown in Equation (1) is performed on the intermediate feature x output by the ReLU layer 305.
[0045]
Number
[0046] Here, β is 2 15 -1. By this processing, the output becomes a 15-bit grayscale real value having a value range of [0, 1]. Also, in this embodiment, β is 2 15 -1 for normalization. However, x may be clipped with arbitrary minimum and maximum values, and normalized with the difference between the aforementioned minimum and maximum values to obtain a real value having a value range of [0, 1] and a grayscale less than 15 bits. inter
[0047] In S312, the normalized intermediate feature obtained by S311 is converted into an 8-bit unsigned integer value. Specifically, the processing as shown in Equation (2) is applied to the output of S311.
[0048]
Number
[0049] Here, s inter is 2 8 -1, and the right parenthesis represents the process of rounding off the decimal part. After setting the scale of the real value in the range of [0, 2 8 -1], the decimal part is rounded off to obtain an 8-bit unsigned integer value. By this processing, the 15-bit unsigned integer value output by the ReLU layer 305 is converted into an 8-bit unsigned integer. In this embodiment, the processing of the bit depth conversion layer 306 is a uniform quantization method that does not perform non-linear processing when quantizing, but a non-uniform quantization method as described in Non-Patent Document 2 may also be used.
[0050] FIG. 3(c) is a diagram for explaining the processing in the final bit depth conversion layer 308. This processing is to convert an input represented in a certain bit depth into a different bit depth. At this time, tone conversion is performed so that non-uniform tone expression (tones in a certain range are expressed finely, and tones in other ranges are expressed coarsely) is performed. This tone conversion corresponds to a non-uniform quantization method as performed in Non-Patent Document 2.
[0051] In S321, the final bit depth conversion layer 308 normalizes the intermediate features obtained in the convolutional layer 307. Here, let the intermediate features be x, and assume that x is a map with a width W, a height H, and a channel number of 1. The normalization is, for example, a process of taking the absolute value of the intermediate features as in Equation (3), clipping so that the value becomes α or less, and then normalizing to the range of [0, 1].
[0052]
Equation
[0053] Here, the parameter α of the clipping range corresponds to 2σ of the noise distribution as described above, and is set to -1. The parameter α can also be determined by 3σ or the like, or can be optimized from a plurality of candidates by Bayesian optimization or the like so that the image quality of the prepared evaluation image improves. At that time, a general quantification index such as PSNR can be used as an index of the image quality to be optimized, but it is not limited to this. 9 -1. The parameter α can also be determined by 3σ or the like, or can be optimized from a plurality of candidates by Bayesian optimization or the like so that the image quality of the prepared evaluation image improves. At that time, a general quantification index such as PSNR can be used as an index of the image quality to be optimized, but it is not limited to this.
[0054] In S322, the final bit depth conversion layer 308 applies a non-linear transformation f Θ to the normalized intermediate features obtained in S321.
[0055]
Equation
[0056] FIG. 4 is a diagram for explaining the non-linear conversion (S322) process in the bit depth conversion process. In the present embodiment, a case where the non-linear conversion can be represented by a tone curve as shown in FIG. 4(a) will be described. First, the normalized map x' obtained in S321 is input to the tone curve to obtain a non-linearly converted map. The tone curve is converted so that the tones of low values become finer, and the tones become coarser as the value increases. The value range of the non-linearly converted map is [0,1], and it takes a 9-bit real value.
[0057] In S323, the final bit depth conversion layer 308 converts the output of S322 into an unsigned 7-bit integer. Specifically, the following formula (5) is used.
[0058]
Equation
[0059] Here, s 1 = 2 7 -1, and the right parenthesis represents the process of rounding off the decimal part. After setting the scale of the 7-bit real value to the range of [0, 2 7 -1], an unsigned 7-bit integer value is obtained by rounding off the decimal part. Since the absolute value of x is taken in S321, an unsigned 7-bit integer value is obtained here instead of an 8-bit value.
[0060] Also, although the 16-bit data has been converted to 7 bits in the processes of S321 to S323 so far, since the parameter α (2 9 -1) is clipped in S321, it is actually 9-bit data that is being converted to 7 bits. Furthermore, by performing non-linear processing in S322, noise with a small absolute value that greatly contributes to the image quality is converted with finer tones, thereby suppressing the degradation of the image quality that occurs when converting to a low bit.
[0061] In S324, the 7-bit integer value obtained in S323 is normalized again. The normalization coefficient is s of S323 1Using the same value, the value range of the normalized map is set to a 7-bit real value in the range of [0, 1].
[0062]
Number
[0063] In S325, the final bit-depth conversion layer 308 applies the inverse transformation f Θ -1 of the non-linear transformation used in S322 to the output obtained in S323. Applying f Θ -1 linearizes the values non-linearized in S322. The value range of the linearized map is [0, 1], and it takes 7-bit real values.
[0064]
Number
[0065] In S326, the 7-bit real value output in S325 is converted into an 8-bit grayscale signed 10-bit integer. Specifically, the following mathematical formula (8) is used.
[0066]
Number
[0067] Here, s 2 = 2 9 - 1, and the parentheses on the right side represent the process of rounding off the digits after the decimal point. After setting the scale of the real value in the range of [0, 2 9 - 1] and then rounding off the digits after the decimal point, an integer value of 7-bit grayscale in the range of [0, 2 9 - 1] is obtained. Since sign(x) outputs the sign of x, the finally obtained value is an integer value of 8-bit grayscale with a value range of [-2 9 , 2 9 - 1].[[]END]
[0068] Since the differential estimation NN represents the input noisy image 301 and the weights and feature amounts of the intermediate layer in 8 bits, it is difficult to accurately infer tones of 9 bits or more as the final output with a high-speed model. Therefore, like the processes of S321 to S326, after applying non-linear processing to convert to a low bit depth and then applying inverse non-linear processing to return the value range to the original bit depth. By this process, while reducing the tones of the noise components to a low bit depth, noise with a small absolute value that greatly contributes to the image quality can be represented with finer tones, suppressing the deterioration of the image quality caused by low-bit tone conversion. Note that in this embodiment, the processes of S321 to S323 that perform low-bit conversion by non-linear conversion are converted to 8 bits, but it is not limited to 8 bits, and any bit depth below the parameter α clipped in S321 may be used.
[0069] The original noise components have a value range of [-2 14 , 2 14 -1] and are 15-bit integer values, and the noise components estimated in S326 are 8-bit tone integer values with a value range of [-2 9 , 2 9 -1]. The original noise components and the estimated noise components only have different possible value ranges, and the estimated noise components may also be treated as data having a 15-bit bit depth. That is, when later subtracting the noise components from the noisy image, the subtraction may be performed with the original numerical values.
[0070] Also, the process combining S321 to S326 may be realized by performing arithmetic operations, or may be realized using a look-up table (LUT) as shown in FIG. 4(b). Thereby, these processes can be speeded up. In this LUT, regions with a small absolute value of noise are converted with fine tones, and as the absolute value of the noise increases, they are converted with coarse tones. When using the LUT, the input x is clipped by the positive and negative of the parameter α and then converted by the LUT. By using the LUT shown in FIG. 4(b), it is possible to represent a range of values where the influence of quantization error is relatively large in the noise components with relatively fine tones.
[0071] In S504, the high-quality image estimation unit 205 subtracts the estimated value of the noise component map obtained in S503 from the 14-bit noisy image obtained in S501. As a result, an estimated value of a denoised image, which is an image with noise removed from the noisy image, is derived.
[0072] <Functional configuration during learning process> In the present embodiment, it is assumed that learning is performed in the framework of pseudo quantum chemistry learning as in Non-Patent Document 1. In pseudo quantum chemistry learning, the weights and intermediate features of the model are, unlike during inference, pseudo-quantized to 8-bit gradation using floating-point numbers expressed instead of integers. When calculating the loss during forward propagation, values quantized to 8-bit gradation are used, and when performing backpropagation, values before quantization such as 32 bits are used, enabling minute updates of the parameters and reducing the error during inference. After the model is trained in the framework of pseudo quantum chemistry learning, it is converted to an integer using a parameter integerization unit 209 described later and used during inference.
[0073] FIG. 2(b) is a diagram showing the functional configuration of the information processing apparatus during learning. The information processing apparatus 1 includes a storage unit 201, a learning data acquisition unit 206, an image quantization unit 203, a difference estimation unit 204, an error calculation unit 207, a parameter update unit 208, and a parameter integerization unit 209. Since the storage unit 201 and the image quantization unit 203 are the same as those in the configuration during inference (FIG. 2(a)), the description thereof is omitted.
[0074] The learning data acquisition unit 206 acquires a clean image, which is an ideal image without noise, from the storage unit 201. Then, by adding artificially generated noise components to the clean image, a noisy image, which is an image to be subjected to noise removal, is generated. The clean image and the noisy image have a 14-bit depth. Note that when generating the noisy image, the portion exceeding the 14-bit upper limit value due to the addition of noise components is clipped.
[0075] The difference estimation unit 204 acquires the model of the difference estimation NN from the storage unit 201. Then, it inputs the 8-bit deep noisy image obtained from the image quantization unit 203 into the 8-bit deep NN, and estimates a noise component map with an integer value range of signed 10 bits in 8-bit gradation.
[0076] The weights and intermediate features of the difference estimation NN model are pseudo-quantized to 8-bit gradation and used, which are represented by floating-point numbers instead of integers, different from the inference time.
[0077] The error calculation unit 207 calculates the loss with respect to the estimation result of the noise component map. Specifically, it calculates the error between the estimated value of the noise component map with an integer value range of signed 10 bits in 8-bit gradation estimated by the difference estimation unit 204 and the GT (Ground Truth) obtained from the learning data acquisition unit 206. The specific calculation method will be described later.
[0078] The parameter update unit 208 updates the parameters of the difference estimation NN shown in Fig. 3(a) based on the error obtained by the error calculation unit 207, and stores them in the storage unit 201.
[0079] The parameter quantization unit 209 quantizes the weights and outputs of the difference estimation NN with pseudo-quantization learning and converts them into integers. For details, a known NN quantization method may be applied, and the description is omitted. As a result, the same output can be obtained before and after conversion to integers.
[0080] <Operation during learning process> Fig. 6 is a flowchart of the learning process of the NN executed by the information processing device. However, the information processing device does not necessarily perform all the steps described in this flowchart.
[0081] In S601, the learning data acquisition unit 206 obtains, from the storage unit 201, a clean image that is an ideal image without noise, and a GT of a noise component map having the same size as the clean image and to be added to the clean image. The noise component map may be generated, for example, by calculating the noise intensity with a function (or table) that takes the luminance of the clean image as an input. Then, a noisy image is obtained by adding each pixel of the noise component map and the clean image. Here, the noisy image is a RAW image, and the bit depth is 14 bits.
[0082] In S602, the image quantization unit 203 converts the 14-bit depth noisy image obtained in S501 into an 8-bit depth noisy image and outputs it.
[0083] In S603, the difference estimation unit 204 obtains an estimated value of the noise component map with a 14-bit depth in the same procedure as S503. That is, a noise component map with an integer value range of signed 10 bits in 8-bit gradation is estimated from the 8-bit depth noisy image obtained in S502.
[0084] In S604, the error calculation unit 207 calculates a loss Loss 1 with respect to the estimation result of the noise component map. The purpose is to advance the learning so that the clean image, which is the difference between the noisy image and the noise, can be correctly estimated by correctly estimating the noise components in the noisy image. In this embodiment, as shown in Equation (9), Loss 1 is calculated as the L1 distance. The L1 distance is the sum of the absolute values of the differences of each element between the estimation result C inf of the noise component map obtained from S603 and the noise component map C gt that is the GT obtained in S601. However, the type of loss is not limited to this.
[0085]
Number
[0086] In S605, the parameter update unit 208 updates the parameters of the NN using the error backpropagation method based on the loss Loss calculated in S604. The parameters to be updated here refer to the weights of the convolutional layers 304 and 307 that make up the NN shown in Fig. 3(a). 1 Here, the parameters to be updated refer to the weights of the convolutional layers 304 and 307 that make up the NN shown in Fig. 3(a).
[0087] In S606, the parameter update unit 208 stores the updated parameters of the NN in the storage unit 201. Then, the weights are loaded into the NN. S601 to S606 are regarded as one iteration of learning.
[0088] In S607, the parameter update unit 208 determines whether to end the learning. The end determination of the learning may be performed, for example, by detecting that the value of the loss obtained by Equation (9) is smaller than a predetermined threshold. Also, it may be determined to end when learning has been performed a predetermined number of times. Note that when the learning loss converges and the learning ends, the parameter quantization unit 209 quantizes the NN into integers.
[0089] As described above, according to the first embodiment, during the inference process, the noise component is estimated by an NN with a bit depth lower than the bit depth of the image to be processed. Then, the denoised image is derived by subtracting the estimated noise component from the noisy image. At this time, the clip value of the noise component is set according to the noise model. Thereby, in the high-quality processing using an NN with a low bit depth, it becomes possible to maintain high noise removal performance. Also, in the final layer of the NN, by applying the non-uniform quantization method, the noise component can be accurately expressed.
[0090] (Modification 1) In Modification 1, a form in which a piecewise linear function is used in the final bit depth conversion layer 308 of the final layer 303 will be described. That is, as the non-linear transformation f Θ a piecewise linear function is used. By using the piecewise linear function, it becomes possible to more freely set which range of tones of the input is to be made finer.
[0091] Note that, as in Non-Patent Document 2, a function that defines the slopes of each section divided at equal intervals may be used as the piecewise linear function. In this case, the larger the slope of a section, the finer the gradation will be represented.
[0092] FIG. 7 is a diagram for explaining the piecewise linear function used in the bit depth conversion process of the final bit depth conversion layer 308. This piecewise linear function has five sections that divide the input domain [0, 1] at equal intervals, and the slope γ i (i = 1 to 5) of which, the slope γ 2 of the second section is the largest. When this piecewise linear function is used, the noise component map output by the final bit depth conversion layer 308 becomes a map in which the gradation in the range of the second section is represented in the finest detail.
[0093] Alternatively, when a function obtained by piecewise linearly approximating the tone curve of the first embodiment is used, the output finally obtained from the final bit depth conversion layer 308 is converted such that the gradation for a small input is fine and the gradation for a large input is coarse. The slopes of each section of the piecewise linear function may be obtained by, for example, Bayesian optimization, or a plurality of candidates may be determined and optimized so that the image quality of a prepared evaluation image improves. In that case, a general quantification index such as PSNR may be used as the image quality index for optimization.
[0094] Also, as in Non-Patent Document 2, the parameters of the piecewise linear function may be learned by the error backpropagation method. Also, the slopes of each section of the piecewise linear function may be determined in consideration of the relationship between the magnitude of the noise component of a certain pixel and the degree of influence on the image quality of that pixel (such as the N / S ratio). For example, when a graph with the magnitude of the noise component on the horizontal axis and the image quality index on the vertical axis (referred to as a noise component - image quality index graph) has a maximum value instead of being monotonically increasing, the gradation in the range near the noise component that gives the maximum value may be converted finely.
[0095] As described above, according to Modification 1, the non-linear transformation f ΘBy using a piecewise linear function, the degree of freedom in shape increases, and the degree of freedom in tone expression is higher than that in the first embodiment. As a result, it becomes possible to effectively suppress image quality degradation due to quantization. Further, by using the method disclosed in Non-Patent Document 2, parameters such as the slope of the piecewise linear function can be learned by error backpropagation together with the weights of the NN, and an optimal tone expression for image quality improvement can be efficiently obtained.
[0096] In the above-described Modification 1, when learning the piecewise linear function and the NN weights, in S604, the error calculation unit 207 calculates the loss Loss with respect to the estimation result of the noise component map 1 as follows. Specifically, for the loss that brings C inf obtained in S603 and C gt acquired in S601 gt close to each other, a weighting map w having the same width and height as the clean image used for the generation of C inf and C gt and having different values for each pixel is prepared. Then, pixel-by-pixel weighting is performed on the loss for bringing C inf and C gt close to each other. An example when the L1 distance is used for the loss is shown in Equation (10).
[0097]
Equation
[0098] Here, the weighting map w i may be determined according to the relationship between the image quality index and the pixel value I. For example, when the image quality index is represented by a function g(I) of the pixel value I, each pixel value of the clean image acquired from the storage unit 201 in S601 may be input to the function g(I) to obtain a map having the same width and height. Further, a map obtained by normalizing the value of the obtained map by dividing it by the maximum value of the map may be used as the weighting map.
[0099] For example, when the graph with the pixel value on the horizontal axis and the image quality index g(I) on the vertical axis is not monotonically increasing and has a maximum value, pixels having pixel values closer to the maximum value of the graph are weighted more in the loss calculation of Equation (10). iIt becomes larger. Therefore, learning for these pixels proceeds preferentially. As a result, in learning the weights of the NN and the parameters of the non-linear transformation, learning is promoted to improve the image quality in the region where noise that affects the image quality is prominent.
[0100] As described above, according to the modification example, the loss is weighted so that the higher the pixel value with a greater contribution to the image quality, the higher the noise estimation accuracy. As a result, it is possible to focus on improving the denoising accuracy in the region with a high image quality improvement effect.
[0101] (Modification Example 2) In Modification Example 2, during the learning process, S322 of the final bit-depth conversion layer 308 that constitutes the final layer 303 is replaced with an identity mapping so that non-linear transformation is implicitly performed within the NN. That is, unlike the first embodiment, the non-linear transformation of S322 is not explicitly performed. As a result, during the inference process, it is possible to avoid an increase in the processing load due to non-linear transformation and accurately represent the noise component even with a small number of gradations, and an improvement in denoising accuracy can be expected. Hereinafter, the parts different from the processing of the first embodiment will be described.
[0102] <Operation during learning process> In S601, the learning data acquisition unit 206 acquires, from the storage unit 201, a clean image that is an ideal image without noise and a noise component map having the same size as the clean image and to be added to the clean image. Then, a noisy image is acquired by adding each pixel of the noise component map and the clean image. Here, the noisy image is a RAW image and the bit depth is 14 bits.
[0103] In S603, the difference estimation unit 204 obtains an estimated value of the noise component map with an integer value range of signed 10 bits in 8-bit gradation in the same procedure as S503. However, in this embodiment, when performing the processing of the final bit-depth conversion layer 308 of the difference estimation NN in S503, the non-linear transformation applied in the non-linear transformation process of S322 is replaced with an identity mapping. Also, the processing of S324 to S326 is performed only during the inference process and not during the learning process.
[0104] In S604, the error calculation unit 207 calculates the loss with respect to the estimation result of the noise component map. 1 The GT of the noise component map used when calculating Loss 1 is subjected to non - linear transformation in advance and converted into a signed 8 - bit integer. Specifically, the noise component map obtained in S601 is subjected to non - linear transformation of the noise component map and conversion into a signed 8 - bit integer in the same manner as the processing of S321 - S323. This is used as the GT of the noise component map.
[0105] The type of non - linear transformation may be a tone curve as used in the first embodiment, but is not limited thereto. The loss Loss 1 is defined such that it becomes smaller as the estimated value of the noise component map obtained in S603 approaches the GT of the noise component map. For example, the L1 distance, which is the sum of the absolute value differences of each element, may be calculated as in the first embodiment, but the type of loss is not limited thereto.
[0106] <Operation during inference processing> In S503, the difference estimation unit 204 changes the processing in the final bit - depth conversion layer 308 of the difference estimation NN. Specifically, the processing of S322 performed in the first embodiment is not carried out. This is because, by performing the above - mentioned learning processing of this embodiment, the NN is learned so that the result of non - linear transformation is directly output at the start time of FIG. 3(c).
[0107] Also, as described above, the processing of S324 - S326 that was not performed in the learning processing is performed during the inference processing.
[0108] As described above, according to Modification Example 2, during the learning processing, the non - linear transformation is configured to be implicitly performed within the NN in the final bit - depth conversion layer 308. Thereby, during the inference processing, while avoiding an increase in processing load due to non - linear transformation, the noise components can be accurately represented with a small number of gradations, and an improvement in denoising accuracy can be expected. (Modification Example 3) In Modification 3, a method of obtaining 8-bit unsigned values using a non-uniform quantization method by applying non-linear processing to a 14-bit noisy image in the image quantization unit 203 will be described.
[0109] In the difference estimation unit 204, in order to represent noise with small absolute values that greatly contribute to image quality with finer gradations, it is desirable that the input data be converted to 8 bits in a suitable state. More specifically, it is desirable to represent regions with low luminance values where the ratio of noise to pixel values is large with finer gradations.
[0110] FIG. 9(a) is a flowchart of the processing of the image quantization unit 203 in the present embodiment.
[0111] In S901, a 14-bit noisy image is normalized. Specifically, the process of Equation (11) is performed on the 14-bit noisy image.
[0112]
Equation
[0113] Here, γ is 2 14 -1. By this process, the output becomes a 14-bit gradation real value having a value range of [0, 1].
[0114] In S902, non-linear conversion f Φ is applied to the normalized noisy image acquired in S901.
[0115]
Equation
[0116] FIG. 9(b) is the non-linear conversion f in the present embodiment Φis a diagram. The non-linear conversion is performed so that the gradation near the black level (OB level) becomes finer. The black level is a value that serves as a reference for black within the range of 14 bits. Pixel values below the black level are finally determined to be black. The image is converted into a digital signal by an image sensor, but if the amount of negative noise generated by the image sensor is large, the pixel values of the subject in the low-luminance part may fall below the black level. When estimating the noise component from a noisy image, low-luminance pixels with a large ratio of noise to the pixel value are important for image quality, and in the input image, the vicinity of the black level corresponds to this. Therefore, it is important to convert pixel values close to the black level into finer gradations. The black level in this embodiment is set to 2048. The value range of the non-linearly converted noisy image is [0, 1], and it takes 14-bit real values.
[0117] In S903, the non-linearly converted noisy image acquired in S902 is converted into an unsigned 8-bit integer value. Specifically, the process as shown in Equation (13) is applied to the output of S902.
[0118] [Number]
[0119] Here, s input = 2 8 - 1, and the parentheses on the right side represent the process of rounding off the decimal part. After setting the scale of the 14-bit real value to the range of [0, 2 8 - 1], an unsigned 8-bit integer value is obtained by rounding off the decimal part.
[0120] As described above, according to Modification 3, the image quantization unit 203 has a configuration in which a non-linear process is applied to a 14-bit noisy image to obtain an unsigned 8 bits by a non-uniform quantization method. As a result, the noise component can be expressed accurately, and an improvement in denoising accuracy can be expected.
[0121] The disclosure of this specification includes the following information processing apparatus, learning apparatus, and information processing method. (Item 1) Conversion means for converting an input image with a first bit depth into a low-bit-depth image with a second bit depth lower than the first bit depth; Estimation means for estimating a noise component map in the input image from the low-bit-depth image using a neural network (NN) with a third bit depth lower than the first bit depth and equal to or higher than the second bit depth; Derivation means for deriving a noise-removed image corresponding to the input image based on the input image and the noise component map; An information processing apparatus comprising the above. (Item 2) The estimation means estimates an intermediate noise component map with the third bit depth from the low-bit-depth image using the NN, and performs bit-depth conversion of the intermediate noise component map to the first bit depth, thereby estimating the noise component map. The derivation means derives the noise-removed image by subtracting the noise component map from the input image. The information processing apparatus according to item 1, characterized by the above. (Item 3) The NN includes a conversion layer for the bit-depth conversion. The bit-depth conversion includes a non-linear conversion. The information processing apparatus according to item 2, characterized by the above. (Item 4) The bit-depth conversion of the NN non-linearly converts the intermediate noise component map clipped by a threshold value to the third bit depth. The information processing apparatus according to item 2 or 3, characterized by the above. (Item 5) The bit-depth conversion of the NN performs non-linear conversion on the intermediate noise component map non-linearly converted to the third bit depth using the inverse function of the non-linear conversion, and converts it to an intermediate noise component map with a tone of the third bit depth having the same value range as the threshold value. The information processing apparatus according to item 4, characterized by the above. (Item 6) The non-linear transformation is performed using a look-up table (LUT) or by arithmetic operations, wherein the arithmetic operations include operations using a piecewise linear function The information processing apparatus according to any one of items 3 to 5, characterized in that. (Item 7) The third bit depth is equal to the second bit depth The information processing apparatus according to any one of items 1 to 6, characterized in that. (Item 8) When the conversion means converts the input image with the first bit depth into a low-bit depth image with the second bit depth lower than the first bit depth, the conversion means executes a process including a non-linear transformation The information processing apparatus according to any one of items 1 to 7, characterized in that. (Item 9) In the process including the non-linear transformation, values closer to the black level are converted to finer gradations The information processing apparatus according to item 8, characterized in that. (Item 10) A learning apparatus for learning the NN of the information processing apparatus according to any one of items 1 to 9, A first acquisition means for acquiring a clean image with the first bit depth without noise and a noise component map to be added to the clean image, A second acquisition means for acquiring a noisy image with the first bit depth obtained by adding the noise component map to the clean image, A second conversion means for converting the noisy image into a low-bit depth image with the second bit depth, A second estimation means for estimating an estimation map, which is an estimation result of the noise component map, from the low-bit depth image using the NN, An update means for updating the parameters of the NN based on an error between the estimation map and the noise component map, A learning apparatus, characterized by comprising. (Item 11) The update means updates the parameters of the NN by error backpropagation The learning device according to item 10, characterized in that... (Item 12) An information processing method in an information processing apparatus, comprising: a conversion step of converting an input image with a first bit depth into a low-bit-depth image with a second bit depth lower than the first bit depth; an estimation step of estimating a noise component map in the input image from the low-bit-depth image using a neural network (NN) with a third bit depth lower than the first bit depth and equal to or higher than the second bit depth; a derivation step of deriving a noise-removed image corresponding to the input image based on the input image and the noise component map; and characterized by including the above steps.
[0122] (Other embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or apparatus via a network or a storage medium, and having one or more processors in a computer of the system or apparatus read and execute the program. Further, it can also be realized by a circuit (for example, ASIC) that realizes one or more functions.
[0123] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Therefore, the claims are attached to disclose the scope of the invention.
Explanation of reference numerals
[0124] 202 Image acquisition unit; 203 Image quantization unit; 204 Difference estimation unit; 205 High-quality image estimation unit; 206 Learning data acquisition unit; 207 Error calculation unit; 208 Parameter update unit; 209 Parameter integerization unit
Claims
1. a conversion means for converting an input image having a first bit depth into a low bit depth image having a second bit depth lower than the first bit depth; an estimation means for estimating a noise component map in the input image from the low bit-depth image using a neural network (NN) having a third bit-depth lower than the first bit-depth and equal to or greater than the second bit-depth; a derivation means for deriving a noise-removed image corresponding to the input image based on the input image and the noise component map; An information processing device comprising:
2. the estimation means estimates an intermediate noise component map of the third bit depth from the low bit depth image using the neural network and performs bit depth conversion of the intermediate noise component map to the first bit depth, thereby estimating the noise component map; The derivation means derives the noise-removed image by subtracting the noise component map from the input image.
2. The information processing apparatus according to claim 1,
3. The neural network includes a conversion layer for the bit depth conversion; The bit depth conversion includes a non-linear conversion.
3. The information processing apparatus according to claim 2.
4. The bit depth conversion of the neural network nonlinearly converts the thresholded intermediate noise component map to the third bit depth.
3. The information processing apparatus according to claim 2.
5. The bit depth conversion of the NN performs a nonlinear conversion on the intermediate noise component map that has been nonlinearly converted to the third bit depth using an inverse function of the nonlinear conversion, and converts the intermediate noise component map into an intermediate noise component map having gradations of the third bit depth having the same value range as the threshold value.
5. The information processing apparatus according to claim 4.
6. The nonlinear transformation is performed using a look-up table (LUT) or by arithmetic operations; The arithmetic operation includes an operation using a piecewise linear function.
4. The information processing apparatus according to claim 3.
7. The third bit depth is equal to the second bit depth.
2. The information processing apparatus according to claim 1,
8. The conversion means performs processing including a nonlinear conversion when converting an input image having the first bit depth into a low bit depth image having the second bit depth lower than the first bit depth.
2. The information processing apparatus according to claim 1,
9. In the process including the nonlinear conversion, the closer the value is to the black level, the finer the gradation is.
9. The information processing apparatus according to claim 8,
10. A learning device for learning the neural network of the information processing device according to claim 1, a first obtaining means for obtaining a clean image of the first bit depth free of noise and a noise component map to be added to the clean image; a second acquisition means for acquiring a noisy image of the first bit depth obtained by adding the noise component map to the clean image; second conversion means for converting the noisy image into a low bit depth image having the second bit depth; a second estimation means for estimating an estimated map, which is an estimation result of the noise component map, from the low bit depth image using the neural network; an update means for updating parameters of the neural network based on an error between the estimation map and the noise component map; A learning device comprising:
11. The updating means updates the parameters of the neural network by backpropagation of errors. The learning device according to claim 10 .
12. An information processing method in an information processing device, a conversion step of converting an input image having a first bit depth into a low bit depth image having a second bit depth lower than the first bit depth; an estimation step of estimating a noise component map in the input image from the low bit-depth image using a neural network (NN) having a third bit-depth lower than the first bit-depth and equal to or greater than the second bit-depth; deriving a denoised image corresponding to the input image based on the input image and the noise component map; 13. An information processing method comprising: