Image processing apparatus, learning method, and inference method

By using an image processing device in an edge device, the device includes a preprocessing unit and a network unit, and adopts nonlinear transformation and convolutional computing technology, the problem of reducing accuracy in low-quality image processing is solved, and efficient image quality improvement is achieved.

CN120129922APending Publication Date: 2025-06-10MAXELL LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202380075289.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-31
Filing Date
2023-09-19
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

When image processing of low-quality images in edge devices to make them high-quality images, it is difficult to achieve a balance between model size and accuracy, resulting in a decrease in accuracy.

Method used

An image processing device is adopted, which includes a preprocessing unit and a network unit. The preprocessing unit converts the pixel value of the input image into a bit with a lower number of bits using a predetermined nonlinear function, and the network unit performs convolutional operations to improve image quality. The device also includes a pooling layer, an upsampling layer and a jump connection to form a U-Net structure.

Benefits of technology

By improving the processing accuracy and efficiency of low-quality images, it is possible to efficiently improve image quality in edge devices while maintaining miniaturization of model size.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120129922A_ABST
    Figure CN120129922A_ABST
Patent Text Reader

Abstract

An image processing device includes: a preprocessing unit that converts a pixel value of an input image to a number of bits lower than the number of bits of the pixel value by using a predetermined function having nonlinearity; and a network unit that performs a convolution operation using the data converted by the preprocessing unit as an input.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing apparatus, a learning method, and an inference method.

[0002] This application claims the priority of Japanese Patent Application No. 2022-174815 filed on October 31, 2022, and incorporates its content herein. Background Art

[0003] When an image is captured by an imaging device, sometimes the amount of surrounding light is insufficient, or a low-quality image is obtained due to settings of the imaging device such as shutter speed, aperture, or ISO sensitivity. There is a technology for transforming a captured low-quality image into a high-quality image through image processing. For example, a technology for performing image processing on a low-quality image using machine learning to make it a high-quality image is known (for example, refer to Patent Document 1).

[0004] Prior Art Documents

[0005] Patent Documents

[0006] Patent Document 1: U.S. Patent No. 10623756 Specification Summary of the Invention

[0007] When applying the prior art as described above to an edge device, miniaturization of the model size is required. However, if the model size is made too small, sometimes a sufficient high-quality image cannot be obtained. That is, one of the problems in the case of miniaturizing the model size is a decrease in accuracy. Therefore, when performing image processing on a low-quality image to make it a high-quality image in an edge device, it is important to achieve a balance between the model size and accuracy.

[0008] Therefore, an object of the present invention is to provide a technology capable of improving the accuracy and efficiency when performing image processing on a low-quality image using machine learning to make it a high-quality image.

[0009] [1] To solve the above problems, one aspect of the present invention is an image processing apparatus including: a preprocessing unit that transforms pixel values of an input input image into a smaller number of bits than the number of bits of the pixel values using a predetermined function having non-linearity; and a network unit that performs a convolution operation with the data transformed by the preprocessing unit as an input.

[0010] [2] Further, one aspect of the present invention is the image processing apparatus according to [1] above, wherein the network unit includes: a pooling layer that performs pooling processing on the result of convolution operation; and an upsampling layer having a structure symmetric to the pooling layer that performs upsampling on the result of convolution operation, and the network unit has a U-Net structure connected by a skip connection.

[0011] [3] Further, one aspect of the present invention is the image processing apparatus according to [1] or [2] above, wherein a post-processing unit is further provided, and the post-processing unit generates an image with higher image quality than the image input to the pre-processing unit based on the result of convolution operation performed by the network unit and the image input to the pre-processing unit.

[0012] [4] Further, one aspect of the present invention is the image processing apparatus according to any one of [1] to [3] above, wherein it is configured to approximate with a plurality of linear functions instead of a predetermined non-linear function used in the transformation of the pre-processing unit.

[0013] [5] Further, one aspect of the present invention is the image processing apparatus according to any one of [1] to [4] above, wherein the predetermined function used by the pre-processing unit in the transformation of the number of bits is determined according to the gamma function used in the gamma processing of the input image.

[0014] [6] Further, one aspect of the present invention is the image processing apparatus according to any one of [1] to [5] above, wherein the network unit performs convolution operation after performing batch normalization processing, performing an operation of an activation function, and performing scaling processing by multiplying a predetermined function, and the batch normalization processing normalizes the data distribution.

[0015] [7] Further, one aspect of the present invention is the image processing apparatus according to any one of [1] to [6] above, wherein the network unit transforms the result of convolution operation into data of 16 bits or more, and quantizes the data of 16 bits or more obtained as the result of convolution operation into 8 bits or less.

[0016] [8] Further, one aspect of the present invention is the image processing apparatus according to [7] above, wherein the network unit quantizes the data of 16 bits or more obtained as the result of convolution operation into 8 bits or less by any one of comparison with a plurality of thresholds or transformation using a predetermined function.

[0017] [9] Further, one aspect of the present invention is the image processing apparatus according to any one of [1] to [8] above, wherein the preprocessing unit transforms the pixel values into 8-bit data, and the network unit performs a convolution operation with the 8-bit data transformed by the preprocessing unit as input.

[0018]

[10] Further, one aspect of the present invention is a learning method, comprising: a preprocessing step of using a predetermined function having non-linearity to transform the pixel values of a pair of a high-quality image and a low-quality image included in the supervised data into a number of bits lower than the number of bits of the pixel values; and a learning step of performing learning related to the extraction of the noise component superimposed on the low-quality image with the data transformed by the preprocessing step as input.

[0019]

[11] Further, one aspect of the present invention is an inference method, comprising: a preprocessing step of using a predetermined function having non-linearity to transform the pixel values of an input image into a number of bits lower than the number of bits of the pixel values; an inference step of performing inference related to the extraction of the noise component with the data transformed by the preprocessing step as input; and a post-processing step of performing processing to eliminate non-linearity using the inverse function of the predetermined function having non-linearity with respect to the inferred noise component, and subtracting the noise component from which non-linearity has been eliminated from the input image to generate an output image with higher image quality than the input image.

[0020]

[12] Further, one aspect of the present invention is a learning method, comprising: a preprocessing step of using a predetermined function having non-linearity to transform the pixel values of a pair of a high-quality image and a low-quality image included in the supervised data into a number of bits lower than the number of bits of the pixel values; and a learning step of performing learning related to the extraction of the noise component superimposed on the low-quality image and the transformation using the inverse function of the predetermined function having non-linearity with the data transformed by the preprocessing step as input.

[0021]

[13] Further, one aspect of the present invention is an inference method, comprising: a preprocessing step of using a predetermined function having non-linearity to transform the pixel values of an input image into a number of bits lower than the number of bits of the pixel values; an inference step of performing inference related to the extraction of the noise component and the transformation using the inverse function of the predetermined function having non-linearity with the data transformed by the preprocessing step as input; and a post-processing step of subtracting the inferred noise component from the input image to generate an output image with higher image quality than the input image.

[0022] According to the present invention, it is possible to improve the accuracy and efficiency when performing image processing on a low-quality image using machine learning to make it a high-quality image. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is a block diagram showing an example of the functional structure of the image processing system according to the embodiment.

[0024] Figure 2 It is a diagram for explaining the functional blocks of the processing unit according to the embodiment.

[0025] Figure 3 It is a diagram showing an example of the data input to the preprocessing unit and the data output from the preprocessing unit according to the embodiment.

[0026] Figure 4 It is a diagram showing a first example of the function used by the preprocessing unit in the transformation according to the embodiment.

[0027] Figure 5 It is a diagram showing a second example of the function used by the preprocessing unit in the transformation according to the embodiment.

[0028] Figure 6 It is a diagram for explaining the skip connection included in the network unit according to the embodiment.

[0029] Figure 7 It is a block diagram showing an example of the functional structure of the operation block included in the network unit according to the embodiment.

[0030] Figure 8 It is a diagram showing an example of the activation function according to the embodiment.

[0031] Figure 9 It is a flowchart showing a first example of the processing in the learning stage according to the embodiment.

[0032] Figure 10 It is a flowchart showing a first example of the processing in the inference stage according to the embodiment.

[0033] Figure 11 It is a flowchart showing a second example of the processing in the learning stage according to the embodiment.

[0034] Figure 12 It is a flowchart showing a second example of the processing in the inference stage according to the embodiment.

[0035] Figure 13 It is a block diagram showing an example of the internal structure of the image processing apparatus, learning apparatus, and inference apparatus according to the embodiment.

[0036] (REFERENCE SIGNS)

[0037] 1: Image processing system; 10: Image sensor; 20: Processing unit; 21: Preprocessing unit; 22: Network unit; 220: Arithmetic block; 221: BN layer; 222: PReLU layer; 223: Scale layer; 224: Quantization layer; 225: Convolution layer; 226: Pooling layer / Upsampling layer; 23: Postprocessing unit; 30: ISP; 40: Memory; 51: First image; 52: Second image; 53: Third image; SC: Skip connection; GSC: Global skip connection. Detailed implementation

[0038] [Embodiment]

[0039] Hereinafter, preferred embodiments of the image processing apparatus, learning method, and inference method according to the aspect of the present invention will be given and described in detail with reference to the accompanying drawings. In addition, the aspect of the present invention is not limited to these embodiments, and also includes aspects with various modifications or improvements. That is, the elements that can be easily conceived by those skilled in the art and the substantially identical elements are included in the components described below, and the components described below can be appropriately combined. In addition, various omissions, substitutions, or changes of the components can be made without departing from the gist of the present invention. In addition, in the following drawings, in order to make each component easy to understand, sometimes the scale and quantity in each structure are different from the scale and quantity in the actual structure.

[0040] First, matters that are the premise of this embodiment will be described. The image processing apparatus, learning method, and inference method of this embodiment are used in embedded devices such as IoT (Internet of Things) devices. As an example of an IoT device, a camera that is an edge device for capturing still images or moving images can be exemplified. Since the image processing apparatus, learning method, and inference method of this embodiment are applied to edge devices, lightweight and high-speed processing are required. Edge devices such as cameras may have functions such as image recognition and object detection. In addition, this embodiment is not limited to this example, and can also be implemented by a plurality of devices connected via a network.

[0041] The images with improved quality processed using the image processing apparatus, learning method, and inference method of this embodiment can be used for appreciation. In addition, object detection can be performed based on the images with improved quality processed using the image processing apparatus, learning method, and inference method of this embodiment. At this time, compared with the case of performing object detection based on low-quality images, object detection can be performed with higher accuracy.

[0042] Figure 1It is a block diagram showing an example of the functional structure of an image processing system according to an embodiment. Referring to this figure, an example of the functional structure of the image processing system 1 will be described. The image processing system 1 includes an image sensor 10, a processing unit 20, an ISP 30, and a memory 40.

[0043] The image sensor 10 outputs an electrical signal corresponding to the intensity of the incident light in units of pixels. That is, the image sensor 10 performs photoelectric conversion on the image of the subject imaged by the optical system. Specifically, the image sensor 10 is configured to include a CCD image sensor, a CMOS image sensor, etc. The image sensor 10 outputs a first image 51 representing the image of the captured subject. Specifically, the first image 51 may be a digital image signal in RAW format (hereinafter, referred to as RAW image data). For example, the RAW image data output by the image sensor 10 may be data representing the pixel values of each pixel in 12 [bit (bits)] or 14 [bit]. In the present embodiment, the case of expressing the pixel value in bits may include the case of expressing the effective amount of information contained in the data by bit values. That is, even if the data originally represented in 12 [bit] or 14 [bit] is represented in 16 [bit (bits)] by performing bit shift or other processing in part of the operations, in the present embodiment, it may be represented in 12 [bit] or 14 [bit].

[0044] The processing unit 20 acquires the first image 51 output from the image sensor 10. The processing unit 20 performs a predetermined process on the first image 51. Specifically, the process performed by the processing unit 20 may be a process of transforming a low-quality image into a high-quality image (noise reduction process). The processing unit 20 outputs a second image 52 obtained as a result of the process. The second image 52 is a high-quality image obtained by removing noise from the image captured by the image sensor 10.

[0045] The ISP (Image Signal Processor) 30 acquires the second image 52 output from the processing unit 20. The ISP 30 performs a predetermined process on the second image 52. The predetermined process performed by the ISP 30 may be, for example, black level adjustment, HDR (High Dynamic Range) synthesis, exposure adjustment, pixel defect correction, black point correction, demosaicing, white balance adjustment, color correction, gamma correction, etc. The ISP 30 outputs a third image 53 obtained as a result of the process. The third image 53 is a high-quality image obtained by further enhancing the quality of the second image 52.

[0046] The memory 40 includes storage devices such as a non-volatile ROM (Read Only Memory) or a volatile RAM (Random Access Memory). The memory 40 acquires the third image 53 output from the ISP 30. The memory 40 stores the acquired third image 53. A predetermined process is performed on the third image 53 stored in the memory 40 by a CPU (Central Processing Unit) or the like (not shown). The predetermined process may be display on the display unit, output to an external device, or the like.

[0047] Figure 2 FIG. is a diagram showing functional blocks of a processing unit for explaining an embodiment. Referring to this figure, details of each functional block included in the processing unit 20 will be described. In the following description, a device having the structure included in the processing unit 20 may be referred to as an image processing device. The processing unit 20 includes a preprocessing unit 21, a network unit 22, and a postprocessing unit 23. The preprocessing unit 21, the network unit 22, and the postprocessing unit 23 are connected in series. The first image 51 output from the image sensor 10 is input to the preprocessing unit 21. In addition, the first image 51 input to the preprocessing unit 21 is also input to the postprocessing unit 23. A path in which the first image 51 output from the image sensor 10 is input to the postprocessing unit 23 by skipping the preprocessing unit 21 and the network unit 22 is shown as a global skip connection GSC.

[0048] Figure 3 FIG. is a diagram showing an example of data input to the preprocessing unit and data output from the preprocessing unit according to the embodiment. Referring to this figure, the input / output data of the preprocessing unit 21 will be described. The first image 51 output from the image sensor 10 is input to the preprocessing unit 21. As shown in the figure, the first image 51 is data representing respective pixel values in 12 [bit] or 14 [bit]. The preprocessing unit 21 performs a process of converting each pixel value to 8 [bit]. As shown in the figure, the preprocessing unit 21 outputs the 8 [bit] data as a result of the conversion to the subsequent stage. When converting to 8 [bit] data, the preprocessing unit 21 preferably uses a predetermined function for the conversion. In addition, in the present embodiment, the case of 8 [bit] is exemplified as the conversion of the pixel value in the preprocessing unit 21, but it is not limited thereto. For example, it may be converted to a smaller bit value such as 4 [bit] or 2 [bit].

[0049] Figure 4This is a diagram showing a first example of the function used by the preprocessing unit in the transformation according to an embodiment. Referring to this diagram, a first example of the function used by the preprocessing unit 21 in the transformation will be described. The horizontal axis of this diagram represents the pixel value before transformation (14 [bit]), and the vertical axis represents the pixel value after transformation (8 [bit]). The preprocessing unit 21 applies the illustrated function to each pixel value to perform the transformation. Specifically, when the pixel value before transformation is x1, the preprocessing unit 21 transforms it into y1; when the pixel value before transformation is x2, the preprocessing unit 21 transforms it into y2; when the pixel value before transformation is x3, the preprocessing unit 21 transforms it into y3.

[0050] When the horizontal axis (pixel value before transformation) is set as x and the initial value of the vertical axis (pixel value after transformation) is y0, the illustrated function is specifically represented by y = x^γ - y0 (γ < 1). As shown in the figure, the function used by the preprocessing unit 21 in the transformation preferably has nonlinearity. That is, the preprocessing unit 21 can also use a predetermined function with nonlinearity to transform the pixel value of the input input image (the first image 51) into a number of bits lower than the number of bits of the pixel value of the input image. As shown in the figure, according to the function used by the preprocessing unit 21 in the transformation, in the region where the input signal value is low (that is, the region where the image is dark), more bit values are allocated after the transformation. This function corresponds to the nonlinear processing used in the gamma processing performed by the ISP 30.

[0051] In one example shown in the figure, the range of the vertical axis represents -128 to +127. However, the function of this embodiment is not limited to this example, and the range of the vertical axis can be arbitrarily changed. In addition, in one example shown in the figure, one pixel value is transformed into one pixel value based on the input one pixel value and a predetermined function, but it can also be transformed into multiple pixel values based on multiple functions. These multiple pixel values are represented in the form of a vector. That is, the preprocessing unit 21 can generate a vectorized output value based on the input image and multiple functions.

[0052] In addition, the predetermined function used by the preprocessing unit 21 in the transformation of the number of bits can be determined in advance, or it can be configured to be switched by selecting from among multiple function candidates. For example, the function can be switched at the timing of switching the gamma function (gamma curve) used by the ISP 30 in the gamma processing. That is, the predetermined function used by the preprocessing unit 21 in the transformation of the number of bits can be determined according to the gamma function used in the gamma processing of the input image performed by the ISP 30.

[0053] Figure 5This is a diagram showing a second example of the function used by the preprocessing unit in the transformation in the embodiment. Referring to this diagram, a second example of the function used by the preprocessing unit 21 in the transformation will be described. The horizontal axis of this diagram represents the pixel value before transformation (14 [bit]), and the vertical axis represents the pixel value after transformation (8 [bit]). The function in the second example is obtained by approximating the function in the first example with a plurality of linear functions (in one example shown in the figure, the straight lines L1, L2, and L3). That is, it can be considered that the function used by the preprocessing unit 21 in the transformation is a piecewise linear function composed of a plurality of linear functions. In other words, it can be considered that it is configured to approximate with a plurality of linear functions to replace the predetermined non-linear function used by the preprocessing unit 21 in the transformation.

[0054] Regarding the function in the second example, similar to the function in the first example, 14 [bit] data is transformed into 8 [bit]. In addition, regarding the function in the second example, similar to the function in the first example, in the region where the input signal value is low (that is, the region where the image is dark), more bit values are assigned after transformation. Furthermore, in one example shown in the figure, an example of the case where the function in the second example is a piecewise linear function composed of three linear functions is described, but this function may also be composed of three or more functions, or may be a function obtained by combining non-linear functions.

[0055] Return Figure 2 , the details of the network unit 22 will be described. The network unit 22 takes the 8 [bit] data transformed by the preprocessing unit 21 as input and performs a convolution operation. The network unit 22 is a neural network (CNN: Convolutional Neural Network) having a plurality of operation blocks 220. In one example shown in the figure, the network unit 22 has operation blocks 220-1 to 220-7. The operation blocks 220-1 to 220-7 are connected to each other. Each operation block 220 includes an input layer, a convolutional layer, a pooling layer, a sampling layer, an output layer, and the like. Each operation block 220 includes at least a convolutional layer. In each operation block 220, after performing a convolution operation (or a deconvolution operation), the data of the operation result is quantized as 16 [bit] data, and thus the 16 [bit] data is transformed into 8 [bit] data.

[0056] Specifically, the network unit 22 has a U-Net structure. According to U-Net, as shown in the figure, it has a left-right symmetric encoder-decoder structure. The multiple operation blocks 220 from the left side to the lower center in the figure are an encoder that at least includes a pooling layer for performing pooling processing on the result of a convolution operation, and performs downsampling. The multiple operation blocks 220 from the lower center to the right side in the figure are a decoder that at least includes an upsampling layer for performing upsampling on the result of a convolution operation, and performs upsampling. It can be considered that the encoder and the decoder have a symmetric structure, or it can be considered that the pooling layer and the upsampling layer have a symmetric structure. According to U-Net, the feature map generated by the encoder is connected or added to the feature map of the decoder, etc. Specifically, the feature map generated by the encoder is copied, cropped, and concatenated with the feature map of the decoder. The connection with the feature map of the decoder can be a simple addition. The path of connecting the feature map generated by the encoder with the feature map of the decoder is illustrated as a skip connection SC. In other words, the operation block 220 constituting the encoder and the operation block 220 constituting the decoder are connected by the skip connection SC. In addition, the network unit 22 may also have a structure other than the U-NET structure. As another example, it may also have a Visual Transformer structure.

[0057] Figure 6 It is a diagram for explaining the skip connection of the network unit in the embodiment. Referring to this diagram, a generalized skip connection is explained. As shown in the figure, the input (x) skips the operation until the output and is added to the operation result of each layer (in the illustrated example, it is F(x)). By adding such a skip connection SC between each layer, the feature of anti-gradient disappearance can be obtained.

[0058] Figure 7 It is a block diagram showing an example of the functional structure of the operation block of the network unit in the embodiment. Referring to this diagram, an example of the functional structure of the operation block 220 of the network unit 22 is explained. In addition, the functional structure shown in this diagram is an example, and it may also be different for each of the multiple operation blocks 220 of the network unit 22. The operation block 220 includes a BN layer 221, a PReLU layer 222, a Scale layer 223, a quantization layer 224, a convolution layer 225, and a pooling / upsampling layer 226. The output data of the previous-stage operation block 220 is input to the BN layer 221, and the data output from the pooling / upsampling layer 226 is input to the subsequent stage. In addition, the input from the preprocessing unit 21 is input to the convolution layer 225.

[0059] 16-bit data is input to the BN (Batch Normalization) layer 221. The BN layer 221 normalizes the distribution of the input data. A predetermined mathematical formula can be used in the normalization process. For example, the BN layer 221 adds (add) a constant and multiplies (multiply) a constant to each element so that the average of the values of each element in the batch is 0 and the variance of the values of each element is 1. In one example shown in the figure, multiplication is performed after addition, but the order of addition and multiplication can also be reversed (that is, addition can also be performed after multiplication). The constant used in addition and the constant used in multiplication can be 32-bit or 16-bit values of floating-point type respectively. The BN layer 221 outputs 32-bit or 16-bit data of floating-point type to the subsequent stage.

[0060] 32-bit or 16-bit data of floating-point type is input to the PReLU layer 222. The PReLU layer 222 performs an operation of an activation function on the input data.

[0061] Figure 8 It is a diagram showing an example of an activation function of an embodiment. Referring to this diagram, an example of the activation function will be described. The horizontal axis represents the input (x), and the vertical axis represents the output (y). In one example shown in the figure, in the range where x < 0, y = px, and in the range where x > 0, y = px. In addition, the activation function is set as PReLU (Parametric Rectified Linear Unit), but the activation function can also be ReLU (Rectified Linear Unit) or Identity (identity function). In PReLU, when slope (p) is set to 0, it becomes ReLU, and when slope (p) is set to 1, it becomes Identity. The range of slope (p) can be a real value from 0 to 1 (32-bit or 16-bit of floating-point type).

[0062] In addition, when the network unit 22 is installed in hardware such as an FPGA (Field Programmable Gate Array) or an ASI (Application Specific Integrated Circuit), the activation function (that is, the PReLU layer 222) can also include the BN layer 221. In addition, the activation function (that is, the PReLU layer 222) can further include quantization processing.

[0063] Return Figure 7, the Scale layer 223 receives data of 32 [bit] or 16 [bit] in floating-point type. The Scale layer 223 performs a scaling process. The scaling process is a process of restoring the normalized data (the opposite process of Batch Normalization). The Scale layer 223 performs addition (add) of a constant and multiplication (multiply) of a constant in the same way as the BN layer 221. In one example shown in the figure, multiplication is performed after addition of a constant, but the order of addition and multiplication can be reversed (that is, addition can also be performed after multiplication). The constant used in addition and the constant used in multiplication can be values of 32 [bit] or 16 [bit] in floating-point type respectively. The Scale layer 223 outputs data of 32 [bit] or 16 [bit] in floating-point type to the subsequent stage.

[0064] In the arithmetic block 220 according to the present embodiment, the BN layer 221 exists before the PReLU layer 222, and the Scale layer 223 exists after the PReLU layer 222. In other words, normalization processing (encoding) of data distribution is performed before the operation of the activation function, and restoration processing (decoding) is performed using a predetermined function after the operation of the activation function. After performing these processes, the convolution operation described later is performed. That is, the network unit 22 of the present embodiment performs a convolution operation after performing batch normalization processing, performing an operation of an activation function, and performing a scaling process of multiplying by a predetermined function, and the batch normalization processing normalizes the data distribution.

[0065] , the quantization layer 224 receives data of 32 [bit] or 16 [bit] in floating-point type. The quantization layer 224 quantizes the input data of 16 [bit] or more into low bits (for example, 8 [bit] or less). Among them, since the output from the preprocessing unit 21 is input to the convolutional layer 225, it can be considered that the data input to the quantization layer 224 is the result of at least one convolution operation. That is, it can be considered that the quantization layer 224 quantizes the data of 16 [bit] or more obtained as a result of the convolution operation into low bits (for example, 8 [bit] or less). The quantization process performed by the quantization layer 224 can be performed by any one of (1) comparison with multiple thresholds or (2) transformation using a predetermined function. In addition, the quantization process of the present embodiment is not limited to this example, and quantization can also be performed by other quantization methods. As a result of performing the quantization process, the quantization layer 224 outputs data of 8 [bit] in integer type to the subsequent stage.

[0066] The convolutional layer 225 is input with integer 8-bit data. The convolutional layer 225 performs a convolution operation related to the input data. Specifically, the convolutional layer 225 performs a convolution operation using weights on the input data. Specifically, the convolutional layer 225 performs a multiplication and accumulation operation with the input data and weights as inputs. The weights (filters, convolutional kernels) of the convolutional layer 225 can be multi-dimensional data having elements as learnable parameters. The weights of the convolutional layer 225 can be low-bit (for example, they can be 1-bit signed integers (i.e., -1, 1)). As a result of performing the convolution operation, the convolutional layer 225 outputs integer 16-bit data to the subsequent stage.

[0067] The pooling / upsampling layer 226 is input with integer 16-bit data. The pooling / upsampling layer 226 performs pooling (downsampling) or upsampling (transposed convolution or deconvolution). The pooling / upsampling layer 226 is a pooling layer in the encoder and an upsampling layer in the decoder. As a result of performing the pooling process (or upsampling process), the pooling / upsampling layer 226 outputs integer 16-bit data to the subsequent stage. In addition, the operations or their outputs of the convolutional layer 225 and the pooling / upsampling layer 226 may not be integer 16-bit, for example, they can be fixed-point.

[0068] Return Figure 2 , and the details of the post-processing unit 23 will be described. The result of the convolution operation performed by the network unit 22 and the image (first image 51) input to the pre-processing unit are input to the post-processing unit 23. The result of the convolution operation performed by the network unit 22 is information related to the noise components included in the first image 51. In other words, the network unit 22 is pre-learned to extract the noise components included in the first image 51. The post-processing unit 23 subtracts the noise components from the first image 51 to generate a high-quality image. That is, the post-processing unit 23 generates an image with higher quality than the image input to the pre-processing unit 21 based on the result of the convolution operation performed by the network unit 22 and the image input to the pre-processing unit 21.

[0069] Among them, the network unit 22 processes based on the values transformed into low bits by the pre-processing unit 21 using a predetermined function with non-linearity. The post-processing unit 23 may perform a process of deforming the output of the network unit 22 from a non-linear value to a linear value before the process of subtracting the noise components from the first image 51. In this transformation process, the inverse function of the function shown in Figure 4 or Figure 5 can be used.

[0070] In addition, the network unit 22 can perform learning and inference including this transformation process. In this case, the process of transforming from a non-linear value to a linear value performed by the post-processing unit 23 can be omitted.

[0071] Next, with reference to Figures 9 to 12 , an example of a series of operations in the learning phase and the inference phase of the image processing system 1 according to the present embodiment will be described. First, with reference to Figure 9 and Figure 10 , a first example will be described. In the first example, in the post-processing, learning is performed on the premise of transformation processing from a non-linear value to a linear value. Therefore, in the first example, in the post-processing, transformation processing from a non-linear value to a linear value is required.

[0072] Figure 9 is a flowchart showing a first example of the processing in the learning phase of the embodiment. Referring to this figure, a first example of the processing in the learning phase of the image processing system 1 will be described.

[0073] (Step S11) First, the preprocessing unit 21 preprocesses the RAW image output from the image sensor 10, which is the supervised data. The supervised data includes a pair of high-quality images and low-quality images. The pair of high-quality images and low-quality images are images obtained by photographing the same object, and noise is superimposed on the low-quality image. The low-quality image may be an image obtained by photographing the same subject as the high-quality image with different settings, or an image generated by performing image processing on the high-quality image. Both the high-quality image and the low-quality image included in the supervised data are RAW images of 12 [bit] or 14 [bit]. Specifically, the preprocessing unit 21 uses a predetermined function with non-linearity to transform the pixel values of each of the pair of high-quality images and low-quality images included in the supervised data into low-bit data. When the pixel value of the image included in the supervised data is 12 [bit] or 14 [bit], the preprocessing unit 21 transforms it into data of a lower bit number than the bit number of the pixel value of the image included in the supervised data, that is, 8 [bit] data. Sometimes the process performed by the preprocessing unit 21 is referred to as a preprocessing process.

[0074] (Step S13) Then, the data transformed by the preprocessing process is input to the network unit 22. The network unit 22 performs learning based on the data transformed by the preprocessing process. Sometimes, the process of learning by the network unit 22 is recorded as the learning process. In the learning process, the data transformed by the preprocessing process is used as the input, and learning related to the extraction of the noise component superimposed on the low-quality image is performed. Among them, in the learning process of the first example, learning is performed based on the data transformed using a predetermined function with nonlinearity in the preprocessing process. That is, in the inference stage of the first example, after inference by the network unit 22, a transformation for eliminating nonlinearity is required. Specifically, the transformation for eliminating nonlinearity can be a transformation using the inverse function of the predetermined function with nonlinearity used in the preprocessing process. In addition, in the learning process in the first example, the preprocessing process can be included as the object of learning. As an example, parameters such as the coefficients and constants of the predetermined function with nonlinearity in the preprocessing process can be learned.

[0075] Figure 10 It is a flowchart showing the first example of the processing in the inference stage of the embodiment. Referring to this figure, the first example of the processing in the inference stage of the image processing system 1 will be described.

[0076] (Step S21) First, the preprocessing unit 21 performs preprocessing on the RAW image output from the image sensor 10, which is the object of image processing. Preferably, the image that is the object of image processing is a low-quality image superimposed with noise. The image that is the object of image processing is a 12 [bit] or 14 [bit] RAW image. Specifically, the preprocessing unit 21 uses a predetermined function with nonlinearity to transform the pixel values of the image that is the object of image processing into low-bit data. When the pixel value of the image that is the object of image processing is 12 [bit] or 14 [bit], the preprocessing unit 21 transforms it into data with a lower number of bits than the pixel value of the image that is the object of image processing, that is, 8 [bit] data.

[0077] (Step S23) Then, using the learning model generated in Step S13, the data transformed by the preprocessing process is used as the input to perform inference on the noise component. Sometimes, the process of performing inference on the noise component is recorded as the inference process. The data transformed by the preprocessing process is input to the network unit 22, and the network unit 22 outputs the inference result of the noise component to the post-processing unit 23.

[0078] (Step S25) Then, the post-processing unit 23 performs a process of eliminating nonlinearity on the noise component inferred by the inference process using the inverse function of the predetermined function with nonlinearity. The inverse function of the predetermined function with nonlinearity can be the inverse function of the function used in Step S21.

[0079] (Step S27) Then, the input image that is the object of image processing is input to the post-processing unit 23 through the global skip connection GSC. The post-processing unit 23 subtracts the noise component from which non-linearity has been eliminated from the input image that is the object of image processing (i.e., the low-quality image with noise superimposed), removes the noise from the low-quality image, and generates an output image with higher quality (higher image quality) than the input image. In addition, the processes performed in step S25 and step S27 are sometimes recorded as post-processing steps.

[0080] Next, with reference to Figure 11 and Figure 12 , a second example will be described. In the second example, learning including transformation processing from a non-linear value to a linear value is performed. Therefore, in the second example, in post-processing, transformation processing from a non-linear value to a linear value is not required.

[0081] Figure 11 is a flowchart showing a second example of the processing in the learning stage of the embodiment. With reference to this figure, a second example of the processing in the learning stage of the image processing system 1 will be described.

[0082] (Step S31) First, the preprocessing unit 21 preprocesses the RAW image output from the image sensor 10, which is the supervised data. The supervised data includes a pair of high-quality images and low-quality images. The pair of high-quality images and low-quality images are images obtained by photographing the same object, and noise is superimposed on the low-quality image. The low-quality image can be an image obtained by photographing the same subject as the high-quality image with different settings, or an image generated by performing image processing on the high-quality image. Both the high-quality image and the low-quality image included in the supervised data are 12[bit] or 14[bit] RAW images. Specifically, the preprocessing unit 21 uses a predetermined function with non-linearity to transform the pixel values of each of the pair of high-image-quality images and low-image-quality images included in the supervised data into low-bit data. When the pixel value of the image included in the supervised data is 12[bit] or 14[bit], the preprocessing unit 21 transforms it into data with a lower number of bits than the number of bits of the pixel value of the image included in the supervised data, that is, 8[bit] data.

[0083] (Step S33) Then, the data transformed by the preprocessing process is input to the network unit 22 to perform the learning process. In the learning process, the data transformed by the preprocessing process is used as the input to perform learning related to the extraction of the noise component superimposed on the low-quality image. Additionally, in the learning process of the second example, further learning related to the transformation using the inverse function of a predetermined function with nonlinearity is performed. That is, in the inference stage of the second example, learning including the transformation for eliminating nonlinearity is performed. Therefore, the elimination process of nonlinearity in the post-processing process is not required. Furthermore, in the learning process of the second example, the preprocessing process can be included as the object of learning. As an example, parameters such as the coefficients and constants of the predetermined function with nonlinearity in the preprocessing process can be learned.

[0084] Figure 12 FIG. is a flowchart showing a second example of the processing in the inference stage of the embodiment. Referring to this figure, a second example of the processing in the inference stage of the image processing system 1 will be described.

[0085] (Step S41) First, the preprocessing unit 21 preprocesses the RAW image output from the image sensor 10, which is the object of image processing. Preferably, the image that is the object of image processing is a low-quality image with superimposed noise. The image that is the object of image processing is a 12-bit or 14-bit RAW image. Specifically, the preprocessing unit 21 uses a predetermined function with nonlinearity to transform the pixel values of the image that is the object of image processing into low-bit data. When the pixel values of the image that is the object of image processing are 12 bits or 14 bits, the preprocessing unit 21 transforms them into data with a lower number of bits, i.e., 8 bits, than the number of bits of the pixel values of the image that is the object of image processing.

[0086] (Step S43) Then, using the learning model generated in Step S33, the data transformed by the preprocessing process is used as the input to perform inference of the noise component. The learning model generated in Step S33 performs learning including the transformation for eliminating nonlinearity. Therefore, it can be considered that the inference result output in the inference process of the second example is the inference result after the transformation for eliminating nonlinearity has been performed. The data transformed by the preprocessing process is input to the network unit 22, and the network unit 22 outputs the inference result of the noise component to the post-processing unit 23.

[0087] (Step S45) Then, the input image that is the object of image processing is input to the post-processing unit 23 through the global skip connection GSC. The post-processing unit 23 subtracts the inference result of the noise component output from the network unit 22 from the input image that is the object of image processing (i.e., the low-quality image superimposed with noise), removes the noise from the low-quality image, and generates an output image of higher quality (higher picture quality) than the input image. Further, in the second example, Step S45 corresponds to the post-processing step.

[0088] Figure 13 FIG. is a block diagram showing an example of the internal structure of the image processing apparatus, learning apparatus, and inference apparatus of the present embodiment. At least part of the functions of the image processing apparatus, learning apparatus, and inference apparatus can be implemented using a computer. As shown in the figure, the computer is configured to include a central processing unit 901, a RAM 902, an input / output port 903, input / output devices 904, 905, etc., and a bus 906. The computer itself can be implemented using existing technologies. The central processing unit 901 executes commands included in a program read from the RAM 902 or the like. The central processing unit 901 writes data to the RAM 902, reads data from the RAM 902, or performs arithmetic operations and logical operations according to each command. The RAM 902 stores data and programs. Each element included in the RAM 902 has an address, and can be accessed using the address. The input / output port 903 is a port for the central processing unit 901 to exchange data with external input / output devices or the like. The input / output devices 904, 905 are input / output devices. The input / output devices 904, 905 exchange data with the central processing unit 901 via the input / output port 903. The bus 906 is a common communication path used inside the computer. For example, the central processing unit 901 reads or writes data to / from the RAM 902 via the bus 906. Additionally, for example, the central processing unit 901 accesses the input / output port via the bus 906. All or part of the functional units included in the image processing apparatus, learning apparatus, and inference apparatus can be implemented using hardware such as an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field-Programmable Gate Array).

[0089] [Summary of the Present Embodiment]

[0090] According to the embodiment described above, the image processing apparatus includes a preprocessing unit 21 that uses a predetermined function having nonlinearity to transform the pixel values of an input input image into a number of bits lower than the number of bits of the pixel values of the input image. In addition, the image processing apparatus includes a network unit 22 that takes the data transformed by the preprocessing unit 21 as an input and performs a convolution operation. That is, according to the image processing apparatus of the present embodiment, the input image is transformed into a nonlinear form and input to the network. Among them, the image data acquired by the image sensor 10 such as a CMOS sensor has a linear characteristic with respect to the input (light amount). The image processing apparatus performs a transformation using a predetermined function having nonlinearity, so that in a region where the input signal value is low (i.e., a region where the image is dark), more bit values can be allocated. In a region where the image is dark, noise is likely to occur, and higher-precision processing is required. According to the image processing apparatus of the present embodiment, a transformation using a predetermined function having nonlinearity is performed, so that more bit values are allocated in a region where the image is dark. Therefore, noise components can be extracted with high precision. In addition, according to the image processing apparatus of the present embodiment, it is transformed into a low bit in the preprocessing before the network, so that processing can be performed efficiently. Therefore, even if the image processing apparatus of the present embodiment is assembled in an edge device, it can operate efficiently. Thus, according to the image processing apparatus of the present embodiment, the accuracy and efficiency of image processing of a low-quality image using machine learning to make it a high-quality image can be improved.

[0091] In addition, according to the embodiment described above, the network unit 22 includes: a pooling layer that performs pooling processing on the result of the convolution operation; and an upsampling layer that has a structure symmetric to the pooling layer and performs upsampling on the result of the convolution operation. The network unit 22 has a U-Net structure connected by skip connections. According to the image processing apparatus of the present embodiment, the U-Net structure is adopted, so that the vanishing gradient is resisted, and learning and inference can be performed efficiently.

[0092] In addition, according to the embodiment described above, the image processing apparatus further includes a postprocessing unit 23 connected to the preprocessing unit 21 by a global skip connection GSC. The image processing apparatus further includes the postprocessing unit 23 to generate an image with higher image quality than the image input to the preprocessing unit 21 based on the result of the convolution operation performed by the network unit 22 and the image input to the preprocessing unit 21. Therefore, according to the image processing apparatus of the present embodiment, the extracted noise components are subtracted from the original input image, so that an image with high image quality can be easily generated.

[0093] In addition, according to the embodiments described above, the predetermined function with non-linearity used by the preprocessing unit 21 in the transformation is composed of a plurality of functions with linearity. That is, it can be considered that the function used in the transformation is a combination of multiple straight lines. Therefore, according to the image processing apparatus of the present embodiment, the arithmetic processing can be made lightweight. Thus, according to the image processing apparatus of the present embodiment, the efficiency of performing image processing on a low-quality image using machine learning to make it a high-quality image can be improved.

[0094] In addition, according to the embodiments described above, the predetermined function used by the preprocessing unit 21 in the transformation of the number of bits is determined (switched) according to the gamma function used in the gamma processing of the input image in the ISP 30. That is, according to the image processing apparatus of the present embodiment, preprocessing is performed using a function corresponding to the gamma function used in the gamma processing of the input image in the ISP 30, so as to extract noise components in consideration of the gamma processing. Therefore, according to the image processing apparatus of the present embodiment, noise components can be extracted with high accuracy. Thus, according to the image processing apparatus of the present embodiment, the accuracy of performing image processing on a low-quality image using machine learning to make it a high-quality image can be improved.

[0095] In addition, according to the embodiments described above, after the network unit 22 performs batch normalization processing, arithmetic operation of an activation function, and scaling processing of multiplying by a predetermined function, a convolution operation is performed, and the batch normalization processing normalizes the data distribution. In other words, batch normalization processing and scaling processing are performed before and after the arithmetic operation of the activation function performed by the network unit 22. According to the image processing apparatus of the present embodiment, the arithmetic operation of the activation function is performed based on the normalized data, so that the accuracy related to the extraction of noise components can be improved. Thus, according to the image processing apparatus of the present embodiment, the accuracy of performing image processing on a low-quality image using machine learning to make it a high-quality image can be improved.

[0096] In addition, according to the embodiments described above, as a result of performing the convolution operation, the network unit 22 transforms into data of 16 bits or more, and quantizes the data of 16 bits or more obtained as a result of the convolution operation to 8 bits or less. That is, the network unit 22 repeatedly performs convolution operation and quantization to extract noise components. Thus, according to the image processing apparatus of the present embodiment, the accuracy and efficiency of performing image processing on a low-quality image using machine learning to make it a high-quality image can be improved.

[0097] In addition, according to the embodiments described above, the network unit 22 quantizes data of 16 bits or more obtained as a result of performing a convolution operation into 8 bits or less by any one of (1) comparison with a plurality of thresholds or (2) transformation using a predetermined function. Therefore, according to the image processing apparatus of the present embodiment, quantization can be easily performed. Thus, according to the image processing apparatus of the present embodiment, the efficiency of performing image processing on a low-quality image using machine learning to make it a high-quality image can be improved.

[0098] In addition, according to the embodiments described above, the preprocessing unit 21 transforms the pixel values into 8-bit data, and the network unit 22 performs a convolution operation with the 8-bit data transformed by the preprocessing unit 21 as input. That is, according to the image processing apparatus of the present embodiment, data of a lower bit than the input image is input to the network unit 22. Therefore, according to the image processing apparatus of the present embodiment, the network unit 22 can be made lightweight. Thus, according to the image processing apparatus of the present embodiment, the efficiency of performing image processing on a low-quality image using machine learning to make it a high-quality image can be improved.

[0099] In addition, according to the embodiments described above, the learning method of the present embodiment has a preprocessing step of transforming the pixel values of each of a pair of high-quality image and low-quality image included in the supervised data into a number of bits lower than the number of bits of the pixel values of the images included in the supervised data using a predetermined function having non-linearity. In addition, the learning method of the present embodiment has a learning step of performing learning related to extraction of a noise component superimposed on the low-quality image with the data transformed by the preprocessing step as input. That is, according to the learning method of the present embodiment, learning is performed on the premise of non-linearity elimination processing in a post-processing step. Therefore, according to the learning method of the present embodiment, the processing of the network unit 22 can be made lightweight. Thus, according to the learning method of the present embodiment, the efficiency of performing image processing on a low-quality image using machine learning to make it a high-quality image can be improved.

[0100] In addition, according to the embodiments described above, the inference method of the present embodiment has a preprocessing step of using a predetermined function with nonlinearity to transform the pixel values of the input input image into a number of bits lower than the number of bits of the pixel values of the input image. Further, the inference method of the present embodiment has an inference step of using, as input, the data transformed by the preprocessing step to perform inference related to the extraction of noise components. Further, the inference method of the present embodiment has a postprocessing step of using the inverse function of a predetermined function with nonlinearity for the inferred noise components to perform a nonlinearity elimination process, subtracting the noise components with eliminated nonlinearity from the input image, thereby generating an output image with higher image quality than the input image. That is, according to the learning method of the present embodiment, inference is performed on the premise of the nonlinearity elimination process in the postprocessing step. Therefore, according to the inference method of the present embodiment, the processing of the network unit 22 can be lightened. Thereby, according to the inference method of the present embodiment, the efficiency of performing image processing on a low-quality image using machine learning to make it a high-quality image can be improved.

[0101] In addition, according to the embodiments described above, the learning method of the present embodiment has a preprocessing step of using a predetermined function with nonlinearity to transform the pixel values of each of a pair of high-quality image and low-quality image included in the supervised data into a number of bits lower than the number of bits of the pixel values of the images included in the supervised data. Further, the learning method of the present embodiment has a learning step of using, as input, the data transformed by the preprocessing step to perform learning related to the extraction of noise components superimposed on the low-quality image and the transformation using the inverse function of a predetermined function with nonlinearity. That is, according to the learning method of the present embodiment, learning including a transformation process using the inverse function of a predetermined function with nonlinearity is performed. Therefore, according to the learning method of the present embodiment, the processing of the postprocessing unit 23 can be lightened. Thereby, according to the learning method of the present embodiment, the efficiency of performing image processing on a low-quality image using machine learning to make it a high-quality image can be improved.

[0102] In addition, according to the embodiment described above, the inference method of this embodiment has a preprocessing step, thereby using a predetermined function with nonlinearity to transform the pixel values of the input image into a lower number of bits than the number of bits of the pixel values of the input image. In addition, the inference method of this embodiment has an inference step, thereby taking the data transformed by the preprocessing step as input and performing inferences related to the extraction of noise components and the transformation using the inverse function of a predetermined function with nonlinearity. In addition, the inference method of this embodiment has a postprocessing step, thereby subtracting the inferred noise components from the input image to generate an output image with higher image quality than the input image. That is, according to the inference method of this embodiment, inferences are performed using a learning model that has learned including transformation processing, and this transformation processing uses the inverse function of a predetermined function with nonlinearity. Therefore, according to the inference method of this embodiment, the processing of the post-processing unit 23 can be lightened. Thereby, according to the inference method of this embodiment, the efficiency of image processing of low-quality images into high-quality images using machine learning can be improved.

[0103] In addition, the learning objects of the image processing device, learning device, and inference device of this embodiment can be weights, quantization parameters, batch normalization processing, scale processing, etc.

[0104] In addition, the functions of each part of the image processing device, learning device, and inference device of the above embodiment as a whole or a part thereof can also be realized by recording a program for realizing these functions in a computer-readable recording medium, and causing a computer system to read and execute the program recorded in the recording medium. In addition, the "computer system" mentioned here includes hardware such as an OS and peripheral devices.

[0105] In addition, the "computer-readable recording medium" refers to removable media such as floppy disks, optical disks, ROMs, CD-ROMs, and storage parts such as hard disks built into a computer system. Furthermore, the "computer-readable recording medium" can include a medium that dynamically holds a program for a short time, such as a communication line when sending a program via a network such as the Internet or a communication line such as a telephone line, and a medium that holds a program for a certain time, such as a volatile memory inside a computer system that becomes a server or a client at this time. In addition, the above program can be a program for realizing a part of the above functions, or a program for realizing the above functions by combining with a program already recorded in a computer system.

[0106] As described above, the embodiments have been used to illustrate the ways for implementing the present invention, but the present invention is not limited to such embodiments, and various modifications and substitutions can be made without departing from the gist of the present invention.

[0107] Industrial Applicability

[0108] According to the present invention, it is possible to improve the accuracy and efficiency when performing image processing on a low-quality image using machine learning to make it a high-quality image.

Claims

1. An image processing apparatus, comprising: a preprocessing unit that transforms pixel values of an input input image into a number of bits lower than the number of bits of the pixel values using a predetermined function having non-linearity; and a network unit that performs a convolution operation with the data transformed by the preprocessing unit as an input.

2. The image processing apparatus according to claim 1, wherein, the network unit includes: a pooling layer that performs a pooling process on a result of the convolution operation; and an upsampling layer that has a structure symmetric to the pooling layer and performs upsampling on the result of the convolution operation, and the network unit has a U-Net structure connected by skip connections.

3. The image processing apparatus according to claim 1 or 2, wherein, it further includes a postprocessing unit that generates an image with higher image quality than the image input to the preprocessing unit based on the result of the convolution operation performed by the network unit and the image input to the preprocessing unit.

4. The image processing apparatus according to claim 1 or 2, wherein, it is configured to approximate with a plurality of linear functions instead of the predetermined function having non-linearity used in the transformation by the preprocessing unit.

5. The image processing apparatus according to claim 1 or 2, wherein, the predetermined function used by the preprocessing unit in the transformation of the number of bits is determined based on a gamma function used in gamma processing of the input image.

6. The image processing apparatus according to claim 1 or 2, wherein, the network unit performs a convolution operation after performing batch normalization processing, performing an operation of an activation function, and performing scaling processing by multiplying a predetermined function, and the batch normalization processing normalizes a data distribution.

7. The image processing apparatus according to claim 1 or 2, wherein, the network unit transforms data of 16 bits or more as a result of the convolution operation, and quantizes the data of 16 bits or more obtained as a result of the convolution operation to 8 bits or less.

8. The image processing apparatus according to claim 7, wherein, the network unit quantizes the data of 16 bits or more obtained as a result of the convolution operation to 8 bits or less by any one of comparison with a plurality of thresholds or transformation using a predetermined function.

9. The image processing apparatus according to claim 1 or 2, wherein, the preprocessing unit transforms the pixel values into 8-bit data, and the network unit performs a convolution operation with the 8-bit data transformed by the preprocessing unit as an input.

10. A learning method, having: a preprocessing step of transforming pixel values of a pair of a high-quality image and a low-quality image included in supervised data into a number of bits lower than the number of bits of the pixel values using a predetermined function having non-linearity; and a learning step of performing learning related to extraction of a noise component superimposed on the low-quality image with the data transformed by the preprocessing step as an input.

11. An inference method comprising: a preprocessing step of transforming pixel values of an input input image into a number of bits lower than the number of bits of the pixel values using a predetermined function having nonlinearity; an inference step of taking, as an input, the data transformed by the preprocessing step and performing inference related to extraction of a noise component; and a postprocessing step of performing a process of eliminating nonlinearity using an inverse function of the predetermined function having the nonlinearity with respect to the inferred noise component, and subtracting the noise component from which nonlinearity has been eliminated from the input image, thereby generating an output image with higher image quality than the input image.

12. A learning method comprising: a preprocessing step of transforming pixel values of each of a pair of a high-image-quality image and a low-image-quality image included in supervised data into a number of bits lower than the number of bits of the pixel values using a predetermined function having nonlinearity; and a learning step of taking, as an input, the data transformed by the preprocessing step and performing learning related to extraction of a noise component superimposed on the low-image-quality image and transformation using an inverse function of the predetermined function having the nonlinearity.

13. An inference method comprising: a preprocessing step of transforming pixel values of an input input image into a number of bits lower than the number of bits of the pixel values using a predetermined function having nonlinearity; an inference step of taking, as an input, the data transformed by the preprocessing step and performing inference related to extraction of a noise component and transformation using an inverse function of the predetermined function having the nonlinearity; and a postprocessing step of subtracting the inferred noise component from the input image, thereby generating an output image with higher image quality than the input image.

Citation Information

Patent Citations

  • Information processing device, system, method and program

    JP2022174815A

  • Interpolating visual data

    US10623756B2