Image processing apparatus, control method thereof, and program

The image processing apparatus addresses the challenge of maintaining inference accuracy by storing parameters for each image process in association with specific correction conditions, thereby adapting to changing image characteristics.

JP7696745B2Active Publication Date: 2025-06-23CANON KK
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2021063580
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-07-17
Filing Date
2021-04-02
Publication Date
2025-06-23
Estimated Expiration
2041-04-02

AI Technical Summary

Technical Problem

Existing image processing technologies struggle to maintain inference accuracy when the characteristics of the image change due to different correction conditions during learning and inference.

Method used

An image processing apparatus that includes image processing means, learning means, and control means. The apparatus performs multiple image processes on training and teacher images, learns a model using these images, and stores parameters for each image process in association with the corresponding correction conditions.

Benefits of technology

The solution improves inference accuracy even when image characteristics change due to varying correction conditions, ensuring accurate noise removal and other image processing tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007696745000002
    Figure 0007696745000002
  • Figure 0007696745000003
    Figure 0007696745000003
  • Figure 0007696745000004
    Figure 0007696745000004
Patent Text Reader

Abstract

To improve the inference accuracy of an image even in the case that an image characteristic is changed by correction to be applied to the image.SOLUTION: An image processing device includes image processing means for executing plural image processing of a training image and a teacher image, learning means for performing machine learning of a learning model by using the training image and the teacher image respectively undergoing the plural image processing, and control means for performing control to store a plurality of parameters about the learning model undergoing learning of the respective plural image processing in association with a correction condition of image processing.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing apparatus, a control method of the image processing apparatus, and a program.

Background Art

[0002] In recent years, a method of inferring an image with improved resolution, contrast, etc. using a neural network has been used. As a related technique, the technique of Patent Document 1 has been proposed. The technique of Patent Document 1 calculates the error between a ground truth image and an output image that have been gamma-corrected respectively, and updates the parameters of the neural network based on the calculated error.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Non-Patent Documents

[0004]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] In the technology of Patent Document 1 described above, there is one type of network parameter of the neural network used when making inferences for correcting blurring due to aberration and diffraction. Also, in the technology of Non-Patent Document 1, it is premised that inferences such as noise removal are made using one type of network parameter. Therefore, when using image processing whose characteristics change depending on correction conditions, if the conditions during learning and the correction conditions during inference are different, appropriate inferences cannot be made in noise removal and the like.

[0006] An object of the present invention is to improve the inference accuracy of an image even when the characteristics of the image change depending on the correction applied to the image.

Means for Solving the Problems

[0007] In order to achieve the above object, the image processing apparatus according to claim 1 of the present application includes: image processing means for performing a plurality of image processes on a training image and a teacher image; learning means for performing machine learning of a learning model using the training image and the teacher image on which the plurality of image processes have been respectively performed; and control means for performing control to store a plurality of parameters for the learning model learned in each of the plurality of image processes in association with the correction conditions of the image process. , the image processing includes at least any one of sensitivity correction for correcting analog gain according to the combination of a column amplifier and a lamp, correction for multiplying a gain to an image according to the aperture value so that the relationship between the aperture value of a lens and the amount of light approaches linearity, color suppression processing for correcting saturation, color skew correction for correcting color change, color balance correction, peripheral light fall correction, correction based on the gain difference between channels of a sensor, flicker correction, and compression processing It is characterized by that.

Effects of the Invention

[0008] According to the present invention, even when the characteristics of the image change depending on the correction applied to the image, the inference accuracy can be improved.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Embodiments for Carrying Out the Invention

[0010] Hereinafter, each embodiment of the present invention will be described in detail with reference to the drawings. However, the configurations described in the following embodiments are merely examples, and the scope of the present invention is not limited by the configurations described in each embodiment.

[0011] Hereinafter, the configurations common to each embodiment will be described. FIG. 1 is a diagram showing an example of a system configuration including an image processing apparatus. In the system configuration of FIG. 1, the image processing apparatus 100 is connected to an imaging apparatus 120, a display apparatus 130, and a storage apparatus 140. The system configuration is not limited to the example of FIG. 1. For example, the image processing apparatus 100 may be configured to include either or both of the display apparatus 130 and the storage apparatus 140. Further, the image processing apparatus 100 may be incorporated in the imaging apparatus 120. The image processing apparatus 100 may be an edge computer, a local server, a cloud server, or the like.

[0012] The image processing apparatus 100 performs various image processes, as well as learning processes and inference processes. Hereinafter, the learning process will be described as a process by deep learning. However, the image processing apparatus 100 may perform machine learning using a machine learning method other than deep learning. For example, the image processing apparatus 100 may perform a learning process using any machine learning method such as a support vector machine, a decision tree, or logistic regression.

[0013] In the learning process, a learning set in which a training image and a teacher image are paired is input to the neural network, and network parameters such as weights and biases of the neural network (hereinafter, parameters) are adjusted. By inputting many learning sets to the neural network and performing machine learning, the parameters are optimized so that the feature distribution of the image output by the neural network approaches the feature distribution of the teacher image. When an unknown image is input to the neural network with the learned parameters set, an inference result (inference image) is output from the neural network.

[0014] The imaging device 120 will be described. The imaging device 120 includes a lens unit 121, an imaging element 122, and a camera control circuit 123. The imaging device 120 may include other elements. The lens unit 121 is configured to include a diaphragm, an optical lens, a motor, and the like. The motor drives the diaphragm and the optical lens. The lens unit 121 may be configured to include a plurality of optical lenses. The lens unit 121 operates based on a control signal output by the camera control circuit 123, and can optically magnify / reduce an image, adjust a focal length, and the like. The diaphragm of the lens unit 121 can control the aperture area. Thereby, the diaphragm value (F value) is controlled, and the incident light amount can be adjusted.

[0015] The light that has passed through the lens unit 121 is imaged on the imaging device 122. The imaging device 122 converts the incident light into an electrical signal, such as a CCD sensor or a CMOS sensor. The imaging device 122 is driven based on a control signal from the camera control circuit 123. The imaging device 122 has functions such as resetting the charge in each pixel, controlling the timing of reading, performing gain processing on the read signal, and converting an analog signal into a digital signal. The imaging device 120 transmits the digital image signal output by the imaging device 122 to the image signal receiving circuit 101 of the image processing device 100. In the example of FIG. 1, the imaging device 120 and the image processing device 100 are connected so as to be communicable wirelessly or by wire, and the image processing device 100 receives the image signal transmitted by the imaging device 120. The camera control circuit 123 performs drive control of the lens unit 121 and the imaging device 122 based on the drive control signal received from the camera communication connection circuit 102 of the image processing device 100.

[0016] Next, the image processing device 100 will be described. The image signal receiving circuit 101 receives an image signal from the imaging device 120. The camera communication connection circuit 102 transmits a drive control signal for driving the lens unit 121 and the imaging device 122 to the imaging device 120. The image processing circuit 103 performs image processing on the image signal received by the image signal receiving circuit 101. This image processing may be performed by the CPU 105. Further, the image processing circuit 103 has a function of performing noise reduction processing using a neural network. The image processing circuit 103 may be realized by a circuit such as a microprocessor, a DSP, an FPGA, or an ASIC that performs image processing, or may be realized by the CPU 105.

[0017] The frame memory 104 is a memory that temporarily stores image signals. The frame memory 104 is an element that can temporarily store image signals and read out the stored image signals at high speed. When the data volume of the image signal increases, it is preferable to apply a high-speed and large-capacity memory to the frame memory 104. For example, as the frame memory 104, DDR3-SDRAM (Dual Data Rate 3 - Synchronous Dynamic RAM) or the like can be applied. When DDR3-SDRAM is applied to the frame memory 104, various processes become possible. For example, DDR3-SDRAM is a suitable element for performing image processing such as synthesis of images at different times and extraction of necessary regions. However, the frame memory 104 is not limited to DDR3-SDRAM.

[0018] The CPU 105 reads out the control programs and various parameters stored in the ROM 106 and expands them in the RAM 107. By executing the programs expanded in the RAM 107 by the CPU 105, the processing of each embodiment is realized. The metadata extraction circuit 108 extracts metadata such as lens drive conditions and sensor drive conditions. The GPU 109 is a graphics processing unit and is a processor capable of performing arithmetic processing at high speed. The GPU 109 is used to generate a screen to be displayed on the display device 130, and is also suitably used for arithmetic processing of deep learning. The image generated by the GPU 109 is output to the display device 130 via the display drive circuit 110 and the display device connection circuit 111. Thereby, an image is displayed on the display device 130.

[0019] The storage drive circuit 112 drives the storage device 140 via the storage connection circuit 113. The storage device 140 stores a large amount of image data. The image data stored in the storage device 140 includes training images (learning images) and teacher images (correct answer images) corresponding to the training images. A training image and a teacher image constitute one pair (learning set). A plurality of learning sets are stored in the storage device 140. The storage device 140 may store learned parameters generated by the learning process. Hereinafter, when the learning process is performed under the control of the CPU 105, the image data pre-stored in the storage device 140 is used, and when the inference process is performed, the image signal acquired from the imaging device 120 is used for explanation. However, it is not limited to the above example. For example, when the inference process is performed, an image signal stored in the storage device 140 or an external storage device may be used.

[0020] The training images used in the learning process may be images of a Bayer array, images captured by a three-panel imaging sensor, images captured by an imaging sensor of a vertical color separation method, etc. Also, the training images may be images of an array other than the Bayer array (honeycomb structure, color filter array with low periodicity, etc.). Also, when the training image is an image of the Bayer array, it may be an image of 1ch of the Bayer array, or an image separated for each color channel. The same applies to the teacher images. The training images and the teacher images are not limited to the above-described examples.

[0021] Next, the learning process will be described. FIGS. 2(A) and (B) are block diagrams showing the functions of the learning process of the image processing apparatus 100. First, FIG. 2(A) will be described. The training image holding unit 202 holds the training images acquired from the storage device 201. The teacher image holding unit 203 holds the teacher images acquired from the storage device 201. The training image holding unit 202 and the teacher image holding unit 203 may be realized by, for example, the RAM 107 or the like.

[0022] The teacher image and the training image are images in which the same subject is depicted. The training image is an image containing noise, and the teacher image is a correct image without noise. However, the teacher image may contain some noise. For example, the CPU 105 or the image processing circuit 103 may generate a training image by adding noise to the teacher image. Also, the training image may be an image of the same subject as the correct image taken in a situation where noise can occur (for example, an image taken with a high-sensitivity setting). The training image may be an image in which sensitivity correction has been performed to the same extent as the teacher image for an image taken under low illumination. Also, the teacher image may be an image taken with low sensitivity (lower sensitivity than the training image).

[0023] The neural network 204 is a neural network with a multi-layer structure. Details will be described later. When learning (deep learning) is performed on the neural network 204, in order to improve the inference accuracy, it is preferable to use training images containing various noise patterns and subjects. By using training images containing various noise patterns and subjects to perform learning of the neural network 204, the accuracy of the inference process when an unknown image containing a noise pattern or subject not included in the training image is input is improved. If the number of training images used for deep learning of the neural network 204 is insufficient, for example, augmentation processing such as cropping, rotation, and inversion may be performed on the training image. In this case, the same augmentation processing is also performed on the teacher image. Also, it is preferable that the training image and the teacher image are each divided by the upper limit value of the signal (saturation luminance value) and normalized. When the neural network 204 receives a training image as input, it generates and outputs an output image.

[0024] The image processing unit 205 performs image processing on the output image of the neural network 204 and the teacher image. Examples of image processing include ISO sensitivity correction, F-number correction, color suppression processing, color distortion correction, color balance correction, peripheral light quantity drop correction, inter-channel gain correction, flicker correction, and compression and decompression processing. When the learning process of the neural network 204 is completed, learned parameters are generated. Thereby, an inference process applying the learned parameters to the neural network 204 can be performed. The image processing unit 205 preferably matches the correction conditions in the learning process and the correction conditions in the inference process. Thereby, the inference accuracy when the inference process is performed is improved. In FIG. 2(A), the image processing unit 205 is on the output side of the neural network 204, but as shown in FIG. 2(B), the image processing unit 205 may be on the input side of the neural network 204. The image processing unit 205 is realized by, for example, the image processing circuit 103, the CPU 105, or the like.

[0025] The error evaluation unit 206 calculates the error between the output image corrected by the image processing unit 205 and the teacher image. The teacher image and the training image have the same arrangement of color components. The error evaluation unit 206 may calculate the error (loss error) using, for example, the mean squared error for each pixel or the sum of the absolute values of the differences between the pixels. Further, the error evaluation unit 206 may evaluate the error by calculating the error using other indexes such as the coefficient of determination. The parameter adjustment unit 207 updates each parameter of the neural network 204 so that the error calculated by the error evaluation unit 206 becomes smaller. The error evaluation unit 206 and the parameter adjustment unit 207 may be realized by, for example, the CPU 105, may be realized by the GPU 109, or may be realized by the cooperative operation of the CPU 105 and the GPU 109. The error evaluation unit 206 and the parameter adjustment unit 207 correspond to the learning means.

[0026] The parameter adjustment unit 207 may update each parameter of the neural network 204, for example, using the error backpropagation method. The parameter adjustment unit 207 may update each parameter of the neural network 204 using other methods. At this time, the parameter adjustment unit 207 may fix or vary the update amount of each parameter of the neural network 204. As described above, each parameter of the neural network 204 is updated so that the error between the output image processed by the image processing unit 205 and the teacher image becomes small. Thereby, the inference accuracy when an unknown image is input to the neural network 204 is improved.

[0027] The parameter storage areas 208-1 to 208-n (n is an integer of 2 or more) are areas for storing the updated parameters. The parameter storage areas 208-1 to 208-n (hereinafter collectively referred to as the parameter storage area 208) are areas for storing the parameters of the neural network 204 that has been learned according to the correction conditions of the image processing. That is, in each parameter storage area 208, the learned parameters learned under different image processing correction conditions are stored in association with their respective image processing correction conditions. The parameter storage area 208 may be a part of the storage area of the RAM 107 or a part of the storage area of the storage device 201. The control unit 209 performs various controls related to the learning process. The control unit 209 is realized by the CPU 105.

[0028] Next, the process of learning will be described. FIG. 3 is a flowchart showing an example of the process of learning. In FIG. 3, the steps are denoted as S, which is the same for other flowchart diagrams. In S301, the control unit 209 acquires a pair of a training image and a teacher image (learning set) from the storage device 301. In S302, the control unit 209 inputs the training image of one of the acquired plurality of learning sets into the neural network 204. An output image is generated by the process performed by the neural network 204. The amount of noise and the noise pattern of the training image may be different from or the same as those of other training images.

[0029] In S304, the error evaluation unit 206 calculates the error between the output image subjected to image processing and the teacher image. In S305, the parameter adjustment unit 207 updates each parameter of the neural network 204 using the error backpropagation method so that the calculated error becomes smaller. In S306, the control unit 209 determines whether a predetermined end condition is satisfied. The control unit 209 may determine that the predetermined end condition is satisfied when the number of learning times reaches a specified number. Also, the control unit 209 may determine that the predetermined end condition is satisfied when the calculated error becomes equal to or less than a predetermined value. Further, the control unit 209 may determine that the predetermined end condition is satisfied when the degree of decrease in the above error becomes equal to or less than a predetermined degree or when there is an end instruction from the user. If the control unit 209 determines NO in S306, the flow returns to S301. In this case, the processes of S301 to S305 are performed using a new learning set of a training image and a correct answer image. On the other hand, if the control unit 209 determines YES in S306, the flow proceeds to S307.

[0030] In S303, the image processing unit 205 performs image processing on the output image from the neural network 304 and the teacher image. As described above, the image processing unit 205 may be on the input side or the output side of the neural network 204. That is, the process of S303 may be performed before the process of S302.

[0031] In S307, the control unit 209 stores the updated parameters (learned parameters) in the parameter storage area 208 that is different for each image processing condition. At this time, the control unit 209 may store not only the parameters updated by the learning process but also information regarding the model structure of the neural network 204 etc. associated with the parameters in the parameter storage area 208. In S308, the control unit 209 determines whether it has acquired the parameters for all the correction conditions of the image processing. If the control unit 209 determines YES in S308, it ends the processing of the flowchart in FIG. 3. On the other hand, if the control unit 209 determines NO in S308, it advances the flow to S309. In S309, the control unit 209 changes the correction condition of the image processing. Then, the control unit 209 returns the flow to S301. Thereby, a learning process with the changed image processing conditions is newly performed.

[0032] Next, the neural network 204 will be described. FIG. 4 is a diagram showing an example of the neural network 204. The neural network 204 of each embodiment will be described as being a convolutional neural network (CNN). However, the neural network 204 is not limited to CNN. For example, the neural network 204 may be a GAN (Generative Adversarial Network), a skip connection, an RNN (Recurrent Neural Network), etc. Also, instead of the neural network 204, a learning model learned by another machine learning method may be used. As the learning model, for example, the above-described support vector machine, decision tree, etc. can be applied.

[0033] When an input image 401 is input to the neural network 204, a convolution operation 402 is performed. The input image 401 may be image data or data representing a feature map of the image data. The convolution matrix 403 is a filter that performs a convolution operation on the input image 401. A bias 404 is added to the result output by the convolution operation of the input image 401 and the convolution matrix 403. The feature map 405 is the operation result of the convolution operation to which the bias 404 is added. The number of layers, the number of neurons, the coupling coefficient, the weight, etc. of the intermediate layer of the neural network 204 may be arbitrary values. Also, when the neural network 204 is implemented in a programming circuit such as an FPGA, the connections and weights between neurons may be reduced. This also applies when the GPU 109 performs the processing of the neural network 204. Also, the learning process and the inference process may be executed collectively for a plurality of color channels or individually for each of the plurality of colors.

[0034] In a CNN, by performing a convolution operation on an input image with a filter, a feature map of the input image is extracted. The size of the filter may be any size. In a subsequent layer, by performing a convolution operation on the feature map extracted in the previous layer with another filter, different feature maps are sequentially extracted. In each layer of the intermediate layer, the weight of the filter representing the connection strength is multiplied by the input signal, and a bias is further added. By applying an activation function to this operation result, an output signal at each neuron is obtained. The weights and biases in each layer are parameters (network parameters), and the values of the parameters are updated by the learning process. The parameters updated by machine learning are the learned parameters. As the activation function, any function such as a sigmoid function or a ReLU function can be applied. The activation function applied in each embodiment will be described as the Leaky ReLU function represented by Equation 1 below. However, as the activation function, any activation function such as a sigmoid function or a tanh function may be applied.

[0035]

Number

[0036] As described above, in the parameter storage areas 208-1 to 208-n, parameters (learned parameters) updated by the learning process are stored for each correction condition of the image processing. When the inference process is performed, the parameters corresponding to the correction condition of the image processing are acquired from one of the parameter storage areas 208-1 to 208-n. Then, the acquired parameters are applied to the neural network 204, and the inference process is performed. Thereby, the inference accuracy in the case of performing image processing with different correction conditions can be improved.

[0037] In the above example, the learning process of the neural network 204 is performed using a training image including noise and a teacher image without noise or with small noise. Thereby, when an image including unknown noise is input to the neural network 204 that has been machine-learned, an image with reduced noise can be inferred by the inference process by the neural network 204. The above example is a learning process for noise reduction processing, but it is also applicable to learning processes for other processes than noise reduction. For example, when the inference process of super-resolution is performed by a neural network, the learning process can be performed using a training image obtained by downsampling the teacher image and the teacher image. At this time, the training image and the teacher image may be sized to match.

[0038] Next, the inference process will be described. FIG. 5 is a block diagram showing the functions of the inference process of the image processing apparatus 100. The acquisition unit 501 acquires a captured image and information used when selecting parameters from the imaging apparatus 120. The information used when selecting parameters is used to specify the parameters to be selected from among the parameters stored in the parameter storage areas 208-1 to 208-n. The information used when selecting parameters may be metadata. The metadata may be information added to the captured image (information on setting conditions such as the lens unit 121 and the imaging apparatus 120), or may be information extracted from the metadata extraction circuit 108. Here, the information used when selecting parameters is information related to image processing. When the information used when selecting parameters is the above-described metadata, the acquisition unit 501 does not necessarily need to separately acquire the information used when selecting parameters.

[0039] The parameter selection unit 502 as the selection means acquires a parameter corresponding to the information used when selecting the acquired parameters from any one of the parameter storage areas 208-1 to 208-n. Each parameter stored in the parameter storage areas 208-1 to 208-n is a parameter generated by the above-described learning process, but may be a parameter learned by an external device other than the image processing apparatus 100. In this case, the external device is a neural network having the same network structure as the neural network 204 and performs the above-described learning process. Then, the image processing apparatus 100 may acquire the parameters generated by the learning process and store them in the parameter storage areas 208-1 to 208-n.

[0040] The learned parameters selected by the parameter selection unit 502 are applied to the neural network 204. Then, the captured image acquired from the imaging device 120 is input to the neural network 204. When the captured image is input, the neural network 204 as an inference means performs an inference process and generates an inference image as an inference result. The inference image output unit 503 outputs the generated inference image to the storage device 201. The control unit 504 performs various controls related to the inference process. The acquisition unit 501, the parameter selection unit 502, the inference image output unit 503, and the control unit 504 are realized by, for example, the CPU 105.

[0041] Figure 6 is a flowchart showing an example of the flow of the inference process of the present embodiment. In S601, the acquisition unit 501 acquires information used when selecting parameters. In S602, the parameter selection unit 502 selects parameters corresponding to the acquired information from the parameter storage areas 208-1 to 208-n. Each parameter stored in the parameter storage areas 208-1 to 208-n is a parameter updated under different correction conditions for image processing. The information used when selecting the acquired parameters is information regarding the correction conditions for image processing. The parameter selection unit 502 acquires parameters whose correction conditions for image processing represented by the acquired information used when selecting parameters are the same or close to the correction conditions for image processing.

[0042] In S603, the control unit 504 applies the selected parameters to the neural network 204. As a result, the neural network 204 can perform an inference process. The acquisition unit 501 acquires a captured image from the imaging device 120. It is assumed that the captured image acquired in S604 is a RAW image. When the RAW image is encoded, the CPU 105 or the image processing circuit 103 performs a decoding process. The acquisition unit 501 may acquire the RAW image from the storage device 140, the ROM 106, the RAM 107, or the like.

[0043] In S605, the control unit 504 inputs the acquired captured image into the neural network 204. At this time, the control unit 504 may convert the acquired captured image into an input image for inputting into the neural network 204. The size of the captured image input into the neural network 204 when performing the inference process may be the same as or different from the size of the training image input into the neural network 204 when performing the learning process. When converting the captured image into an input image, the control unit 504 may perform signal normalization, separation for each color component, etc.

[0044] In S606, the neural network 204 performs an inference process. The inference process by the neural network 204 may be executed by the GPU 109, may be executed by the CPU 105, or may be executed by the cooperative operation of the CPU 105 and the GPU 109. As a result of the inference process in S606, an inference image is generated as the inference result of the neural network 204. In S607, the inference image output unit 503 outputs the generated inference image to the storage device 140. The inference image output unit 503 may output the generated inference image to the ROM 106, the RAM 107, the display device 130, etc. In S604, when converting from the captured image to the input image, the control unit 504 may perform an inverse conversion on the inference image based on the conversion performed in S604 to restore it.

[0045] When the inference process for other captured images is performed, the image processing apparatus 100 executes each of the processes in S601 to S607. As described above, in the parameter storage areas 208-1 to 208-n, the parameters of the neural network 204 that have been learned for each correction condition of the image processing are stored. Then, the parameter selection unit 502 selects the corresponding parameters from the parameter storage areas 208-1 to 208-n based on the correction condition of the image processing. Thereby, the inference process of the captured image can be performed using the optimal parameters. As a result, an image with reduced noise can be obtained.

[0046] As described above, the image processing apparatus 100 performs both learning processing and inference processing. In this regard, the image processing apparatus that performs learning processing and the image processing apparatus that performs inference processing may be separate apparatuses. In this case, for example, the image processing apparatus that performs learning processing (hereinafter referred to as the learning apparatus) has the configuration shown in FIG. 2(A) or FIG. 2(B), and the image processing apparatus that performs inference processing (hereinafter referred to as the inference apparatus) has the configuration shown in FIG. 5. In this case, the image processing apparatus that performs learning processing and the image processing apparatus that performs inference processing are communicably connected.

[0047] The inference apparatus acquires the learned parameters from the learning apparatus and stores the acquired parameters in different parameter storage areas 208, similar to the control unit 209 described above. Then, the inference apparatus performs each process of the flowchart in FIG. 6. As described above, even if the image processing apparatus that performs learning processing and the image processing apparatus that performs inference processing are separate apparatuses, the same effects can be obtained as in the case where the image processing apparatus that performs learning processing and the image processing apparatus that performs inference processing are the same apparatus.

[0048] (First Embodiment) In the first embodiment, taking ISO sensitivity correction as an example of image processing, a specific example of the parameters of a learning model in which learning processing is performed using teacher images and training images that have been subjected to ISO sensitivity correction with correction values corresponding to each of a plurality of ISO sensitivity corrections will be described. Also, in the first embodiment, a specific example of the parameters of a learning model in which learning processing is performed taking into account temperature conditions will be described.

[0049] In the first embodiment, the configuration described in the above-described embodiment is used. First, ISO sensitivity correction will be described. ISO sensitivity correction is a correction process that performs sensitivity correction for each combination of a column amplifier and a ramp so that the analog gain determined by the combination of the column amplifier and the ramp becomes the target ISO brightness. The target ISO brightness is the sensitivity standard defined by the International Organization for Standardization. Here, there are a plurality of combinations of the column amplifier and the ramp, and the optimal ISO sensitivity correction value differs for each combination. Therefore, when performing ISO sensitivity correction by inference processing, if only one type of parameter that has been learned is stored, the inference accuracy will decrease. Therefore, the image processing apparatus 100 stores the parameters that have been learned for each combination of the column amplifier and the ramp.

[0050] Also, the optimal correction value of the analog gain also differs depending on the temperature. That is, when performing ISO sensitivity correction by inference processing, the optimal parameters also differ depending on the temperature. Therefore, in this embodiment, the image processing apparatus 100 can also store parameters according to the ISO sensitivity and the temperature.

[0051] FIG. 7 is a block diagram showing the functions of the image processing unit 205 of the present embodiment. FIG. 7(A) is a diagram when the neural network 204 is provided on the input side of the image processing unit 205. FIG. 7(B) is a diagram when the neural network 204 is provided on the output side of the image processing unit 205. In the case of FIG. 7(A), a teacher image is input to the image processing unit 205 from the teacher image holding unit 203, and an output image is input from the neural network 204. In the case of FIG. 7(B), a training image is input to the image processing unit 205 from the training image holding unit 202, and a teacher image is input from the teacher image holding unit 203. In FIG. 7(A), an output image processed by the neural network 204 with the training image as an input is input to the image processing unit 205. On the other hand, in FIG. 7(B), a training image that has not been processed by the neural network 204 is input. The following processing is the same in FIGS. 7(A) and 7(B). As shown in FIGS. 7(A) and 7(B), the image processing unit 205 includes an extraction unit 702, a correction value acquisition unit 703, and a correction unit 704.

[0052] The extraction unit 702 extracts metadata from the teacher image. The metadata is information added to the teacher image and includes ISO sensitivity information about the ISO sensitivity and temperature information about the temperature. In the present embodiment, the extraction unit 702 extracts ISO sensitivity information and temperature information as the metadata of the teacher image. The metadata may be information added to the image in, for example, the EXIF (Exchange Image File Format) format. However, the metadata is not limited to data in the EXIF format.

[0053] The correction value acquisition unit 703 acquires data on ISO sensitivity correction corresponding to the extracted ISO sensitivity information and temperature information from the ROM 106. The ROM 106 stores data on correction values corresponding to combinations of ISO sensitivity and temperature. The ROM 106 may store data on correction values corresponding to ISO sensitivity instead of the above combinations. The correction unit 704 performs ISO sensitivity correction on the output image and the teacher image output by the neural network 204 using the data on ISO sensitivity correction acquired by the correction value acquisition unit 703. The output image and the teacher image subjected to ISO sensitivity correction are output to the error evaluation unit 206.

[0054] Next, a first example of ISO sensitivity correction will be described. FIG. 8 is a diagram showing a first example of analog gain. FIG. 8(A) is a diagram showing an example of an ideal analog gain. FIG. 8(B) is a diagram showing an example of an actual analog gain. The analog gain is determined by the combination of a column amplifier and a ramp. The analog gain shown in FIG. 8(B) has variations compared to the analog gain in FIG. 8(A). Therefore, the correction unit 704 corrects the analog gain.

[0055] For example, when focusing on ISO 1, as shown in the example of FIG. 8(A), when the column amplifier is "0.6 times" and the ramp is "0.85 times", the analog gain "0.51 times (= 0.6 × 0.85)" is the ideal value. On the other hand, for the actual analog gain, in the example of FIG. 8(B), since the column amplifier is "0.59 times" and the ramp is "0.84 times", the analog gain is "0.50 times (= 0.59 × 0.84)". Therefore, the correction value for ISO sensitivity correction for ISO 1 is "1.02 times (= 0.51 / 0.50)".

[0056] When focusing on ISO9, as shown in the example of Fig. 8(A), when the column amplifier is "8.0 times" and the lamp is "1.00 times", the analog gain "8.00 times (= 8.0 × 1.00)" is the ideal value. On the other hand, for the actual analog gain, in the example of Fig. 8(B), since the column amplifier is "7.88 times" and the lamp is "0.97 times", the analog gain is "7.64 times (= 7.88 × 0.97)". Therefore, the correction value for ISO sensitivity correction for ISO1 is "1.047 times (= 8.00 / 7.64)".

[0057] As described above, the correction value for ISO sensitivity correction differs for each ISO sensitivity. Therefore, the correction unit 704 performs ISO sensitivity correction using the correction value corresponding to the ISO sensitivity information of the metadata extracted from the teacher image. In the ROM106, the correction value for ISO sensitivity correction for each ISO sensitivity is stored. The correction value for ISO sensitivity correction for each ISO sensitivity may be stored in the RAM107, the storage device 140, or the like. In the example of Fig. 8(B), an example where the actual analog gain is lower than the ideal analog gain is shown, but the actual analog gain may also be higher than the ideal analog gain.

[0058] Next, a second example of ISO sensitivity correction will be described. Fig. 9 is a diagram showing a second example of the analog gain. The analog gain shown in Fig. 9 has variations in characteristics depending on the temperature. For this reason, the correction value for ISO sensitivity correction also differs depending on the temperature. This is due to the temperature characteristics of the channel amplifier. Therefore, the ROM106 may store the correction value for each combination of each ISO sensitivity and each temperature.

[0059] Here, when the ROM106 stores the correction values for all temperatures, the number of correction values stored in the ROM106 also increases. In this case, the number of combinations of each ISO sensitivity and each temperature increases, and the number of parameters stored in the parameter storage area 208 also increases. As a result, the amount of data used for the parameter storage area 208 becomes large. Therefore, the ROM106 may store the correction values for temperatures in groups for each predetermined range.

[0060] For example, the ROM 106 may store common correction values for three types of temperature ranges: a first temperature range (low temperature below 0°C), a second temperature range (temperature from 0°C to 40°C), and a third temperature range (high temperature above 40°C). When correction values for three types of temperature ranges are used for each of the above-mentioned ISOs as described above, a total of 33 (= 11×3) correction values are stored in the ROM 106. Thereby, for a plurality of temperatures (for example, from 0°C to 40°C), since one common correction value is used, the number of parameters stored in the parameter storage area 208 can be reduced, and the data usage amount can be decreased. Note that the method of setting the temperature range, the number of types, etc. are not limited to the above example. For example, the temperature range may be variably set according to the degree of change in the analog gain due to temperature. In this case, the temperature range may be set narrowly in a section where the degree of the above change is equal to or more than a predetermined degree, and may be set widely in a section where the degree is less than the predetermined degree.

[0061] Next, an example of the flow of the learning process in the first embodiment will be described. FIG. 10 is a flowchart showing an example of the flow of the learning process in the first embodiment. The flowchart of FIG. 10 corresponds to the block diagram of FIG. 7(A). In S1001, the control unit 209 acquires a pair of a training image and a teacher image (learning set) from the storage device 301. The process of S1001 corresponds to the process of S301 in FIG. 3. In S1002, the control unit 209 inputs the training image of one learning set among the acquired plurality of learning sets to the neural network 204. An output image is generated by the process performed by the neural network 204. The process of S1002 corresponds to the process of S302 in FIG. 3.

[0062] In S1003, the extraction unit 702 of the image processing unit 205 extracts metadata from the teacher image. In S1004, the correction value acquisition unit 703 acquires a correction value corresponding to the combination of the ISO sensitivity information and the temperature information of the extracted metadata from the ROM 106. In S1005, the correction unit 704 performs ISO sensitivity correction on the teacher image and the output image output from the neural network 204 with the acquired correction value. The output image output from the neural network 204 is an image obtained by processing the training image by the neural network 204.

[0063] In S1006, the error evaluation unit 206 calculates the error between the output image subjected to image processing and the teacher image. The process of S1006 corresponds to the process of S304 in FIG. 3. In S1007, the parameter adjustment unit 207 updates each parameter of the neural network 204 using the error backpropagation method so that the calculated error becomes smaller. The process of S1006 corresponds to the process of S305 in FIG. 3. In S1008, the control unit 209 determines whether a predetermined end condition is satisfied. The process of S1008 corresponds to the process of S306 in FIG. 3. If the control unit 209 determines NO in S1008, since the predetermined end condition is not satisfied, the flow returns to S1001. On the other hand, if the control unit 209 determines YES in S1008, since the predetermined end condition is satisfied, the flow proceeds to S1009.

[0064] In S1009, the control unit 209 stores the updated parameters (learned parameters) in parameter storage areas 208 that are different for each correction condition of ISO sensitivity correction. The correction conditions of the ISO sensitivity correction in the present embodiment are correction values corresponding to combinations of each ISO sensitivity and each temperature range. When temperature information is not used, the correction conditions of the ISO sensitivity correction in the present embodiment are correction values corresponding to each ISO sensitivity. The control unit 209 stores the updated parameters in parameter storage areas 208 that are different for each combination of each ISO sensitivity and each temperature range. The process of S1009 corresponds to S307 in FIG. 3. In S1010, the control unit 209 determines whether it has acquired the parameters for all conditions. At this time, the control unit 209 makes the determination in S1010 based on whether it has acquired all combinations for each ISO and each temperature range. If the control unit 209 determines YES in S1010, it ends the process of the flowchart in FIG. 10. On the other hand, if the control unit 209 determines NO in S1010, it advances the flow to S1011. In S1011, the control unit 209 changes the conditions of the combination of ISO and temperature range. Then, the control unit 209 returns the flow to S1001. Thereby, a learning process with the combination of ISO and temperature range changed is newly performed.

[0065] Next, another example of the processing flow in the first embodiment will be described. FIG. 11 is a flowchart showing another example of the processing flow of the learning process in the first embodiment. The flowchart in FIG. 11 corresponds to the block diagram in FIG. 7(B). In the processing of the flowchart in FIG. 10, ISO sensitivity correction is performed on the output image output from the neural network 204, and the learning process is performed. On the other hand, in the processing of the flowchart in FIG. 11, ISO sensitivity correction is performed on the training image input to the neural network 204, and the learning process is performed.

[0066] The process of S1101 corresponds to the process of S1001. In S1102, the extraction unit 702 of the image processing unit 205 extracts metadata from the teacher image. In S1103, the correction value acquisition unit 703 acquires, from the ROM 106, a correction value corresponding to the combination of the ISO sensitivity information and the temperature information of the extracted metadata. The process of S1102 corresponds to the process of S1003, and the process of S1103 corresponds to the process of S1004. In S1104, the correction unit 704 performs ISO sensitivity correction on the teacher image and the training image. That is, the correction unit 704 performs ISO sensitivity correction on the training image on which the process of the neural network 204 has not been performed.

[0067] In S1105, the control unit 209 inputs the training image subjected to ISO sensitivity correction to the neural network 204. An output image is generated by the process performed by the neural network 204 using the training image subjected to ISO sensitivity correction as an input. In S1106, the error evaluation unit 206 calculates the error between the teacher image subjected to ISO sensitivity correction and the output image generated in S1105. In S1107, the parameter adjustment unit 207 updates each parameter of the neural network 204 using the error backpropagation method so that the calculated error becomes smaller. Since the processes of S1108 to S1111 are the same as the processes of S1008 to S1011 in FIG. 10, the description thereof is omitted.

[0068] The flowchart of FIG. 10 and the flowchart of FIG. 11 show the flow of the learning process regarding the combination of the ISO sensitivity information and the temperature information in the first embodiment. In S1004 of FIG. 10 or S1103 of FIG. 11, the correction value acquisition unit 703 acquires, from the ROM 106, a correction value corresponding to the combination of the ISO sensitivity information and the temperature information of the extracted metadata. At this time, the correction value acquisition unit 703 may acquire the ISO sensitivity information from the ROM 106 without acquiring the temperature information. For example, when the ISO sensitivity information is not stored in the ROM 106, the correction value acquisition unit 703 acquires the ISO sensitivity information from the ROM 106 and does not acquire the temperature information. Even in this case, ISO sensitivity correction can be performed using the respective correction values for each ISO.

[0069] Next, the inference process of the first embodiment will be described. FIG. 12 is a flowchart showing an example of the flow of the inference process of the first embodiment. In S1201, the acquisition unit 501 acquires ISO sensitivity information and temperature information. The acquisition unit 501 can acquire ISO sensitivity information and temperature information from the imaging device 120. For example, based on the setting conditions of the imaging device 120, the acquisition unit 501 can acquire ISO sensitivity information. Further, when the imaging device 120 has a temperature measurement function such as a thermistor, the acquisition unit 501 can acquire temperature information from the temperature measured by the thermistor of the imaging device 120. The acquisition unit 501 may acquire ISO sensitivity information and temperature information from the metadata of the image data stored in the storage device 140.

[0070] In S1202, the parameter selection unit 502 selects a parameter corresponding to the combination of the acquired ISO sensitivity information and temperature information from the parameter storage areas 208-1 to 208-n. In S1203, the control unit 504 applies the selected parameter to the neural network 204. As a result, it becomes possible to perform the inference process by the neural network 204. Since the processes in S1204 to S1207 are the same as those in S604 to S607 in FIG. 6, the description thereof will be omitted.

[0071] In the first embodiment, parameters of the neural network 204 that have undergone learning processing according to each ISO sensitivity are stored in the parameter storage areas 208-1 to 208-n. Then, the parameter selection unit 502 selects the corresponding parameters from the parameter storage areas 208-1 to 208-n based on the ISO sensitivity condition. As a result, inference processing of a captured image using parameters that match the ISO sensitivity condition can be performed, improving the inference accuracy. Consequently, an image with reduced noise can be obtained. Also, in the first embodiment, learning processing that takes into account temperature conditions is performed. That is, parameters of the neural network 204 that have undergone learning processing according to each combination of each ISO and each temperature range can also be stored in the parameter storage areas 208-1 to 208-n. Then, the parameter selection unit 502 selects the corresponding parameters from the parameter storage areas 208-1 to 208-n based on the conditions of both ISO sensitivity and temperature. As a result, inference processing of a captured image using parameters that match the characteristics of ISO sensitivity and temperature characteristics can be performed, further improving the inference accuracy. Consequently, the noise reduction effect of the image can be further improved.

[0072] (Second Embodiment) In the second embodiment, taking F-value correction as an example of image processing, a specific example of the parameters of a learning model that has undergone learning processing using a teacher image and a training image in which F-value correction is performed with correction values corresponding to each of a plurality of F-value corrections will be described.

[0073] F-value correction is a correction that applies a gain to an image so that the relationship between the F-value, which is the aperture value, and the amount of light (luminance) approaches linearity when the relationship is not linear. FIG. 13 is a graph showing an example of the relationship between the F-value and the amount of light, and the relationship between the F-value and the correction value in the second embodiment. FIG. 13(A) is a graph showing the relationship between the F-value and the amount of light. The horizontal axis represents the F-value, and the vertical axis represents the amount of light received by the imaging element 122 of the imaging device 120. As described above, F-value correction is a process of applying a gain according to the aperture value (F-value) of the lens unit 121. Here, when the aperture is changed from the small-aperture side to the open side, the relationship between the F-value and the amount of light (luminance) is linear on the small-aperture side, but the relationship between the F-value and the amount of light (luminance) is non-linear on the open side. One of the factors causing the non-linear relationship between the F-value and the amount of light (luminance) on the open side is that the light incident obliquely on the imaging element 122 cannot be sufficiently captured on the open side. Therefore, in the section where the F-value is on the open side, F-value correction is performed to multiply the image signal by a gain in order to approach an ideal linearity. Since F-value correction is performed in the section on the open side, noise occurs in the image.

[0074] In FIG. 13(A), the relationship between the F-value and the amount of light in the section of the F-value from Fx to Fs changes linearly as represented by the solid line. On the other hand, the relationship between the F-value and the amount of light in the section from Fe to Fx changes non-linearly as represented by the dashed-dotted line. Therefore, as described above, a correction value for removing noise is multiplied in the section from Fe to Fx (hereinafter, the non-linear section) so that the relationship between the F-value and the amount of light becomes linear even in the non-linear section. The correction value in the non-linear section is not a constant value as shown in FIGS. 13(A) and 13(B). On the other hand, a correction value of a constant value is multiplied in the section from Fx to Fs (hereinafter, the linear section). That is, the correction value in the non-linear section needs to be changed to an appropriate value according to the F-value so that the relationship between the F-value and the amount of light becomes linear. In FIG. 13(B), the correction value in the non-linear section is indicated by the dotted line.

[0075] Next, an example of the flow of the learning process in the second embodiment will be described. FIG. 14 is a flowchart showing an example of the flow of the learning process in the second embodiment. The flowchart of FIG. 14 corresponds to the block diagram of FIG. 7(A), and the correction unit 704 of the image processing unit 205 performs F-value correction. In S1401, the control unit 209 newly sets the F-value when performing the learning process. In S1402, the control unit 209 acquires a pair of training image and teacher image (learning set) from the storage device 301. In S1403, the control unit 209 inputs the training image of one of the acquired plurality of learning sets to the neural network 204. An output image is generated by the process performed by the neural network 204. The noise amount and noise pattern of the training image may be different from or the same as those of other training images.

[0076] In S1404, the image processing unit 205 performs F-value correction on the teacher image and the output image from the neural network 204. At this time, the image processing unit 205 performs F-value correction with a correction value corresponding to the F-value set in S1401. As described above, the image processing unit 205 may be on the output side of the neural network 204 as shown in FIG. 2(A), or may be on the input side as shown in FIG. 2(B). That is, the process of S1401 may be performed before the process of S1403. Note that if the extraction unit 702 of the image processing unit 205 is configured to acquire the F-value from the metadata of the teacher image, S1401 can be omitted. In S1405, the error evaluation unit 206 calculates the error between the output image subjected to F-value correction and the teacher image. In S1406, the parameter adjustment unit 207 updates each parameter of the neural network 204 using the error backpropagation method so that the calculated error becomes small. In S1407, the control unit 209 determines whether a predetermined end condition is satisfied. If the control unit 209 determines NO in S1407, the flow returns to S1402. In this case, the processes of S1402 to S1406 are performed using a new learning set of training image and correct image. On the other hand, if the control unit 209 determines YES in S1407, the flow proceeds to S1408.

[0077] In S1408, the control unit 209 determines whether the F value set in S1401 is in a linear section. The linear section will be described later. If the control unit 209 determines YES in S1408, the flow proceeds to S1409. In S1409, the control unit 209 stores the updated parameters (learned parameters) in the same parameter storage area 208. On the other hand, if the control unit 209 determines NO in S1408, the flow proceeds to S1410. In S1410, the control unit 209 stores the updated parameters (learned parameters) in different parameter storage areas 208 according to the F value set in S1401. In S1409 and S1410, the control unit 209 may store information regarding the parameters updated by machine learning, the model structure of the neural network 204, etc., in association with the conditions of the F value set in S1401, in the parameter storage area 208. In S1411, the control unit 209 determines whether it has acquired the parameters for all the conditions of the F values. If the control unit 209 determines NO in S1411, the flow returns to S1401 to newly set the conditions of the F value. If the control unit 209 determines YES in S1411, the processing of the flowchart in FIG. 14 is terminated.

[0078] In this way, the control unit 209 performs control to change the F value to be corrected and cause the neural network 204 to perform machine learning (deep learning). Thereby, machine learning is performed reflecting the correction value for each F value, and the parameters are updated. Then, in S1408 of FIG. 14, the control unit 209 determines whether the F value set when performing machine learning is within a linear section. The control unit 209 can perform the determination in S1408, for example, based on whether the set F value is equal to or greater than the above-described Fx. The value of Fx (predetermined threshold) may be set to any value, but is preferably set at the boundary where the relationship between the F value and the light amount changes from linear to non-linear. However, the value of Fx may be set to a value within a predetermined range from the boundary.

[0079] Then, when the control unit 209 determines YES in S1408, it stores the updated learned parameters in the same parameter storage area 208 among the parameter storage areas 208-1 to 208-n. Even if the updated parameters are for different F values, when the control unit 209 determines YES in S1408, it stores them in the same parameter storage area 208-1. Alternatively, if the control unit 209 has already stored the updated parameters once when it determines YES in S1408, it may omit the process of S1409 for the second and subsequent times. On the other hand, when the control unit 209 determines NO in S1408, it sets different F values and stores the updated learned parameters in different parameter storage areas 208 respectively. Each time the control unit 209 executes the process of S1410, it sequentially stores them in different parameter storage areas 208 (for example, parameter storage areas 208-2 to 208-n).

[0080] Therefore, in the section where the F value and the light amount (luminance) have a non-linear relationship, the inference process can be executed using the parameters of the neural network 204 corresponding to the F values with fine granularity. Thereby, the inference accuracy is improved. On the other hand, in the section where the F value and the light amount (luminance) are linear, the inference process is executed using the parameters of one neural network 204. As a result, since the number of parameters to be stored can be reduced, the amount of data used can be reduced, and the amount of use of hardware resources such as the storage device 140 and the RAM 107 can be reduced. That is, since the image processing apparatus 100 changes the granularity of the parameters to be stored according to the F value, it can reduce the amount of data used while suppressing a decrease in the inference accuracy.

[0081] In the above example, the image processing apparatus 100 stores the parameters of the machine-learned neural network 204 for each F value in the non-linear section, but the present invention is not limited to this example. For example, in the non-linear section, the image processing apparatus 100 may store the learned parameters for each driving resolution of the aperture provided in the lens unit 121 of the imaging apparatus 120. Further, the image processing apparatus 100 performs both the learning process and the inference process. In this regard, the image processing apparatus that performs the learning process and the image processing apparatus that performs the inference process may be separate apparatuses. This point is the same as that of the first embodiment. Examples in which the image processing apparatus that performs the learning process and the image processing apparatus that performs the inference process are separate apparatuses are the same in each of the following embodiments.

[0082] In this case, the inference apparatus acquires and stores one learned parameter in the linear section, and acquires and stores the learned parameters corresponding to each of a plurality of aperture values in the non-linear section. Then, the inference apparatus performs each process of the flowchart of FIG. 6. As described above, even if the image processing apparatus that performs the learning process and the image processing apparatus that performs the inference process are separate apparatuses, the same effects can be obtained.

[0083] Next, a modification of the second embodiment will be described. FIG. 15 is a graph showing an example of the relationship between the F value and the correction value in the modification of the second embodiment. In the first embodiment, the correction value in the section where the F value is in the range of Fx to Fs is constant, but in the second embodiment, the correction value in the section where the F value is in the range of Fx to Fs changes linearly. The section where the F value is in the range of Fx to Fs is a section where the relationship between the F value and the light amount (luminance) is linear. In the second embodiment, the control unit 209 stores the parameter (first parameter) obtained by setting the F value to Fx and performing the process of FIG. 14 and the parameter (second parameter) obtained by setting the F value to Fs and performing the process of FIG. 14 in the parameter storage area 208.

[0084] Then, as shown in FIG. 15, when the inference process of the captured image when the F value is Fz is performed, the control unit 504 interpolates using the inference image using the first parameter and the inference image using the second parameter. As a result, an inference image for any F value in the section of Fx to Fs where the correction value changes linearly can be obtained. In this case, the number of parameters (machine-learned parameters) stored to obtain the inference image for any F value in the section of Fx to Fs is two.

[0085] Here, the control unit 504 may perform the process of FIG. 14 using any two F values in the section of Fx to Fs where the correction value changes linearly, and store the first parameter and the second parameter. Even in this case, by interpolating using the inference image using the first parameter and the inference image using the second parameter, an inference image for any F value in the section of Fx to Fs where the correction value changes linearly can be obtained.

[0086] FIG. 16 is a flowchart showing an example of the flow of the inference process in a modification of the second embodiment. In S1601, the acquisition unit 501 acquires information on the F value from the imaging device 120. In S1602, the control unit 504 determines whether the F value indicated by the information acquired in S1601 is within a linear section. As described above, when the F value indicated by the information acquired in S1601 is equal to or greater than Fx (equal to or greater than a predetermined threshold), the control unit 504 may determine NO in S1602, and when the F value is less than Fx (less than a predetermined threshold), the control unit 504 may determine YES in S1602. When the control unit 504 determines NO in S1602, in S1603, the parameter selection unit 502 selects a parameter corresponding to the acquired F value information from the parameter storage areas 208-1 to 208-n. Each parameter stored in the parameter storage areas 208-1 to 208-n is a parameter updated under different F value conditions. The parameter selection unit 502 acquires a parameter for an F value that is the same as the acquired F value information or a parameter for the F value with the closest value.

[0087] In S1604, the control unit 504 applies the selected parameters to the neural network 204. As a result, the neural network 204 can perform an inference process. In S1605, the acquisition unit 501 acquires a captured image from the imaging device 120. This captured image is assumed to be a RAW image. If the RAW image has been subjected to an encoding process, the CPU 105 or the image processing circuit 103 performs a decoding process. The acquisition unit 501 may acquire the RAW image from the storage device 140, the ROM 106, the RAM 107, or the like.

[0088] In S1606, the control unit 504 inputs the acquired captured image into the neural network 204. In S1607, the neural network 204 performs an inference process. The inference process by the neural network 204 may be executed by the GPU 109, may be executed by the CPU 105, or may be executed by a cooperative operation of the CPU 105 and the GPU 109. As a result of the inference process in S1607, an inference image is generated as the inference result of the neural network 204. In S1608, the inference image output unit 503 outputs the generated inference image to the storage device 140. The inference image output unit 503 may output the generated inference image to the ROM 106, the RAM 107, the display device 130, or the like. At this time, the parameters selected in S1603 are learned parameters that have been learned using F values within a non-linear section.

[0089] When the control unit 504 determines NO in S1602, the flow proceeds to S1609. In S1609, the parameter selection unit 502 selects the first parameter described above. For example, when the first parameter is stored in the parameter storage area 208-1, the parameter selection unit 502 selects the first parameter from the parameter storage area 208-1. In S1610, the neural network 204 executes an inference process to which the first parameter is applied. The process of S1610 corresponds to the processes of S1604 to S1607. That is, an inference process using the captured image as an input is executed for the neural network 204 to which the first parameter is applied. As a result, an inference image (first inference image) is generated by the inference process in which the first parameter is applied to the neural network 204.

[0090] In S1611, the parameter selection unit 502 selects the second parameter described above. For example, when the second parameter is stored in the parameter storage area 208-2, the parameter selection unit 502 acquires the second parameter from the parameter storage area 208-2. Then, the neural network 204 executes an inference process to which the second parameter is applied. The process of S1612 corresponds to the processes of S1604 to S1607. That is, an inference process using the captured image as an input is executed for the neural network 204 to which the second parameter is applied. As a result, an inference image (second inference image) is generated by the inference process in which the second parameter is applied to the neural network 204.

[0091] In S1613, the control unit 504 executes interpolation processing on the inference image. At this time, the control unit 504 synthesizes the first inference image and the second inference image and executes interpolation processing. Also, the control unit 504 executes interpolation processing by weighting the first inference image and the second inference image based on the F value indicated by the information acquired in S1601. For example, when the F value indicated by the information acquired in S1601 is close to the F value corresponding to the first parameter, the control unit 504 may execute interpolation processing with a heavier weight on the first inference image. On the other hand, when the F value indicated by the information acquired in S1601 is close to the F value corresponding to the second parameter, the control unit 504 may execute interpolation processing with a heavier weight on the second inference image. When interpolating the first inference image and the second inference image, the control unit 504 may perform interpolation for each pixel, or may perform interpolation for each area obtained by arbitrarily dividing the entire image into a certain range. Then, in S1608, the control unit 504 outputs the inference image on which the interpolation processing has been executed.

[0092] As described above, in the second embodiment, even if the correction value is not constant in the linear section, the same effect as in the first embodiment can be obtained. In the second embodiment, two parameters, namely, the first parameter and the second parameter on which the learning process has been performed, are used. The F value corresponding to the first parameter and the F value corresponding to the second parameter when the learning process is performed may be selected according to the linearity of the correction value. Also, in the second embodiment, the control unit 504 may further divide the section (the section of Fx to Fs) in which the F value and the light amount (luminance) are linear into a plurality of sections. In this case, the control unit 504 may use the divided section as one processing unit and execute interpolation processing for the two parameters for each processing unit.

[0093] (Third Embodiment) In the third embodiment, taking color suppression processing and color distortion correction as examples of image processing, specific examples of the parameters of the learning model for which learning processing is performed using a teacher image and a training image on which image processing has been performed with correction values corresponding to each of a plurality of color suppression processes or a plurality of color distortion corrections will be described.

[0094] When a high-brightness subject is photographed by an imaging device, any of the RGB signals output from the image sensor may be in a saturated state. In this case, the ratio of the RGB signals may change from the value representing the color of the original subject. To address this issue, a method is known in which color suppression processing, which is a process of reducing the saturation in an area where any of the RGB signals is saturated, is applied to mitigate the color change in the high-brightness portion caused by signal saturation.

[0095] Also, when the imaging device photographs a low-brightness and high-saturation subject, any of the RGB signals may approach the black level, which is the lower limit. In this case, if gamma correction with a steep rise at the black level is applied to the RGB signals, the noise components superimposed on the RGB signals are emphasized by the gamma correction, causing image quality degradation. To address this issue, a method is known in which color suppression processing is applied to a lower limit region where any of the RGB signals takes a value close to the lower limit, thereby increasing the signal value of the signal with a small signal value among the RGB signals. This can reduce the degree of emphasis of the noise components by gamma correction.

[0096] Since the color suppression processing is a process of reducing the saturation, the noise in the area where the processing is applied is reduced. Therefore, by performing the color suppression processing during learning, a neural network that takes into account the noise reduced by the color suppression processing can be obtained. Also, the characteristics of the color suppression processing change depending on the gamma correction and gamut conversion processes. Therefore, when the output format input during learning and the output format of the image input during inference are different, the characteristics of the color suppression processing are also different, resulting in a decrease in the inference accuracy of the inference image.

[0097] In addition, for the images viewed by the user, gamma correction and color gamut conversion processes are performed according to the output format. At this time, depending on the color of the subject, colors different from the original subject may be output. In order to suppress this color change, the process of correcting by multiplying a gain for each color is called color bending correction. The way the color changes depends on the gamma correction and color gamut conversion processes. That is, the characteristics of color bending correction change depending on the output format. Therefore, when the output format input during learning is different from the output format of the image input during inference, the characteristics of color bending correction are also different, resulting in a decrease in the inference accuracy of the inference image.

[0098] Next, an example of the flow of the learning process in the third embodiment will be described. FIG. 17 is a flowchart showing an example of the flow of the learning process in the third embodiment. The flowchart of FIG. 17 can be realized using, for example, the block diagrams of FIGS. 2 and 7(A). The correction unit 704 of the image processing unit 205 performs the development process and color suppression process described later. Here, the color suppression process will be described as an example of the image process, but the same applies when color bending correction is used instead of the color suppression process.

[0099] In S1701, the control unit 209 acquires a pair of training images and teacher images (learning set) from the storage device 301. In S1702, the control unit 209 inputs the training image of one of the acquired plurality of learning sets into the neural network 204. An output image is generated by the process performed by the neural network 204. In S1703, the extraction unit 702 of the image processing unit 205 extracts metadata regarding gamma, color gamut, or gamma and color gamut from the teacher image. The gamma may be based on performing color adjustment such as color grading later, or may be for the purpose of being viewed by the user (PQ, HLG, etc.). Also, the color gamut may be BT.2020, BT.709, DCI-P3, or may be adjusted according to the display characteristics of the display device 130 (2.2, 2.4, etc.). Also, combinations of these may be used.

[0100] In S1704, the correction value acquisition unit 703 acquires, from the ROM 106, a correction value corresponding to the gamma, color gamut, or a combination of gamma and color gamut of the extracted metadata. In S1705, the correction unit 704 performs development processing on the teacher image and the output image output from the neural network 204, thereby converting them into a color space divided into luminance and hue, such as YCC or ICtCp. Note that the development processing may include debayering processing performed on a Bayer array image. Further, the correction unit 704 performs color suppression correction on the teacher image and the output image to which the development processing has been applied, using the acquired correction value.

[0101] In S1706, the error evaluation unit 206 calculates the error between the output image that has undergone image processing and the teacher image. In S1707, the parameter adjustment unit 207 updates each parameter of the neural network 204 using the error backpropagation method so as to reduce the calculated error. In S1708, the control unit 209 determines whether a predetermined end condition is satisfied. If the control unit 209 determines NO in S1708, since the predetermined end condition is not satisfied, the flow returns to S1701. On the other hand, if the control unit 209 determines YES in S1708, since the predetermined end condition is satisfied, the flow proceeds to S1709.

[0102] In S1709, the control unit 209 stores the updated parameters (learned parameters) in the parameter storage area 208 that is different for each correction condition of the color suppression process. The correction condition of the color suppression process in the present embodiment is a correction value corresponding to the combination of each gamma and each color gamut. The control unit 209 stores the updated parameters in the parameter storage area 208 that is different for each combination of each gamma and each color gamut. If there is no change in the color gamut, the parameters that are different for each gamma are stored in the parameter storage area 208. In S1710, the control unit 209 determines whether it has acquired the parameters for all conditions. At this time, the control unit 209 makes the determination in S1710 based on whether it has acquired all combinations for each gamma and each color gamut. If the control unit 209 determines YES in S1710, it ends the processing of the flowchart in FIG. 17. On the other hand, if the control unit 209 determines NO in S1710, it advances the flow to S1711. In S1711, the control unit 209 changes the conditions of the combination of gamma and color gamut. Then, the control unit 209 returns the flow to S1701. As a result, a learning process with the combination of gamma and color gamut changed is newly performed. Note that in FIG. 17, the training image was input to the neural network 204, and development processing and color suppression processing were performed on the output image from the neural network 204, but it is not limited to this. It may be configured such that after the development processing and the color suppression processing are performed on the training image, the training image on which these processes have been performed is input to the neural network 204. In this case, the output image from the neural network 204 and the teacher image to which the development processing and the color suppression processing have been applied are input to the error evaluation unit 206.

[0103] Next, the inference processing of the third embodiment will be described. FIG. 18 is a flowchart showing an example of the flow of the inference processing of the third embodiment. In S1801, the acquisition unit 501 acquires gamma and color gamut information. Alternatively, the acquisition of the captured image in S1804 may be performed first, and the gamma and color gamut information may be acquired from the metadata of the captured image. In S1802, the parameter selection unit 502 selects parameters corresponding to the acquired gamma and color gamut information from the parameter storage areas 208-1 to 208-n. In S1803, the control unit 504 applies the selected parameters to the neural network 204. As a result, the neural network 204 can perform inference processing. Since the processes of S1804 to S1807 are the same as those of S604 to S607 in FIG. 6, the description thereof will be omitted.

[0104] In addition to gamma correction and color gamut conversion processing, in color balance correction for adjusting exposure correction, gain correction, and saturation correction, etc., by preparing parameters for the neural network according to the combination of these correction values, the same effect can be obtained.

[0105] (Fourth Embodiment) In the fourth embodiment, taking peripheral light fall correction as an example of image processing, a specific example of the parameters of the learning model in which learning processing is performed using teacher images and training images that have been subjected to peripheral light fall correction with correction values corresponding to each of a plurality of peripheral light fall corrections will be described.

[0106] Peripheral light fall correction is a correction in which a gain is applied to an image so as to correct the lens characteristic that the light amount gradually decreases from the center to the periphery of the lens, and the gain becomes a larger value as it goes from the center to the periphery of the lens. Also, the characteristics of peripheral light fall vary depending on the type of lens, aperture, focal length, and zoom position conditions, and the optimal gain varies for each region.

[0107] Figures 19(a) and 19(b) are diagrams showing the state of peripheral light quantity drop. The horizontal axis is the image height which is the distance from the center position of the lens, with the image height at the center position of the lens being 0% and the image height at the outer periphery of the image circle of the lens being 100%. The vertical axis is the light quantity, with the light quantity at the center position of the lens being 100%. In Figures 19(a) and 19(b), the lens conditions (aperture, focus, zoom position) are different. In Figure 19(a), the light quantity at the position corresponding to an image height of 15% is close to 100%, and the light quantity at the position corresponding to an image height of 100% is 80%. In Figure 15(b), the light quantity at the position corresponding to an image height of 70% is close to 100%, and the light quantity at the position corresponding to an image height of 100% is 90%. Thus, depending on the state of the lens, the starting point where peripheral light quantity drop begins and the shape such as the inclination are different.

[0108] Figure 19(c) shows the gain curve for correcting the peripheral light quantity drop shown in Figure 19(a), and Figure 19(d) shows the gain curve for correcting the peripheral light quantity drop shown in Figure 19(b). The horizontal axis is the image height, and the vertical axis is the gain. The gain curve in Figure 19(c) has a shape for making the light quantity 100% at each image height, and a gain of about 1.25 times is set at the position of image height 100%. The gain curve in Figure 19(d) has a gain of about 1.1 times set at the position of image height 100%. Since the noise component included in the image is amplified more as the gain is increased, an upper limit may be set for the gain.

[0109] Next, an example of the flow of the learning process in the fourth embodiment will be described. Figure 20 is a flowchart showing an example of the flow of the learning process in the fourth embodiment. The flowchart in Figure 20 can be realized, for example, using the block diagrams in Figures 2 and 7(A). The correction unit 704 of the image processing unit 205 performs peripheral light quantity drop correction described later.

[0110] In S2001, the control unit 209 acquires a pair of a training image and a teacher image (learning set) from the storage device 301. In S2002, the control unit 209 inputs the training image of one of the acquired plurality of learning sets into the neural network 204. An output image is generated by the process performed by the neural network 204. In S2003, the extraction unit 702 of the image processing unit 205 extracts metadata regarding the type of lens, aperture, zoom position, and focal length from the teacher image.

[0111] In S2004, the correction value acquisition unit 703 acquires a correction value corresponding to the combination of the type of lens, aperture, zoom position, and focal length that has been extracted from the ROM 106. In S2005, the correction unit 704 performs peripheral light fall correction on the teacher image and the output image output from the neural network 204 with the acquired correction value.

[0112] In S2006, the error evaluation unit 206 calculates the error between the output image that has been subjected to image processing and the teacher image. Then, in S2007, the parameter adjustment unit 207 updates each parameter of the neural network 204 using the error backpropagation method so that the calculated error becomes smaller. In S2008, the control unit 209 determines whether a predetermined end condition is satisfied. If the control unit 209 determines NO in S2008, since the predetermined end condition is not satisfied, the flow returns to S1001. On the other hand, if the control unit 209 determines YES in S2008, since the predetermined end condition is satisfied, the flow proceeds to S2009.

[0113] In S2009, the control unit 209 stores the updated parameters (learned parameters) in parameter storage areas 208 that are different for each correction condition of the peripheral light quantity drop correction. The correction conditions for the peripheral light quantity drop correction in the present embodiment are correction values corresponding to combinations of the type of lens, aperture, zoom position, and focal length. The control unit 209 stores the updated parameters in parameter storage areas 208 that are different for each combination of the type of lens, aperture, zoom position, and focal length. Note that if the lens is not an interchangeable lens but is a lens integrated with the imaging device, it is not necessary to store parameters according to the type of lens. Also, if the lens is a single-focus lens without a zoom function, it is not necessary to store parameters according to the zoom position. In S2010, the control unit 209 determines whether it has acquired the parameters for all conditions. If the control unit 209 determines YES in S2010, it ends the processing of the flowchart in FIG. 20. On the other hand, if the control unit 209 determines NO in S2010, it advances the flow to S2011. In S2011, the control unit 209 changes the conditions of the combination of the type of lens, aperture, zoom position, and focal length. Then, the control unit 209 returns the flow to S2001.

[0114] FIG. 21 is a diagram showing an example of table data of peripheral light quantity drop correction data. Table data 2100 is a data group at a focal length of A1. The horizontal axis is the aperture value (S1, S2, S3 ··· Sn), and the vertical axis is the zoom position (T1, T2, T3, ··· Tn). For example, the peripheral light quantity drop correction data at the zoom position T1, aperture value S1, and focal length A1 is Dt1s1a1. Also, table data 2101, 2102, 2103, 2104 are data groups at focal lengths of A1, A2, A3, An, respectively. The peripheral light quantity drop correction data corresponding to each lens condition is stored in the ROM 106. Also, in the case of an interchangeable lens, since the shape of the peripheral light quantity drop is different for each lens, table data is provided for each lens. The interchangeable lens can be identified by the lens ID assigned to each lens.

[0115] Next, the inference process of the fourth embodiment will be described. FIG. 22 is a flowchart showing an example of the flow of the inference process of the fourth embodiment. In S2201, the acquisition unit 501 acquires information on the type of lens, aperture, zoom position, and focal length. Alternatively, the acquisition of the captured image in S2204 may be performed first, and information on the type of lens, aperture, zoom position, and focal length may be acquired from the metadata of the captured image. In S2202, the parameter selection unit 502 selects parameters corresponding to the acquired information on the type of lens, aperture, zoom position, and focal length from the parameter storage areas 208-1 to 208-n. Then, in S2203, the control unit 504 applies the selected parameters to the neural network 204. As a result, the neural network 204 can perform the inference process. Since each process of S2204 to S2207 is the same as S604 to S607 in FIG. 6, the description thereof will be omitted.

[0116] (Fifth Embodiment) In the fifth embodiment, taking inter-channel gain correction as an example of image processing, a specific example of the parameters of a learning model in which learning processing is performed using a teacher image and a training image in which inter-channel gain correction is performed with correction values corresponding to each of a plurality of inter-channel gain corrections will be described.

[0117] Inter-channel gain correction is a process of correcting the gain of the output amplifier arranged at the final stage of the analog output of the sensor. Since this output amplifier is an analog amplifier, there is a gain difference for each amplifier. There is also a difference in gain between a plurality of output amplifiers within the same sensor. In addition, it has the characteristic of changing according to temperature. Therefore, if only one network parameter is used, the noise difference due to the gain difference of the output amplifier caused by the temperature condition is not taken into account, and a sufficient noise removal effect cannot be obtained. That is, an optimal network parameter suitable for the inter-channel gain correction characteristics is required for each temperature condition.

[0118] FIG. 23 is a diagram showing a sensor structure. A vertical line 2303 is connected to a photodiode 2301 that converts the intensity of light into voltage via a switch 2302. The vertical line 2303 is connected to output amplifiers 2305 to 2308 via a switch 2304. By switching the on and off states of these switches 2302 and 2304, the voltage signal obtained from the photodiode 2301 is read from the output amplifiers 2305 to 2308. The inter-channel gain correction in this embodiment is a function of correcting the gains of these four output amplifiers 2305 to 2308, and these four output amplifiers 2305 to 2308 are called channels. For example, here, the output amplifier 2305 is channel A, the output amplifier 2306 is channel B, the output amplifier 2307 is channel C, and the output amplifier 2308 is channel D.

[0119] FIG. 24 is a diagram showing the output voltage of each channel with respect to the amount of light incident on the photodiode 2301.

[0120] In FIG. 24(A), graph 2401 represents the output of channel A, graph 2402 represents the output of channel B, graph 2403 represents the output of channel C, and graph 2404 represents the output of channel D. Since the output amplifier is an analog amplifier, as shown in FIG. 24, there is a difference in output levels. The reasons for this difference in output levels are variations in offset and gain between channels. After the variation in offset is corrected, the variation in gain is corrected. FIG. 24(B) is a diagram showing the output voltages of each channel in a state where the variation in offset is corrected so that the black levels in FIG. 24(A) match. Graphs 2411, 2412, 2413, and 2414 are graphs in which the black levels of graphs 2401, 2402, 2403, and 2404 are corrected, respectively. In the inter-channel gain correction, a sensor is uniformly irradiated with a certain amount of light indicated by line 2415 to obtain one image, and the average of the output voltages of each channel is calculated from a 400-pixel × 400-pixel region at the center position of the acquired pixel output. The offset is subtracted from the average value of the output voltages of each channel, and a correction gain for each channel is calculated so that the subtraction result becomes a constant value. This calculation of the correction gain is performed for each case where the sensor temperature is 60 degrees and 30 degrees. The correction gain thus obtained is associated with the sensor number indicating the individual sensor, the channel number (numbers of channels A to D), and the temperature information, and is stored in ROM106 as a correction value.

[0121] Next, an example of the flow of the learning process in the fifth embodiment will be described. FIG. 25 is a flowchart showing an example of the flow of the learning process in the fifth embodiment. The flowchart of FIG. 25 corresponds to the block diagram of FIG. 7(A), and the correction unit 704 of the image processing unit 205 performs inter-channel gain correction. In S2501, the control unit 209 acquires a pair of a training image and a teacher image (learning set) from the storage device 301. In S2502, the control unit 209 inputs the training image of one of the acquired plurality of learning sets into the neural network 204. An output image is generated by the process performed by the neural network 204. In S2503, the extraction unit 702 of the image processing unit 205 extracts metadata from the teacher image. In S2504, the correction value acquisition unit 703 acquires a correction value corresponding to the combination of the sensor number and temperature information of the extracted metadata from the ROM 106. In S2505, the correction unit 704 performs inter-channel gain correction on the teacher image and the output image output from the neural network 204 with the acquired correction value.

[0122] In S2506, the error evaluation unit 206 calculates the error between the output image that has undergone image processing and the teacher image. In S2507, the parameter adjustment unit 207 updates each parameter of the neural network 204 using the error backpropagation method so that the calculated error becomes smaller. In S2508, the control unit 209 determines whether a predetermined end condition is satisfied. If the control unit 209 determines NO in S2508, since the predetermined end condition is not satisfied, the flow returns to S2501. On the other hand, if the control unit 209 determines YES in S2508, since the predetermined end condition is satisfied, the flow proceeds to S2509.

[0123] In S2509, the control unit 209 stores the updated parameters (learned parameters) in the parameter storage area 208 that is different for each sensor number and temperature information correction condition. In this embodiment, the correction condition for the inter-channel gain correction is a correction value corresponding to the combination of the sensor number and the temperature information. The control unit 209 stores the updated parameters in the parameter storage area 208 that is different for each combination of the sensor number and the temperature information. In S2510, the control unit 209 determines whether it has acquired the parameters for all conditions. At this time, the control unit 209 makes the determination in S2510 based on whether it has acquired all combinations of the sensor number and the temperature information. Note that if different parameters are prepared for each of the plurality of imaging devices, it is only necessary to acquire the parameters corresponding to the sensor number corresponding to the imaging device. If the control unit 209 determines YES in S2510, it ends the processing of the flowchart in FIG. 25. On the other hand, if the control unit 209 determines NO in S2510, it advances the flow to S2511. In S2511, the control unit 209 changes the condition of the combination of the sensor number and the temperature information. Then, the control unit 209 returns the flow to S2501. As a result, a new learning process in which the combination of the sensor number and the temperature information is changed is performed.

[0124] Next, the inference process of the fifth embodiment will be described. FIG. 26 is a flowchart showing an example of the flow of the inference process of the fifth embodiment. In S2601, the acquisition unit 501 acquires the sensor number and temperature information. The acquisition unit 501 can acquire the sensor number and temperature information from the imaging device 120. In S2602, the parameter selection unit 502 selects the parameters corresponding to the acquired sensor number from the parameter storage areas 208-1 to 208-n. In this embodiment, since only the correction value at 30 degrees and the parameters at 60 degrees are stored as temperature information, both the correction value at 30 degrees and the parameters at 60 degrees corresponding to the sensor number are selected. Of course, not only 30 degrees and 60 degrees, but also parameters corresponding to many temperatures may be prepared in advance, and the parameters corresponding to the temperature closest to the acquired temperature may be used. In S2603, the control unit 504 applies each of the selected parameters to the neural network 204. As a result, the neural network 204 can perform the inference process.

[0125] In S2604, the acquisition unit 501 acquires the captured image. In S2605, the control unit 504 inputs the acquired captured image into the neural network 204. In S2606, the neural network 204 performs the inference process. In this embodiment, the inference process using the parameters at 30 degrees is performed to generate the first inference image, and the inference process using the parameters at 60 degrees is performed to generate the second inference image. In S2607, the inference image output unit 503 interpolates the first inference image and the second inference image based on the acquired temperature information to generate an inference image corresponding to the temperature information. Specifically, let t be the acquired temperature information, α be the coefficient of weighted addition, Z1 be the signal level of the coordinate (x, y) in the first inference image, and Z2 be the signal level of the coordinate (x, y) in the second inference image. In this case, the signal level Z OUT (x,y) of the inference image obtained by interpolating the first inference image and the second inference image can be obtained using the following equations 2 and 3. α = (-1 / (60 - 30)) × t + 2 ··· (Equation 2) Z OUT Z(x,y) = α × Z1(x,y) + (1 - α) × Z2(x,y) ···(Equation 3)

[0126] In addition, when the temperature information is less than 30 degrees, it may be regarded as if the temperature information is 30 degrees, and when the temperature information exceeds 60 degrees, it may be regarded as if the temperature information is 60 degrees. In S2608, the inference image output unit 503 outputs the inference image generated in S2206 to the storage device 140. The inference image output unit 503 may output the generated inference image to the ROM 106, the RAM 107, the display device 130, etc.

[0127] As described above, in the first to fifth embodiments, as image processing, specific examples of the parameters of the learning model are described by taking ISO sensitivity correction, F-value correction, color suppression processing, color distortion correction, color balance correction, peripheral light quantity drop correction, and inter-channel gain correction as examples, but it is not limited thereto. For example, as an example of image processing different from the image processing described so far, there is flicker correction. Flicker correction is a process of correcting the difference in luminance values generated between lines or between frames of a sensor due to the flickering of a light source such as a fluorescent lamp. The magnitude of the flicker can be detected from the amount of change (amplitude) of the luminance value changing by line, or the frequency of the flicker can be detected from the period of the luminance value changing for each line. It has the characteristic that the way of generation is different depending on conditions such as flicker, the brightness of the light source, the frequency of flickering, and the accumulation time of the sensor. That is, parameters of the learning model may be prepared for each of these conditions. Even in such a case, the same effects as those of the first to fifth embodiments can be obtained.

[0128] (Sixth Embodiment) In the sixth embodiment, compression processing and decompression processing are taken as examples of image processing, and specific examples of the parameters of the learning model obtained by performing learning processing using a teacher image and training images subjected to compression processing and decompression processing at each of a plurality of compression ratios will be described. In this embodiment, since compression processing is performed on the image input to the neural network 204, an effect different from those of the first to fifth embodiments can be obtained, that is, the circuit scale of the neural network can be suppressed.

[0129] In this embodiment, the image processing unit 205 performs compression processing by reducing the pixel value to 1 / m and decompression processing by reducing the pixel value to 1 / n. Note that the method of compression processing is not limited to this method, and any other method may be used as long as it is the same as the compression processing used in the inference processing step.

[0130] FIG. 27 is a block diagram showing the learning processing functions of the image processing apparatus 100 according to the sixth embodiment. In the sixth embodiment, the difference from the first embodiment is that the image processing unit 205 is arranged on the input side and the output side of the neural network 204. A training image is input to the image processing unit 205, and the image processing unit 205 performs compression processing on the training image at any one of a plurality of compression ratios. The image compressed by the image processing unit 205 is input to the neural network 204, and the neural network 204 generates and outputs an output image. Then, the output image is input to the image processing unit 205 again, and the image processing unit 205 performs decompression processing corresponding to the compression ratio in the compression processing on the output image. Then, the error evaluation unit 206 calculates the error between the output image expanded by the image processing unit 205 and the teacher image.

[0131] The learned parameters compressed at different compression ratios are stored in the respective parameter storage areas 208-1 and 208-2. Here, an example with two compression ratios has been described, but the number of compression ratios may be three or more. When there are n types of compression ratios, the parameter storage areas 208-1 to 208-n are used.

[0132] Next, an example of the flow of the learning process in the sixth embodiment will be described. FIG. 28 is a flowchart showing an example of the flow of the learning process in the first embodiment. In S2801, the control unit 209 acquires a pair of a training image and a teacher image (learning set) from the storage device 301. In S2802, the control unit 209 acquires compression information. In the present embodiment, the compression information is information in which, for each pixel value of the training image, a value of 0 is set when the pixel value is less than a predetermined threshold value, and a value of 1 is set when the threshold value is equal to or more than the threshold value, and is composed of data having the same number as the number of pixels of the training image. Note that the compression information is not limited to information in such a data format, and may be information in another data format as long as it is information that allows the compression processing method of the training image to be understood. Further, the control unit 209 may acquire the compression information stored in advance in either the imaging device 120 or the storage device 140 in step S2802. Further, without acquiring the compression information in step S2802, the image processing unit 205 may directly calculate the compression information from the training image in the following step S2803.

[0133] In step S2803, the control unit 209 causes the image processing unit 205 to perform compression processing on the training image based on the compression information acquired in step S202. In the image processing unit 205, compression processing is performed by setting the pixel value to 1 / m when the compression information is 0, and setting the pixel value to 1 / n when the compression information is 1. Note that the method of compression processing is not limited to this method, and other methods may be used as long as they are the same as the compression processing used in the inference processing step. Further, the training image acquired in step S2801 may already be compressed. In this case, since it is not necessary to perform compression processing on the training image, the process directly proceeds from step S2801 to step S2804 described later. Further, when there are a plurality of operation modes of the imaging device 120 and the compression processing methods are different for each, the compression information acquired in step S2802 may be changed or the presence or absence of compression processing may be switched according to the operation mode.

[0134] In S2804, the control unit 209 inputs the training image compressed in step S2803 into the neural network 204 to generate an output image. Subsequently, in step S2805, the image processing unit 205 performs an expansion process on the output image generated in step S2804, and the error evaluation unit 206 calculates the error between the expanded output image and the correct image.

[0135] In the present embodiment, the image processing unit 205 performs an expansion process that is the reverse process of the compression process performed on the training image in step S2803. Specifically, when the compression information is 0, the pixel value is multiplied by m, and when the compression information is 1, the pixel value is multiplied by n. However, the method of the expansion process is not limited to the method of the present embodiment, and any other method may be used as long as it is the same as the expansion process used in the inference process. Also, when there are multiple operation modes of the imaging device 120 and the processing methods of the compression process and the expansion process are different for each, the processing methods of the compression process and the expansion process, and the presence or absence of the compression process and the expansion process may be switched according to the operation mode. By using the same expansion process method for the expansion process used when expanding the inference image and the expansion process performed on the output image in step S2805, inference can be performed with more stable accuracy regardless of the amount of noise after the expansion process. Subsequently, in step S2806, the parameter adjustment unit 207 updates each parameter of the neural network 204 using the error backpropagation method so that the calculated error becomes smaller. In S2807, the control unit 209 determines whether a predetermined end condition is satisfied. If the control unit 209 determines NO in S2807, the flow returns to S2801. In this case, the processes of S2801 to S2806 are performed using a new learning set of a training image and a correct image. On the other hand, if the control unit 209 determines YES in S2807, the flow proceeds to S2808.

[0136] In step S2808, the control unit 209 stores information regarding the updated parameters, the structure of the neural network, etc. in the parameter storage area 208-1 or 208-2. In S2809, the control unit 209 determines whether it has acquired the parameters for all the compression information. If the control unit 209 determines NO in S2809, it returns the flow to S2801 and acquires other compression information in S2802. If the control unit 209 determines YES in S2809, it ends the processing of the flowchart in FIG. 28. Through the above operations, without increasing the circuit scale, a neural network can be obtained in which the inference accuracy is hardly affected for the decompressed image.

[0137] FIG. 29 is a block diagram showing the inference processing function of the image processing apparatus 100. It is different in that an image processing unit 205 is arranged between the acquisition unit 501 and the neural network 204, and between the neural network 204 and the inference image output unit 503, compared to the blocks shown in FIG. 5. This image processing unit 205 performs compression processing and decompression processing of images.

[0138] FIG. 30 is a flowchart showing an example of the flow of inference processing in the sixth embodiment. In S3001, the acquisition unit 501 acquires compression information. The acquisition unit 501 can acquire compression information from the imaging device 120. For example, based on the setting conditions of the imaging device 120, the acquisition unit 501 can acquire compression information. This compression information is in the same format as the compression processing used in the learning process. Specifically, in this embodiment, for the pixel values of the captured image, information is set such that if the pixel value is less than a predetermined threshold, it is 0, and if it is greater than or equal to the threshold, it is 1. Note that the compression information is not limited to this format, and may be information in other data formats as long as it is the same as the compression processing used in the learning process. In S3002, the parameter selection unit 502 selects the parameters corresponding to the acquired compression information from the parameter storage area 208-1 or 208-2. Then, in S3003, the control unit 504 applies the selected parameters to the neural network 204. Thereby, it becomes possible to perform inference processing by the neural network 204.

[0139] In S3004, the acquisition unit 501 acquires a captured image from the imaging device 120. This captured image is an uncompressed RAW image and is composed of data having the same number of pixels as the training image. Also, the compression information acquired at this time is in the same format as the compression process used in the learning process. The acquisition unit 501 may acquire the RAW image from the storage device 140. In S3005, the compression unit 1901 performs a compression process on the acquired captured image according to the compression information acquired in S3001. When the compression information is 0, the image processing unit 205 performs the compression process by setting the pixel value to 1 / m, and when the compression information is 1, by setting the pixel value to 1 / n. Note that the method of the compression process is not limited to this method, and any other method may be used as long as it can reduce the data amount of the captured image. Also, the captured image acquired in S3004 may already be compressed. In this case, since the compression process of the captured image does not need to be performed, the process proceeds from S3004 to S3006. Further, when there are a plurality of operation modes of the imaging device 120 and the compression process methods are different for each, the compression information acquired in step S3001 may be changed or the presence or absence of the compression process may be switched according to the operation mode.

[0140] In S3006, the control unit 504 inputs the compressed captured image into the neural network 204. In S3007, the neural network 204 performs an inference process to generate an inference image. In S3008, the image processing unit 205 executes an expansion process on the inference image generated in step S3007. In the present embodiment, the image processing unit 205 performs an expansion process that is the reverse of the compression process performed on the captured image in step S3005. Specifically, when the compression information is 0, the pixel value is multiplied by m, and when the compression information is 1, the pixel value is multiplied by n. However, the method of the expansion process is not limited to the method of the present embodiment, and other methods may be used as long as they are the same as the expansion process used in the learning process. Also, if the expansion process is separately performed in a later stage, the expansion process may not be performed here. In S3009, the inference image output unit 503 outputs the expanded inference image to the storage device 140. The inference image output unit 503 may output the generated inference image to the ROM 106, the RAM 107, the display device 130, etc.

[0141] According to the present embodiment, since the compression process is performed on the image input to the neural network 204, the circuit scale of the neural network can be suppressed. Also, after performing the expansion process on the output image from the neural network 204, the update of each network parameter is performed in the learning process, so that the decrease in the noise suppression effect of the neural network for inference due to the above compression process can be suppressed.

[0142] Note that although the first to sixth embodiments have been described, the parameters of the neural network may be used according to the combination of the correction conditions in each embodiment. That is, parameters corresponding to correction conditions combining two or more of ISO sensitivity correction, F-value correction, color suppression processing, color bending correction, color balance correction, peripheral light amount drop correction, inter-channel gain correction, flicker correction, and compression and expansion processing may be used.

[0143] As described above, the preferred embodiments of the present invention have been explained. However, the present invention is not limited to the above-described embodiments, and various modifications and changes are possible within the scope of the gist thereof. The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or apparatus via a network or a storage medium, and having one or more processors of a computer of the system or apparatus read and execute the program. Further, the present invention can also be realized by a circuit (for example, ASIC) that realizes one or more functions.

Explanation of Reference Numerals

[0144] 100 Image processing apparatus 103 Image processing circuit 105 CPU 106 RAM 109 GPU 108 Metadata extraction circuit 120 Imaging device 140 Storage device 204 Neural network 205 Image processing unit 206 Error evaluation unit 207 Parameter adjustment unit 208 Parameter storage area 501 Acquisition unit 502 Parameter selection unit

Claims

1. image processing means for performing a plurality of image processes on training images and teacher images; learning means for performing machine learning of a learning model using the training images and the teacher images on which the plurality of image processes have been respectively performed; control means for performing control to store a plurality of parameters regarding the learning model learned in each of the plurality of image processes in association with correction conditions of the image process; comprising: The image process is at least any one of sensitivity correction for correcting an analog gain according to a combination of a column amplifier and a lamp, correction for multiplying an image by a gain according to the aperture value so that the relationship between the aperture value of the lens and the amount of light approaches linearity, color suppression processing for correcting saturation, color bend correction for correcting color change, color balance correction, peripheral light fall correction, correction based on a gain difference between channels of a sensor, flicker correction, and compression processing. An image processing apparatus characterized by this.

2. selection means for selecting one parameter from the plurality of stored parameters according to information regarding correction conditions of the image process; inference means for applying the selected parameter to the learning model and performing inference processing on an image; The image processing apparatus according to claim 1, further comprising:

3. The training image on which the image process is performed is characterized in that it is being processed by the learning model. The image processing apparatus according to claim 1 or 2.

4. The training image is an image with noise added thereto with respect to the teacher image. The image processing apparatus according to any one of claims 1 to 3.

5. The image processing is at least a combination of any two or more of sensitivity correction for correcting analog gain according to a combination of a column amplifier and a lamp, correction for multiplying a gain to an image according to an aperture value so that the relationship between the aperture value of a lens and the amount of light approaches linearity, color suppression processing for correcting saturation, color bend correction for correcting color change, color balance correction, peripheral light amount drop correction, correction based on a gain difference between channels of a sensor, flicker correction, and compression processing. The image processing apparatus according to any one of claims 1 to 4, characterized in that it is such a combination.

6. The image processing is ISO sensitivity correction for correcting analog gain according to a combination of a column amplifier and a lamp. The image processing apparatus according to any one of claims 1 to 4, characterized in that it is such correction.

7. The correction condition of the image processing is ISO sensitivity, and the control means stores the plurality of parameters in association with the ISO sensitivity. The image processing apparatus according to claim 6, characterized in that it is such storage.

8. The correction condition of the image processing is a combination of ISO sensitivity and temperature, and the control means stores the plurality of parameters in association with the combination of ISO sensitivity and temperature. The image processing apparatus according to claim 6, characterized in that it is such storage.

9. The image processing is correction for multiplying a gain to an image according to an aperture value so that the relationship between the aperture value of a lens and the amount of light approaches linearity. The image processing apparatus according to any one of claims 1 to 4, characterized in that it is such correction.

10. The correction condition of the image processing is the aperture value, and the control means stores the plurality of parameters in association with the aperture value. The image processing apparatus according to claim 9, characterized in that it is such storage.

11. The control means makes the granularity of the parameters to be stored different between a section where the relationship between the aperture value and the gain is linear and a non-linear section. The image processing apparatus according to claim 9 or 10, characterized in that it is such a difference.

12. The image processing according to any one of claims 1 to 4 is characterized in that it is at least one of a color suppression process for correcting saturation and a color bend correction for correcting color changes.

13. The correction condition of the image processing is at least one of gamma and color gamut, and the control means stores the plurality of parameters in association with at least one of gamma and color gamut. The image processing apparatus according to claim 12, characterized in that.

14. The image processing according to any one of claims 1 to 4 is characterized in that it is a peripheral light quantity drop correction.

15. The correction condition of the image processing is at least one of the type of lens, aperture value, zoom position, and focal length, and the control means stores the plurality of parameters in association with at least one of the type of lens, aperture value, zoom position, and focal length. The image processing apparatus according to claim 14, characterized in that.

16. The image processing according to any one of claims 1 to 4 is characterized in that it is a correction based on the gain difference between channels of the sensor.

17. The correction condition of the image processing is the sensor and temperature, and the control means stores the plurality of parameters in association with the individual of the sensor and temperature. The image processing apparatus according to claim 16, characterized in that.

18. The image processing according to any one of claims 1 to 4 is characterized in that it is a flicker correction.

19. The correction condition of the image processing is at least one of the brightness of the light source, the frequency of the light source flickering, and the accumulation time of the sensor, and the control means stores the plurality of parameters in association with at least one of the brightness of the light source, the frequency of the light source flickering, and the accumulation time of the sensor. The image processing apparatus according to claim 18, characterized in that.

20. The image processing apparatus according to claim 1, wherein the image processing is compression processing.

21. The correction condition of the image processing is a compression rate, and the control means stores the plurality of parameters in association with the compression rate. The image processing apparatus according to claim 20.

22. Selection means for selecting one parameter from the plurality of stored parameters according to the compression rate; Inference means for applying the selected parameter to the learning model to perform inference processing on the image compressed at the compression rate; The image processing apparatus according to claim 21, further comprising:

23. Storage means for storing a plurality of parameters for a plurality of learning models learned in each of a plurality of image processings in different areas; Selection means for selecting one parameter from the plurality of stored parameters according to information regarding the correction condition of the image processing; Inference means for applying the selected parameter to the learning model to perform inference processing on the image; Comprising The image processing includes at least one of sensitivity correction for correcting analog gain according to a combination of a column amplifier and a lamp, correction for multiplying a gain to an image according to the aperture value so that the relationship between the aperture value of the lens and the amount of light approaches linearity, color suppression processing for correcting saturation, color bend correction for correcting color change, color balance correction, peripheral light fall correction, correction based on a gain difference between channels of a sensor, flicker correction, and compression processing. An image processing apparatus characterized by this.

24. A step of performing a plurality of image processings on a training image and a teacher image; A step of performing machine learning of a learning model using the training image and the teacher image on which the plurality of image processings have been respectively performed; A step of storing a plurality of parameters for the learning models learned in each of the plurality of image processes in association with correction conditions of the image process; comprising; The image process is at least any one of sensitivity correction for correcting an analog gain according to a combination of a column amplifier and a lamp, correction for multiplying an image by a gain according to the aperture value so that the relationship between the aperture value of the lens and the amount of light approaches linearity, color suppression processing for correcting saturation, color bend correction for correcting color change, color balance correction, peripheral light quantity drop correction, correction based on a gain difference between channels of a sensor, flicker correction, and compression processing. A control method for an image processing apparatus, characterized by this.

25. A step of storing a plurality of parameters for the learning models learned in each of the plurality of image processes in different areas respectively; A step of selecting one parameter from the plurality of stored parameters according to information regarding correction conditions of the image process; A step of applying the selected parameter to the learning model and performing an inference process on an image; comprising; The image process is at least any one of sensitivity correction for correcting an analog gain according to a combination of a column amplifier and a lamp, correction for multiplying an image by a gain according to the aperture value so that the relationship between the aperture value of the lens and the amount of light approaches linearity, color suppression processing for correcting saturation, color bend correction for correcting color change, color balance correction, peripheral light quantity drop correction, correction based on a gain difference between channels of a sensor, flicker correction, and compression processing. A control method for an image processing apparatus, characterized by this.

26. A program for causing a computer to execute each means of the image processing apparatus according to any one of Claims 1 to 23.

Citation Information

Patent Citations

  • Image processing method, image processing apparatus, and storage medium

    CN110021047A

  • Image recognition device

    JP2012068965A

  • Image processing method, image processing apparatus, image processing program and storage medium

    JP2019121252A

  • Image processing method, image processing device, imaging device, lens device, program, and storage medium

    JP2020030569A

  • Improved video stream delivery via adaptive quality enhancement using error correction models

    US20180288440A1