Information processing device, imaging device, information processing method, and program
The information processing device integrates AI and non-AI methods to optimize image quality by synthesizing images processed differently, addressing the limitations of neural networks in adjusting noise and capturing images from varying devices.
Patent Information
- Application Number
- JP2024066779
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-01-29
- Filing Date
- 2024-04-17
- Publication Date
- 2025-09-08
- Estimated Expiration
- 2042-01-18
AI Technical Summary
Existing image processing systems using neural networks struggle to effectively adjust image quality and remove noise, particularly when dealing with images captured by different imaging devices with varying parameters and conditions.
An information processing device that combines AI and non-AI methods to process images, using a neural network for some adjustments and non-neural network methods for others, with weighted synthesis of images processed differently to optimize image quality.
Enhances image quality by adjusting noise using both AI and non-AI methods, allowing for improved image synthesis and compensation for differences in imaging devices, resulting in better image quality and sharpness.
Smart Images

Figure 0007735470000003 
Figure 0007735470000004 
Figure 0007735470000005
Abstract
Description
[Technical Field]
[0001] The technology disclosed herein relates to an information processing device, an imaging device, an information processing method, and a program. [Background technology]
[0002] JP 2018-206382 A discloses an image processing system that uses a neural network having an input layer, an output layer, and an intermediate layer provided between the input layer and the output layer to perform processing on an input image input to the input layer, and an adjustment unit that adjusts at least one internal parameter of one or more nodes included in the intermediate layer, the internal parameter being calculated by learning, based on data related to the input image when processing after learning.
[0003] Furthermore, in the image processing system described in JP 2018-206382 A, the input image is an image containing noise, and the noise is removed or reduced from the input image by processing performed by the processing unit.
[0004] In addition, in the image processing system described in JP 2018-206382 A, the neural network includes a first neural network, a second neural network, a division unit that divides an input image into a high-frequency component image and a low-frequency component image and inputs the high-frequency component image to the first neural network while inputting the low-frequency component image to the second neural network, and a synthesis unit that synthesizes a first output image output from the first neural network and a second output image output from the second neural network, and the adjustment unit adjusts internal parameters of the first neural network based on data related to the input image while not adjusting internal parameters of the second neural network.
[0005] Furthermore, Japanese Patent Application Laid-Open No. 2018-206382 discloses an image processing system that includes a processing unit that uses a neural network to generate a noise-reduced output image from an input image, and an adjustment unit that adjusts the internal parameters of the neural network according to the imaging conditions of the input image.
[0006] Japanese Patent Publication No. 2020-166814 discloses a medical image processing device that includes an acquisition unit that acquires a first image, which is a medical image of a specific part of a subject's body, an image quality improvement unit that uses an image quality improvement engine including a machine learning engine to generate a second image from the first image, which has higher image quality than the first image, and a display control unit that displays a composite image obtained by combining the first image and the second image using a ratio obtained using information regarding at least a portion of the area of the first image on a display unit.
[0007] JP 2020-184300 A discloses an electronic device including a memory that stores at least one instruction word, and a processor electrically connected to the memory that executes the instruction word to obtain a noise map indicating the quality of the input image from the input image, applies the input image and the noise map to a learning network model including multiple layers, and obtains an output image with improved quality of the input image, wherein the processor provides the noise map to at least one intermediate layer of the multiple layers, and the learning network model is a trained artificial intelligence model obtained by learning the relationship between multiple sample images, the noise map for each sample image, and the original image for each sample image through an artificial intelligence algorithm. Summary of the Invention
[0008] One embodiment of the technology disclosed herein provides an information processing device, an imaging device, an information processing method, and a program that can obtain images with adjusted image quality compared to when images are processed only by an AI method using a neural network. [Means for solving the problem]
[0009] A first aspect of the technology of the present disclosure is an information processing device that includes a processor and a memory connected to or built into the processor, in which the processor processes a captured image using an AI method that uses a neural network, and performs a synthesis process that synthesizes a first image obtained by processing the captured image using the AI method and a second image obtained by not processing the captured image using the AI method.
[0010] A second aspect of the technology of the present disclosure is an information processing device according to the first aspect, in which a processor performs an AI noise adjustment process to adjust noise contained in a captured image using an AI method, and adjusts the noise by performing a synthesis process.
[0011] A third aspect of the technology of the present disclosure is an information processing device according to the second aspect, in which the processor performs a non-AI noise adjustment process that adjusts noise using a non-AI method that does not use a neural network, and the second image is an image obtained by adjusting noise in the captured image using the non-AI noise adjustment process.
[0012] A fourth aspect of the technique of the present disclosure is the information processing device according to the second or third aspect, in which the second image is an image obtained without adjusting noise in the captured image.
[0013] A fifth aspect of the technology of the present disclosure is an information processing device according to any one of the second to fourth aspects, in which a processor assigns weights to the first image and the second image and synthesizes the first image and the second image according to the weights.
[0014] A sixth aspect of the technology of the present disclosure is an information processing device according to the fifth aspect, in which weights are classified into a first weight assigned to a first image and a second weight assigned to a second image, and a processor combines the first image and the second image by performing a weighted average using the first weight and the second weight.
[0015] A seventh aspect of the technique of the present disclosure is the information processing device according to the fifth or sixth aspect, in which the processor changes the weight according to related information related to the captured image.
[0016] An eighth aspect of the technique of the present disclosure is the information processing device according to the seventh aspect, in which the related information includes sensitivity related information related to the sensitivity of an image sensor used in capturing an image to obtain the captured image.
[0017] A ninth aspect of the technique of the present disclosure is the information processing device according to the seventh or eighth aspect, in which the related information includes brightness-related information related to the brightness of the captured image.
[0018] A tenth aspect of the technique of the present disclosure is the information processing device according to the ninth aspect, in which the brightness-related information is pixel statistics of at least a part of the captured image.
[0019] An eleventh aspect of the technique of the present disclosure is the information processing device according to any one of the seventh to tenth aspects, in which the related information includes spatial frequency information indicating the spatial frequency of the captured image.
[0020] A twelfth aspect of the technology of the present disclosure is an information processing device according to any one of the fifth to eleventh aspects, in which a processor detects a subject that appears in a captured image based on the captured image and changes a weight according to the detected subject.
[0021] A thirteenth aspect of the technology of the present disclosure is an information processing device according to any one of the fifth to twelfth aspects, in which a processor detects a part of a subject that appears in a captured image based on the captured image and changes a weight according to the detected part.
[0022] A fourteenth aspect of the technology of the present disclosure is an information processing device according to any one of the fifth to thirteenth aspects, in which a neural network is provided for each imaging scene, and a processor switches the neural network for each imaging scene and changes the weights according to the neural network.
[0023] A fifteenth aspect of the technology of the present disclosure is an information processing device according to any one of the fifth to fourteenth aspects, in which a processor changes a weight depending on the degree of difference between the feature values of the first image and the feature values of the second image.
[0024] A 16th aspect of the technology of the present disclosure is an information processing device according to any one of the 2nd to 15th aspects, in which the processor normalizes the image input to the neural network with respect to image characteristic parameters determined according to the image sensor used in the imaging to obtain the image input to the neural network and the imaging conditions.
[0025] A 17th aspect of the technology of the present disclosure is an information processing device according to any one of the 2nd to 16th aspects, in which the training image input to the neural network when training the neural network is an image in which the first RAW image obtained by capturing an image with the first imaging device is normalized with respect to at least one first parameter, namely the number of bits and the offset value of the first RAW image.
[0026] In an eighteenth aspect of the technique of the present disclosure, the captured image is an image for inference, and the first parameter is associated with a neural network to which a training image is input, and when a second RAW image obtained by capturing an image with a second imaging device is input as an image for inference to the neural network that has been trained by inputting the training image, the processor normalizes the second RAW image using a first parameter associated with the neural network to which the training image is input and at least one second parameter selected from the number of bits and offset value of the second RAW image.
[0027] A 19th aspect of the technology of the present disclosure is an information processing device according to the 18th aspect, in which the first image is a normalized noise-adjusted image obtained by adjusting noise in a second RAW image normalized using a first parameter and a second parameter using an AI noise adjustment process using a neural network that has been trained by inputting a learning image, and the processor adjusts the normalized noise-adjusted image to an image of the second parameter using the first parameter and the second parameter.
[0028] A twentieth aspect of the technology of the present disclosure is an information processing device according to any one of the second to nineteenth aspects, in which the processor performs signal processing on the first image and the second image according to specified setting values, and the setting values differ when performing signal processing on the first image and when performing signal processing on the second image.
[0029] A 21st aspect of the technology of the present disclosure is an information processing device according to any one of the 2nd to 20th aspects, in which the processor performs processing on the first image to compensate for sharpness lost by AI noise adjustment processing.
[0030] A 22nd aspect of the technology of the present disclosure is an information processing device according to any one of the 2nd to 21st aspects, in which the first image to be synthesized in the synthesis process is an image represented by a color difference signal obtained by performing AI noise adjustment processing on a captured image.
[0031] A 23rd aspect of the technology of the present disclosure is an information processing device according to any one of the 2nd to 22nd aspects, in which the second image to be synthesized in the synthesis process is an image represented by a luminance signal obtained without performing AI noise adjustment processing on the captured image.
[0032] A 24th aspect of the technology of the present disclosure is an information processing device according to any one of the 2nd to 23rd aspects, in which the first image to be combined in the combination process is an image represented by a color difference signal obtained by performing AI noise adjustment processing on the captured image, and the second image is an image represented by a luminance signal obtained without performing AI noise adjustment processing on the captured image.
[0033] A 25th aspect of the technology of the present disclosure is an imaging device that includes a processor, a memory connected to or built into the processor, and an image sensor, in which the processor processes an image obtained by capturing an image with the image sensor using an AI method that uses a neural network, and performs a synthesis process to synthesize a first image obtained by processing the captured image using the AI method and a second image obtained by not processing the captured image using the AI method.
[0034] A 26th aspect of the technology of the present disclosure is an information processing method that includes processing a captured image obtained by capturing an image with an image sensor using an AI method with a neural network, and performing a synthesis process that synthesizes a first image obtained by processing the captured image using the AI method with a second image obtained without processing the captured image using the AI method.
[0035] A 27th aspect of the technology of the present disclosure is a program for causing a computer to execute processing including processing an image obtained by capturing an image using an image sensor using an AI method that employs a neural network, and performing a synthesis process that synthesizes a first image obtained by processing the captured image using the AI method with a second image obtained by not processing the captured image using the AI method. [Brief explanation of the drawings]
[0036] [Figure 1] 1 is a schematic diagram illustrating an example of the overall configuration of an imaging device. [Figure 2] FIG. 1 is a schematic diagram illustrating an example of a hardware configuration of an optical system and an electrical system of an imaging apparatus. [Figure 3] FIG. 2 is a block diagram illustrating an example of the functions of an image processing engine. [Figure 4] FIG. 1 is a conceptual diagram illustrating an example of the configuration of a learning execution system. [Figure 5] FIG. 1 is a conceptual diagram showing an example of processing contents of an AI processing unit and a non-AI processing unit. [Figure 6] FIG. 10 is a block diagram showing an example of processing content of a weight derivation unit. [Figure 7] FIG. 10 is a conceptual diagram illustrating an example of processing contents of a weighting unit and a combining unit. [Figure 8] FIG. 2 is a conceptual diagram illustrating an example of a function of a signal processing unit. [Figure 9] 10 is a flowchart showing an example of the flow of image quality adjustment processing. [Figure 10] FIG. 10 is a conceptual diagram showing an example of processing contents of a weight derivation unit according to a first modified example. [Figure 11] FIG. 10 is a conceptual diagram showing an example of processing details of a weighting unit and a combining unit according to a first modified example. [Figure 12] FIG. 11 is a conceptual diagram showing an example of processing contents of a weight derivation unit according to a second modified example. [Figure 13] FIG. 10 is a conceptual diagram showing an example of processing details of a weight derivation unit according to the third and fourth modified examples. [Figure 14] FIG. 13 is a conceptual diagram showing an example of processing contents of a weight derivation unit according to a fifth modified example. [Figure 15] FIG. 13 is a block diagram showing an example of the contents stored in the NVM according to a sixth modified example. [Figure 16] FIG. 13 is a conceptual diagram showing an example of processing contents of an AI method processing unit according to a sixth modified example. [Figure 17] FIG. 13 is a block diagram showing an example of processing contents of a weight derivation unit according to a sixth modified example. [Figure 18] FIG. 13 is a conceptual diagram showing an example of the configuration of a learning execution system according to a seventh modified example. [Figure 19] FIG. 13 is a conceptual diagram showing an example of processing contents of an image processing engine according to a seventh modified example. [Figure 20]FIG. 13 is a block diagram showing an example of functions of a signal processing unit and a parameter adjustment unit according to an eighth modified example. [Figure 21] FIG. 13 is a conceptual diagram showing an example of processing contents of an AI processing unit, a non-AI processing unit, and a signal processing unit according to a ninth modified example. [Figure 22] FIG. 13 is a conceptual diagram showing an example of processing content of a first image processing unit according to a ninth modified example. [Figure 23] FIG. 13 is a conceptual diagram showing an example of processing contents of a second image processing unit according to a ninth modified example. [Figure 24] FIG. 13 is a conceptual diagram showing an example of processing contents of a synthesis unit according to a ninth modified example. [Figure 25] FIG. 10 is a conceptual diagram showing a modified example of image quality adjustment processing. [Figure 26] FIG. 1 is a schematic configuration diagram illustrating an example of an imaging system. DETAILED DESCRIPTION OF THE INVENTION
[0037] Hereinafter, exemplary embodiments of an image processing device, an imaging device, an image processing method, and a program according to the techniques of the present disclosure will be described with reference to the accompanying drawings.
[0038] First, the terms used in the following description will be explained.
[0039] CPU is an abbreviation for "Central Processing Unit". GPU is an abbreviation for "Graphics Processing Unit". TPU is an abbreviation for "Tensor processing unit". NVM is an abbreviation for "Non-volatile memory". RAM is an abbreviation for "Random Access Memory". IC is an abbreviation for "Integrated Circuit". ASIC is an abbreviation for "Application Specific Integrated Circuit". PLD is an abbreviation for "Programmable Logic Device". FPGA is an abbreviation for "Field-Programmable Gate Array". SoC is an abbreviation for "System-on-a-chip". SSD is an abbreviation for "Solid State Drive". USB is an abbreviation for "Universal Serial Bus". HDD is an abbreviation for "Hard Disk Drive". EEPROM is an abbreviation for "Electrically Erasable and Programmable Read Only Memory". EL is an abbreviation for "Electro-Luminescence". I / F is an abbreviation for "Interface". UI is an abbreviation for "User Interface". fps is an abbreviation for "frames per second". MF is an abbreviation for "Manual Focus". AF is an abbreviation for "Auto Focus". CMOS is an abbreviation for "Complementary Metal Oxide Semiconductor". CCD is an abbreviation for "Charge Coupled Device". LAN is an abbreviation for "Local Area Network". WAN is an abbreviation for "Wide Area Network". NN is an abbreviation for "Neural Network". CNN is an abbreviation for "Convolutional Neural Network". AI is an abbreviation for "Artificial Intelligence".A / D is an abbreviation for "Analog / Digital". FIR is an abbreviation for "Finite Impulse Response". IIR is an abbreviation for "Infinite Impulse Response". JPEG is an abbreviation for "Joint Photographic Experts Group". TIFF is an abbreviation for "Tagged Image File Format". JPEG XR is an abbreviation for "Joint Photographic Experts Group Extended Range". ID is an abbreviation for "Identification". LSB is an abbreviation for "Least Significant Bit".
[0040] As an example, as shown in FIG. 1 , an imaging device 10 is a device that captures an image of a subject, and includes an image processing engine 12, an imaging device body 16, and an interchangeable lens 18. The image processing engine 12 is an example of an "information processing device" and a "computer" according to the techniques of the present disclosure. The image processing engine 12 is built into the imaging device body 16 and controls the entire imaging device 10. The interchangeable lens 18 is interchangeably attached to the imaging device body 16. The interchangeable lens 18 is provided with a focus ring 18A. The focus ring 18A is operated by a user of the imaging device 10 (hereinafter simply referred to as "user") when the user manually adjusts the focus of the imaging device 10 on a subject.
[0041] 1, a digital camera with an interchangeable lens is shown as an example of the imaging device 10. However, this is merely one example, and the imaging device 10 may be a digital camera with a fixed lens, or may be a digital camera built into various electronic devices such as a smart device, a wearable terminal, a cell observation device, an ophthalmic observation device, or a surgical microscope.
[0042] The imaging device body 16 is provided with an image sensor 20. The image sensor 20 is an example of an "image sensor" according to the technology of the present disclosure. The image sensor 20 is a CMOS image sensor. The image sensor 20 captures an image of an imaging range including at least one subject. When an interchangeable lens 18 is attached to the imaging device body 16, subject light representing the subject passes through the interchangeable lens 18 and is focused on the image sensor 20, and image data representing the image of the subject is generated by the image sensor 20.
[0043] In this embodiment, a CMOS image sensor is exemplified as the image sensor 20, but the technology of the present disclosure is not limited to this, and the technology of the present disclosure is also applicable even if the image sensor 20 is another type of image sensor, such as a CCD image sensor.
[0044] A release button 22 and a dial 24 are provided on the top surface of the imaging device body 16. The dial 24 is operated when setting the operation mode of the imaging system and the operation mode of the playback system, and by operating the dial 24, the imaging device 10 is selectively set to one of the imaging mode, playback mode, and setting mode as its operation mode. The imaging mode is an operation mode that causes the imaging device 10 to capture images. The playback mode is an operation mode that plays back images (e.g., still images and / or moving images) obtained by capturing images for recording in the imaging mode. The setting mode is an operation mode that is set for the imaging device 10 when, for example, setting various setting values used in control related to imaging.
[0045] The release button 22 functions as an imaging preparation instruction unit and an imaging instruction unit, and is capable of detecting two stages of pressing operation: an imaging preparation instruction state and an imaging instruction state. The imaging preparation instruction state refers to a state in which the button is pressed from a standby position to an intermediate position (half-pressed position), for example, and the imaging instruction state refers to a state in which the button is pressed beyond the intermediate position to a final pressed position (fully-pressed position). Note that, hereinafter, the "state in which the button is pressed from the standby position to the half-pressed position" will be referred to as the "half-pressed state," and the "state in which the button is pressed from the standby position to the fully-pressed position" will be referred to as the "fully-pressed state." Depending on the configuration of the imaging device 10, the imaging preparation instruction state may be a state in which the user's finger is in contact with the release button 22, and the imaging instruction state may be a state in which the operating user's finger has moved from a state in which the button is in contact with the release button 22 to a state in which the finger is released.
[0046] On the rear surface of the imaging device main body 16, instruction keys 26 and a touch panel display 32 are provided.
[0047] The touch panel display 32 includes a display 28 and a touch panel 30 (see also FIG. 2). An example of the display 28 is an EL display (e.g., an organic EL display or an inorganic EL display). The display 28 may be a different type of display, such as a liquid crystal display, instead of an EL display.
[0048] The display 28 displays images and / or text information, etc. When the imaging device 10 is in imaging mode, the display 28 is used to capture images for live view images, i.e., to display live view images obtained by continuous imaging. Here, a "live view image" refers to a moving image for display based on image data obtained by imaging by the image sensor 20. The imaging performed to obtain a live view image (hereinafter also referred to as "image capture for live view images") is performed at a frame rate of, for example, 60 fps. 60 fps is merely an example, and the frame rate may be less than 60 fps or may be greater than 60 fps.
[0049] The display 28 is also used to display a still image obtained by capturing a still image when an instruction to capture a still image is given to the imaging device 10 via the release button 22. The display 28 is also used to display a playback image when the imaging device 10 is in playback mode. Furthermore, the display 28 is also used to display a menu screen on which various menus can be selected when the imaging device 10 is in setting mode, and a setting screen for setting various setting values used in imaging-related controls.
[0050] The touch panel 30 is a transmissive touch panel that is overlaid on the surface of the display area of the display 28. The touch panel 30 receives instructions from the user by detecting contact with a pointing object such as a finger or a stylus pen. For ease of explanation, the "full press state" described above will hereinafter also include a state in which the user presses the soft key for starting imaging via the touch panel 30.
[0051] In this embodiment, an out-cell type touch panel display in which the touch panel 30 is overlaid on the surface of the display area of the display 28 is given as an example of the touch panel display 32, but this is merely an example. For example, an on-cell type or an in-cell type touch panel display can also be used as the touch panel display 32.
[0052] The instruction keys 26 accept various instructions. Here, "various instructions" refers to, for example, an instruction to display a menu screen, an instruction to select one or more menus, an instruction to confirm a selection, an instruction to erase a selection, an instruction to zoom in, zoom out, and frame-by-frame advance. These instructions may also be given via the touch panel 30.
[0053] As an example, as shown in FIG. 2, the image sensor 20 includes a photoelectric conversion element 72. The photoelectric conversion element 72 has a light-receiving surface 72A. The photoelectric conversion element 72 is disposed within the imaging device body 16 so that the center of the light-receiving surface 72A coincides with the optical axis OA (see also FIG. 1). The photoelectric conversion element 72 has a plurality of photosensitive pixels arranged in a matrix, and the light-receiving surface 72A is formed by the plurality of photosensitive pixels. Each photosensitive pixel has a microlens (not shown). Each photosensitive pixel is a physical pixel having a photodiode (not shown), which photoelectrically converts received light and outputs an electrical signal according to the amount of received light.
[0054] In addition, the multiple photosensitive pixels have red (R), green (G), or blue (B) color filters (not shown) arranged in a matrix in a predetermined pattern arrangement (e.g., Bayer arrangement, G-stripe R / G complete checkerboard, X-Trans (registered trademark) arrangement, honeycomb arrangement, etc.).
[0055] For ease of explanation, hereinafter, a photosensitive pixel having a microlens and an R color filter will be referred to as an R pixel, a photosensitive pixel having a microlens and a G color filter will be referred to as a G pixel, and a photosensitive pixel having a microlens and a B color filter will be referred to as a B pixel. Also, for ease of explanation, hereinafter, an electrical signal output from an R pixel will be referred to as an "R signal," an electrical signal output from a G pixel will be referred to as a "G signal," and an electrical signal output from a B pixel will be referred to as a "B signal." Also, for ease of explanation, hereinafter, the R signal, G signal, and B signal will also be referred to as "RGB color signals."
[0056] The interchangeable lens 18 includes an imaging lens 40. The imaging lens 40 has an objective lens 40A, a focus lens 40B, a zoom lens 40C, and an aperture 40D. The objective lens 40A, the focus lens 40B, the zoom lens 40C, and the aperture 40D are arranged in this order along the optical axis OA from the subject side (object side) to the imaging device main body 16 side (image side).
[0057] The interchangeable lens 18 also includes a control device 36, a first actuator 37, a second actuator 38, and a third actuator 39. The control device 36 controls the entire interchangeable lens 18 in accordance with instructions from the imaging device main body 16. The control device 36 is a device having a computer including, for example, a CPU, an NVM, and RAM. The NVM of the control device 36 is, for example, an EEPROM. However, this is merely one example, and instead of or together with the EEPROM, an HDD and / or an SSD may be used as the NVM of the system controller 44. The RAM of the control device 36 temporarily stores various types of information and is used as a work memory. In the control device 36, the CPU reads necessary programs from the NVM and executes the read programs on the RAM to control the entire imaging lens 40.
[0058] Although a device having a computer is given here as an example of the control device 36, this is merely an example, and devices including ASIC, FPGA, and / or PLD may also be applied. Furthermore, the control device 36 may be, for example, a device realized by a combination of hardware and software configurations.
[0059] The first actuator 37 includes a focusing slide mechanism (not shown) and a focusing motor (not shown). The focusing slide mechanism has a focus lens 40B attached thereto so as to be slidable along the optical axis OA. The focusing motor is also connected to the focusing slide mechanism, and the focusing slide mechanism operates by receiving power from the focusing motor to move the focus lens 40B along the optical axis OA.
[0060] The second actuator 38 includes a zoom slide mechanism (not shown) and a zoom motor (not shown). The zoom lens 40C is attached to the zoom slide mechanism so that it can slide along the optical axis OA. The zoom motor is also connected to the zoom slide mechanism, and the zoom slide mechanism operates by receiving power from the zoom motor to move the zoom lens 40C along the optical axis OA.
[0061] The third actuator 39 includes a power transmission mechanism (not shown) and an aperture motor (not shown). The aperture 40D has an opening 40D1, and the size of the opening 40D1 is variable. The opening 40D1 is formed, for example, by a plurality of aperture blades 40D2. The plurality of aperture blades 40D2 are connected to the power transmission mechanism. The power transmission mechanism is also connected to an aperture motor, and the power transmission mechanism transmits the power of the aperture motor to the plurality of aperture blades 40D2. The plurality of aperture blades 40D2 operate upon receiving power transmitted from the power transmission mechanism, thereby changing the size of the aperture 40D1. The aperture 40D adjusts exposure by changing the size of the opening 40D1.
[0062] The focus motor, zoom motor, and aperture motor are connected to a control device 36, which controls the driving of each of the focus motor, zoom motor, and aperture motor. In this embodiment, a stepping motor is used as an example of the focus motor, zoom motor, and aperture motor. Therefore, the focus motor, zoom motor, and aperture motor operate in synchronization with pulse signals in response to commands from the control device 36. While an example is shown here in which the focus motor, zoom motor, and aperture motor are provided in the interchangeable lens 18, this is merely an example, and at least one of the focus motor, zoom motor, and aperture motor may be provided in the imaging device body 16. The components and / or operation method of the interchangeable lens 18 can be changed as needed.
[0063] In the imaging mode, the imaging device 10 selectively sets MF mode and AF mode in accordance with instructions given to the imaging device body 16. The MF mode is an operating mode in which the focus is adjusted manually. In the MF mode, for example, when the user operates the focus ring 18A or the like, the focus lens 40B moves along the optical axis OA by an amount corresponding to the amount of operation of the focus ring 18A or the like, thereby adjusting the focus.
[0064] In AF mode, imaging device body 16 calculates the in-focus position according to the subject distance, and adjusts the focus by moving focus lens 40B toward the calculated in-focus position. Here, the in-focus position refers to the position on optical axis OA of focus lens 40B when the subject is in focus.
[0065] The imaging device main body 16 includes an image sensor 20, an image processing engine 12, a system controller 44, an image memory 46, a UI device 48, an external I / F 50, a communication I / F 52, a photoelectric conversion element driver 54, and an input / output interface 70. The image sensor 20 also includes a photoelectric conversion element 72 and an A / D converter 74.
[0066] The input / output interface 70 is connected to the image processing engine 12, the image memory 46, the UI device 48, the external I / F 50, the photoelectric conversion element driver 54, and the A / D converter 74. The input / output interface 70 is also connected to the control device 36 of the interchangeable lens 18.
[0067] The system controller 44 includes a CPU (not shown), an NVM (not shown), and a RAM (not shown). In the system controller 44, the NVM is a non-transitory storage medium that stores various parameters and programs. The NVM of the system controller 44 is, for example, an EEPROM. However, this is merely an example, and instead of or in addition to the EEPROM, an HDD and / or an SSD may be used as the NVM of the system controller 44. The RAM of the system controller 44 temporarily stores various information and is used as a work memory. In the system controller 44, the CPU reads necessary programs from the NVM and executes the read programs on the RAM to control the entire imaging device 10. That is, in the example shown in FIG. 2, the image processing engine 12, the image memory 46, the UI device 48, the external I / F 50, the communication I / F 52, the photoelectric conversion element driver 54, and the control device 36 are controlled by the system controller 44.
[0068] The image processing engine 12 operates under the control of the system controller 44. The image processing engine 12 includes a CPU 62, an NVM 64, and a RAM 66. Here, the CPU 62 is an example of a "processor" according to the technology of the present disclosure, and the NVM 64 is an example of a "memory" according to the technology of the present disclosure.
[0069] The CPU 62, NVM 64, and RAM 66 are connected via a bus 68, which is connected to an input / output interface 70. Although the example shown in Fig. 2 shows a single bus as the bus 68 for convenience of illustration, multiple buses may be used. The bus 68 may be a serial bus or a parallel bus including a data bus, an address bus, a control bus, etc.
[0070] The NVM 64 is a non-transitory storage medium that stores various parameters and programs different from those stored in the NVM of the system controller 44. The various programs include an image quality adjustment processing program 80 (see FIG. 3), which will be described later. The NVM 64 is, for example, an EEPROM. However, this is merely an example, and instead of or together with the EEPROM, an HDD and / or an SSD may be used as the NVM 64. The RAM 66 temporarily stores various information and is used as a work memory.
[0071] The CPU 62 reads out a necessary program from the NVM 64 and executes the read program in the RAM 66. The CPU 62 performs image processing in accordance with the program executed on the RAM 66.
[0072] The photoelectric conversion element 72 is connected to a photoelectric conversion element driver 54. The photoelectric conversion element driver 54 supplies an imaging timing signal that defines the timing of imaging performed by the photoelectric conversion element 72 to the photoelectric conversion element 72 in accordance with an instruction from the CPU 62. The photoelectric conversion element 72 performs resetting, exposure, and output of an electrical signal in accordance with the imaging timing signal supplied from the photoelectric conversion element driver 54. Examples of imaging timing signals include a vertical synchronization signal and a horizontal synchronization signal.
[0073] When the interchangeable lens 18 is attached to the imaging device body 16, subject light incident on the imaging lens 40 is imaged on the light-receiving surface 72A by the imaging lens 40. Under the control of the photoelectric conversion element driver 54, the photoelectric conversion element 72 photoelectrically converts the subject light received by the light-receiving surface 72A and outputs an electrical signal corresponding to the amount of subject light to the A / D converter 74 as analog image data indicating the subject light. Specifically, the A / D converter 74 reads out the analog image data from the photoelectric conversion element 72 in units of one frame and for each horizontal line using an exposure sequential readout method.
[0074] The A / D converter 74 digitizes the analog image data to generate a RAW image 75A. The RAW image 75A is an example of a "captured image" according to the technology of the present disclosure. The RAW image 75A is an image in which R pixels, G pixels, and B pixels are arranged in a mosaic pattern. In the present embodiment, as an example, the number of bits, i.e., the bit length, of each of the R pixels, B pixels, and G pixels included in the RAW image 75A is 14 bits.
[0075] In this embodiment, as an example, the CPU 62 of the image processing engine 12 acquires a RAW image 75A from the A / D converter 74, and performs image processing on the acquired RAW image 75A.
[0076] A processed image 75B is stored in the image memory 46. The processed image 75B is an image obtained by the CPU 62 performing image processing on the RAW image 75A.
[0077] The UI device 48 includes a display 28, and the CPU 62 displays various pieces of information on the display 28. The UI device 48 also includes a reception device 76. The reception device 76 includes a touch panel 30 and a hard key unit 78. The hard key unit 78 is a plurality of hard keys including the instruction keys 26 (see FIG. 1 ). The CPU 62 operates in accordance with various instructions received by the touch panel 30. Note that, although the hard key unit 78 is included in the UI device 48 here, the technology of the present disclosure is not limited to this, and for example, the hard key unit 78 may be connected to the external I / F 50.
[0078] The external I / F 50 controls the exchange of various information with devices (hereinafter also referred to as "external devices") that exist outside the imaging device 10. An example of the external I / F 50 is a USB interface. To the USB interface, external devices (not shown) such as smart devices, personal computers, servers, USB memory, memory cards, and / or printers are directly or indirectly connected.
[0079] The communication I / F 52 is connected to a network (not shown). The communication I / F 52 controls the exchange of information between a communication device (not shown), such as a server on the network, and the system controller 44. For example, the communication I / F 52 transmits information in response to a request from the system controller 44 to the communication device via the network. The communication I / F 52 also receives information transmitted from the communication device and outputs the received information to the system controller 44 via the input / output interface 70.
[0080] As an example, as shown in FIG. 3, an image quality adjustment processing program 80 is stored in the NVM 64 of the image capture device 10. The image quality adjustment processing program 80 is an example of a "program" according to the technology of the present disclosure. A trained neural network 82 is also stored in the NVM 64 of the image capture device 10. Note that, for ease of explanation, "neural network" will also be abbreviated as "NN" below.
[0081] The CPU 62 reads out the image quality adjustment processing program 80 from the NVM 64, and executes the read-out image quality adjustment processing program 80 on the RAM 66. The CPU 62 performs image quality adjustment processing (see FIG. 9 ) in accordance with the image quality adjustment processing program 80 executed on the RAM 66. The image quality adjustment processing is realized by the CPU 62 operating as an AI processing unit 62A, a non-AI processing unit 62B, a weight derivation unit 62C, a weight assignment unit 62D, a synthesis unit 62E, and a signal processing unit 62F in accordance with the image quality adjustment processing program 80.
[0082] As an example, as shown in FIG. 4, a trained NN 82 is generated by a learning execution system 84. The learning execution system 84 includes a storage device 86 and a learning execution device 88. An example of the storage device 86 is a HDD. Note that the HDD is merely an example, and other types of storage devices such as an SSD may also be used. The learning execution device 88 is a device realized by a computer or the like having a CPU (not shown), NVM (not shown), and RAM (not shown).
[0083] The trained NN 82 is generated by executing machine learning on the NN 90 by the learning execution device 88. The trained NN 82 is a trained model generated by optimizing the NN 90 through machine learning. An example of the NN 90 is a CNN.
[0084] The storage device 86 stores a plurality of (e.g., tens of thousands to hundreds of billions) pieces of teacher data 92. The learning execution device 88 is connected to the storage device 86. The learning execution device 88 acquires the plurality of pieces of teacher data 92 from the storage device 86, and causes the NN 90 to perform machine learning using the acquired plurality of pieces of teacher data 92.
[0085] The training data 92 is labeled data. The labeled data is, for example, data in which a training RAW image 75A1 and supervised answer data 75C are associated with each other. The training RAW image 75A1 is, for example, a RAW image 75A obtained by capturing an image using the imaging device 10, and / or a RAW image obtained by capturing an image using an imaging device other than the imaging device 10.
[0086] The supervised data 75C is an image obtained by removing noise from the learning RAW image 75A1. Here, noise refers to, for example, noise caused by image capture by the imaging device 10. Examples of noise include pixel defects, dark current noise, and / or beat noise.
[0087] The learning execution device 88 acquires teacher data 92 one by one from the storage device 86. The learning execution device 88 inputs the learning RAW image 75A1 from the teacher data 92 acquired from the storage device 86 to the NN 90. When the learning RAW image 75A1 is input, the NN 90 performs inference and outputs an image 94 indicating the inference result.
[0088] The learning execution device 88 calculates an error 96 between the image 94 and the supervised answer data 75C associated with the learning RAW image 75A1 input to the NN 90. The learning execution device 88 calculates a plurality of adjustment values 98 that minimize the error 96. The learning execution device 88 then adjusts a plurality of optimization variables in the NN 90 using the plurality of adjustment values 98. Here, the plurality of optimization variables refers to, for example, a plurality of connection weights and a plurality of offset values included in the NN 90.
[0089] The learning execution device 88 repeatedly performs the learning process of inputting the learning RAW images 75A1 to the NN 90, calculating the error 96, calculating the plurality of adjustment values 98, and adjusting the plurality of optimization variables in the NN 90, using the plurality of teacher data 92 stored in the storage device 86. In other words, the learning execution device 88 optimizes the NN 90 by adjusting the plurality of optimization variables in the NN 90 using the plurality of adjustment values 98 calculated so as to minimize the error 96 for each of the plurality of learning RAW images 75A1 included in the plurality of teacher data 92 stored in the storage device 86.
[0090] The learning execution device 88 generates a trained NN 82 by optimizing the NN 90. The learning execution device 88 is connected to the external I / F 50 or the communication I / F 52 (see FIG. 2) of the imaging device main body 16, and stores the trained NN 82 in the NVM 64 (see FIG. 3).
[0091] For example, when a RAW image 75A (see FIG. 2) is input to the trained NN 82, the trained NN 82 outputs an image from which most of the noise has been removed. Due to the characteristics of the trained NN 82, when noise contained in the RAW image 75A is removed, it is conceivable that the fine structure of the subject captured in the RAW image 75A (e.g., the subject's fine contours and / or fine patterns) may also be removed. If the fine structure of the subject is removed, the RAW image 75A may become an image lacking in sharpness. The reason why such an image is obtained from the trained NN 82 is thought to be that the trained NN 82 is not good at distinguishing between noise and the subject's fine structure. In particular, if the number of layers included in the NN 90 is reduced and the trained NN 82 is simplified, it is expected that it will become more difficult for the trained NN 82 to distinguish between noise and the subject's fine structure (hereinafter also referred to as "fine structure").
[0092] In consideration of these circumstances, the imaging device 10 is configured so that the CPU 62 performs image quality adjustment processing (see FIGS. 3 and 6 to 9). By performing the image quality adjustment processing, the CPU 62 processes the RAW image for inference 75A2 (see FIG. 5) by an AI method using the trained NN 82, and performs a synthesis process to synthesize a first image 75D (see FIGS. 5 and 7) obtained by processing the RAW image for inference 75A2 by the AI method with a second image 75E (see FIGS. 5 and 7) obtained by not processing the RAW image for inference 75A2 by the AI method. The RAW image for inference 75A2 is an image inferred by the trained NN 82. In this embodiment, the RAW image 75A obtained by capturing an image by the imaging device 10 is used as the RAW image for inference 75A2. Note that the RAW image 75A is merely an example, and the RAW image for inference 75A2 may be an image other than the RAW image 75A (for example, an image obtained by processing the RAW image 75A).
[0093] 5, an inference-use RAW image 75A2 is input to the AI processing unit 62A. The AI processing unit 62A performs AI noise adjustment processing on the inference-use RAW image 75A2. The AI noise adjustment processing is processing that adjusts noise included in the inference-use RAW image 75A2 by AI. The AI processing unit 62A performs processing using a trained NN 82 as the AI noise adjustment processing.
[0094] In this case, the AI processing unit 62A inputs the RAW image for inference 75A2 to the trained NN 82. When the RAW image for inference 75A2 is input, the trained NN 82 performs inference on the RAW image for inference 75A2 and outputs a first image 75D as the inference result. The first image 75D is an image with reduced noise compared to the RAW image for inference 75A2. The first image 75D is an example of a "first image" according to the technology of the present disclosure.
[0095] Similar to the AI processing unit 62A, the non-AI processing unit 62B also receives the RAW image for inference 75A2. The non-AI processing unit 62B performs non-AI noise adjustment processing on the RAW image for inference 75A2. The non-AI noise adjustment processing is processing that adjusts noise included in the RAW image for inference 75A2 by a non-AI method that does not use an NN.
[0096] The non-AI processing unit 62B has a digital filter 100. The non-AI processing unit 62B performs processing using the digital filter 100 as non-AI noise adjustment processing. The digital filter 100 is, for example, an FIR filter. Note that the FIR filter is merely one example, and other digital filters such as an IIR filter may also be used, as long as they have the function of reducing noise contained in the inference-use RAW image 75A2 in a non-AI manner.
[0097] The non-AI processing unit 62B generates a second image 75E by filtering the RAW image for inference 75A2 using the digital filter 100. The second image 75E is an image obtained by filtering using the digital filter 100, i.e., an image obtained by adjusting noise using non-AI noise adjustment processing. The second image 75E is an image in which noise has been reduced more than the RAW image for inference 75A2, but it also has residual noise compared to the first image 75D. The second image 75E is an example of a "second image" according to the technology of the present disclosure.
[0098] The second image 75E contains noise that was removed from the RAW image for inference 75A2 by the trained NN 82, while also containing the fine structure that was removed from the RAW image for inference 75A2 by the trained NN 82. Therefore, by combining the first image 75D and the second image 75E, the CPU 62 not only reduces noise but also generates an image that avoids the loss of the fine structure (for example, an image that maintains sharpness).
[0099] Incidentally, one of the causes of noise intrusion into the inference-use RAW image 75A2 is the sensitivity (e.g., ISO sensitivity) of the image sensor 20. This is because the sensitivity of the image sensor 20 depends on the analog gain used to amplify analog image data, and increasing the analog gain also increases noise. In addition, in this embodiment, the trained NN 82 and the digital filter 100 differ in their ability to remove noise caused by the sensitivity of the image sensor 20.
[0100] Therefore, CPU 62 assigns different weights to first image 75D and second image 75E to be combined, and combines first image 75D and second image 75E according to the assigned weights. The weights assigned to first image 75D and second image 75E refer to the degree of pixel value of first image 75D and the degree of pixel value of second image 75E used to combine pixels at corresponding pixel positions between first image 75D and second image 75E.
[0101] For example, if the digital filter 100 has a lower ability to remove noise caused by the sensitivity of the image sensor 20 than the trained NN 82, a smaller weight is assigned to the first image 75D than to the second image 75E. The difference in weight assigned to the first image 75D and the second image 75E is determined according to the difference in ability to remove noise caused by the sensitivity of the image sensor 20, etc.
[0102] As an example, as shown in Fig. 6, the NVM 64 stores related information 102. The related information 102 is information related to the RAW image for inference 75A2. The related information 102 includes sensitivity related information 102A. The sensitivity related information 102A is information related to the sensitivity of the image sensor 20 used in capturing the image to obtain the RAW image for inference 75A2. An example of the sensitivity related information 102A is information indicating ISO sensitivity.
[0103] The weight derivation unit 62C acquires related information 102 from the NVM 64. Based on the related information 102 acquired from the NVM 64, the weight derivation unit 62C derives a first weight 104 and a second weight 106 as weights to be assigned to the first image 75D and the second image 75E. The weights assigned to the first image 75D and the second image 75E are categorized into the first weight 104 and the second weight 106. The first weight 104 is the weight assigned to the first image 75D, and the second weight 106 is the weight assigned to the second image 75E.
[0104] The weight derivation unit 62C has a weight calculation formula 108. The weight calculation formula 108 is a calculation formula in which a parameter identified from the related information 102 is an independent variable and the first weight 104 is a dependent variable. Here, the parameter identified from the related information 102 is, for example, a value indicating the sensitivity of the image sensor 20. The value indicating the sensitivity of the image sensor 20 is identified from the sensitivity related information 102A. Note that, an example of the value indicating the sensitivity of the image sensor 20 is a value indicating the ISO sensitivity. However, this is merely an example, and the value indicating the sensitivity of the image sensor 20 may also be a value indicating an analog gain.
[0105] The weight derivation unit 62C calculates the first weight 104 by substituting the value indicating the sensitivity of the image sensor 20 into the weight calculation formula 108. Here, if the first weight 104 is "w", the first weight 104 is a value that satisfies the magnitude relationship of "0 < w < 1". The second weight is "1 - w". The weight derivation unit 62C calculates the second weight 106 from the first weight 104 calculated using the weight calculation formula 108.
[0106] Thus, since the first weight 104 and the second weight 106 are values dependent on the related information 102, the first weight 104 and the second weight 106 calculated by the weight derivation unit 62C are changed according to the related information 102. For example, the first weight 104 and the second weight 106 are changed by the weight derivation unit 62C according to the value indicating the sensitivity of the image sensor 20.
[0107] As shown in FIG. 7 as an example, the weighting unit 62D acquires the first image 75D from the AI method processing unit 62A and acquires the second image 75E from the non-AI method processing unit 62B. The weighting unit 62D assigns the first weight 104 derived by the weight derivation unit 62C to the first image 75D. The weighting unit 62D assigns the second weight 106 derived by the weight derivation unit 62C to the second image 75E.
[0108] The combining unit 62E adjusts the noise included in the inference RAW image 75A2 by combining the first image 75D and the second image 75E. That is, the image obtained by combining the first image 75D and the second image 75E by the combining unit 62E (in the example shown in FIG. 7, the combined image 75F) is an image in which the noise included in the inference RAW image 75A2 is adjusted.
[0109] The composition unit 62E generates a composite image 75F by combining the first image 75D and the second image 75E in accordance with the first weight 104 and the second weight 106. The composite image 75F is an image obtained by combining the pixel values of each pixel between the first image 75D and the second image 75E in accordance with the first weight 104 and the second weight 106. An example of the composite image 75F is a weighted average image obtained by performing a weighted average using the first weight 104 and the second weight 106. The weighted average using the first weight 104 and the second weight 106 refers to, for example, a weighted average using the first weight 104 and the second weight 106 for the pixel values of each pixel whose pixel positions correspond between the first image 75D and the second image 75E. Note that the weighted average image is merely an example, and if the absolute value of the difference between the first weight 104 and the second weight 106 is less than a threshold value (e.g., 0.01), the image obtained by simply averaging the pixel values without using the first weight 104 and the second weight 106 may also be used as the composite image 75F.
[0110] As an example, as shown in FIG. 8, the signal processing unit 62F includes an offset correction unit 62F1, a white balance correction unit 62F2, a demosaic processing unit 62F3, a color correction unit 62F4, a gamma correction unit 62F5, a color space conversion unit 62F6, a luminance processing unit 62F7, a color difference processing unit 62F8, a color difference processing unit 62F9, a resizing processing unit 62F10, and a compression processing unit 62F11, and performs various signal processing on the composite image 75F.
[0111] The offset correction unit 62F1 performs an offset correction process on the composite image 75F. The offset correction process is a process for correcting dark current components contained in the R pixels, G pixels, and B pixels included in the composite image 75F. One example of the offset correction process is a process for correcting the RGB color signals by subtracting the optical black signal value obtained from the light-shielded photosensitive pixels included in the photoelectric conversion element 72 (see FIG. 2) from the RGB color signals.
[0112] The white balance correction unit 62F2 performs white balance correction on the composite image 75F that has undergone offset correction. The white balance correction corrects the influence of the color of the light source type on the RGB color signals by multiplying the RGB color signals by white balance gains set for each of the R, G, and B pixels. The white balance gain is, for example, a gain for white. An example of the gain for white is a gain set so that the signal levels of the R, G, and B signals are equal for a white subject appearing in the composite image 75F. The white balance gain is set, for example, according to a light source type identified by image analysis or a light source type specified by a user, etc.
[0113] The demosaic processing unit 62F3 performs demosaic processing on the composite image 75F that has undergone white balance correction processing. The demosaic processing is a process of dividing the composite image 75F into three R, G, and B images. That is, the demosaic processing unit 62F3 performs color interpolation processing on the R, G, and B signals to generate R image data representing an image corresponding to R, B image data representing an image corresponding to B, and G image data representing an image corresponding to G. Here, color interpolation processing refers to a process of interpolating a color missing from each pixel from surrounding pixels. That is, since each photosensitive pixel of the photoelectric conversion element 72 can only obtain an R signal, a G signal, or a B signal (i.e., a pixel value of one color out of R, G, and B), the demosaic processing unit 62F3 interpolates the other colors missing from each pixel using pixel values of surrounding pixels. Note that hereinafter, the R image data, B image data, and G image data are also referred to as "RGB image data."
[0114] The color correction unit 62F4 performs color correction (here, as an example, color correction using a linear matrix (i.e., color mixing correction)) on the RGB image data obtained by the demosaic process. The color correction is a process for adjusting the hue and color saturation characteristics. One example of the color correction process is applying a color reproduction coefficient (for example, a linear matrix coefficient) to the RGB image data. The color reproduction coefficient is a coefficient determined to bring the spectral characteristics of R, G, and B closer to the human visual sensitivity characteristics.
[0115] The gamma correction unit 62F5 performs gamma correction on the RGB image data that has been color corrected. The gamma correction process corrects the gradation of the image represented by the RGB image data in accordance with a value that indicates the response characteristics of the gradation of the image, i.e., a gamma value.
[0116] The color space conversion unit 62F6 performs color space conversion on the RGB image data that has undergone gamma correction. The color space conversion process converts the color space of the RGB image data that has undergone gamma correction from the RGB color space to the YCbCr color space. That is, the color space conversion unit 62F6 converts the RGB image data into luminance and color difference signals. The luminance and color difference signals are a Y signal, a Cb signal, and a Cr signal. The Y signal is a signal that indicates luminance. Hereinafter, the Y signal may also be referred to as the luminance signal. The Cb signal is a signal obtained by adjusting a signal obtained by subtracting the luminance component from the B signal. The Cr signal is a signal obtained by adjusting a signal obtained by subtracting the luminance component from the R signal. Hereinafter, the Cb signal and Cr signal may also be referred to as color difference signals.
[0117] The luminance processing unit 62F7 performs luminance filtering on the Y signal. The luminance filtering is a process of filtering the Y signal using a luminance filter (not shown). For example, the luminance filter is a filter that reduces high-frequency noise generated in the demosaic processing and enhances sharpness. The signal processing on the Y signal, i.e., filtering by the luminance filter, is performed in accordance with luminance filter parameters. The luminance filter parameters are parameters set for the luminance filter. The luminance filter parameters define the degree to which high-frequency noise generated in the demosaic processing is reduced and the degree to which sharpness is enhanced. The luminance filter parameters are changed in accordance with, for example, the related information 102 (see FIG. 6), the imaging conditions, and / or an instruction received by the receiving device 76.
[0118] The color difference processing unit 62F8 performs first color difference filtering on the Cb signal. The first color difference filtering is filtering of the Cb signal using a first color difference filter (not shown). For example, the first color difference filter is a low-pass filter that reduces high-frequency noise contained in the Cb signal. The signal processing on the Cb signal, i.e., filtering by the first color difference filter, is performed according to specified first color difference filter parameters. The first color difference filter parameters are parameters set for the first color difference filter. The first color difference filter parameters define the degree to which high-frequency noise contained in the Cb signal is reduced. The first color difference filter parameters are changed according to, for example, the related information 102 (see FIG. 6), the imaging conditions, and / or an instruction received by the receiving device 76.
[0119] The color difference processing unit 62F9 performs second color difference filtering on the Cr signal. The second color difference filtering is filtering of the Cr signal using a second color difference filter (not shown). For example, the second color difference filter is a low-pass filter that reduces high-frequency noise contained in the Cr signal. The signal processing on the Cr signal, i.e., filtering by the second color difference filter, is performed according to specified second color difference filter parameters. The second color difference filter parameters are parameters set for the second color difference filter. The second color difference filter parameters define the degree to which high-frequency noise contained in the Cr signal is reduced. The second color difference filter parameters are changed according to, for example, the related information 102 (see FIG. 6), the imaging conditions, and / or an instruction received by the receiving device 76.
[0120] The resizing unit 62F10 performs resizing on the luminance and color difference signals. The resizing process adjusts the luminance and color difference signals so that the size of the image represented by the luminance and color difference signals matches the size specified by the user or the like.
[0121] The compression processing unit 62F11 performs compression processing on the resized luminance and color difference signals. The compression processing is, for example, processing in which the luminance and color difference signals are compressed according to a predetermined compression method. Examples of the predetermined compression method include JPEG, TIFF, and JPEG XR. A processed image 75B is obtained by performing compression processing on the luminance and color difference signals. The compression processing unit 62F11 stores the processed image 75B in the image memory 46.
[0122] Next, the operation of the imaging device 10 will be described with reference to Fig. 9. Fig. 9 shows an example of the flow of image quality adjustment processing executed by the CPU 62.
[0123] 9, first, in step ST100, the AI processing unit 62A determines whether or not the image sensor 20 (see FIG. 2) has generated an inference-use RAW image 75A2 (see FIG. 5). If the image sensor 20 has not generated an inference-use RAW image 75A2 in step ST100, the determination is negative, and the image quality adjustment process proceeds to step ST126. If the image sensor 20 has generated an inference-use RAW image 75A2 in step ST100, the determination is positive, and the image quality adjustment process proceeds to step ST102.
[0124] In step ST102, the AI processing unit 62A acquires a RAW image for inference 75A2 from the image sensor 20. The non-AI processing unit 62B also acquires a RAW image for inference 75A2 from the image sensor 20. After the processing of step ST102 is executed, the image quality adjustment processing proceeds to step ST104.
[0125] In step ST104, the AI processing unit 62A inputs the RAW image for inference 75A2 acquired in step ST102 to the trained NN 82. After the processing of step ST104 is executed, the image quality adjustment processing proceeds to step ST106.
[0126] In step ST106, the weighting unit 62D acquires the first image 75D output from the trained NN 82 as a result of the inference-use RAW image 75A2 being input to the trained NN 82 in step ST104. After the processing of step ST106 is executed, the image quality adjustment processing proceeds to step ST108.
[0127] In step ST108, the non-AI processing unit 62B adjusts noise contained in the RAW image for inference 75A2 by non-AI filtering the RAW image for inference 75A2 acquired in step ST102 using the digital filter 100. After the processing of step ST108 is executed, the image quality adjustment processing proceeds to step ST110.
[0128] In step ST110, the weighting unit 62D acquires the second image 75E obtained by adjusting the noise contained in the RAW image for inference 75A2 using a non-AI method in step ST108. After the processing of step ST110 is executed, the image quality adjustment processing proceeds to step ST112.
[0129] In step ST112, the weight derivation unit 62C acquires the related information 102 from the NVM 64. After the processing of step ST112 is executed, the image quality adjustment processing proceeds to step ST114.
[0130] In step ST114, the weight derivation unit 62C extracts the sensitivity related information 102A from the related information 102 acquired in step ST112. After the process of step ST114 is executed, the image quality adjustment process proceeds to step ST116.
[0131] In step ST116, the weight derivation unit 62C calculates the first weight 104 and the second weight 106 based on the sensitivity related information 102A extracted in step ST114. That is, the weight derivation unit 62C identifies a value indicating the sensitivity of the image sensor 20 from the sensitivity related information 102A, calculates the first weight 104 by substituting the value indicating the sensitivity of the image sensor 20 into the weight calculation formula 108, and calculates the second weight 106 from the calculated first weight 104. After the processing of step ST116 is executed, the image quality adjustment processing proceeds to step ST118.
[0132] In step ST118, the weighting unit 62D applies the first weight 104 calculated in step ST116 to the first image 75D acquired in step ST106. After the process of step ST118 is executed, the image quality adjustment process proceeds to step ST120.
[0133] In step ST120, the weighting unit 62D applies the second weight 106 calculated in step ST116 to the second image 75E acquired in step ST110. After the process of step ST120 is executed, the image quality adjustment process proceeds to step ST122.
[0134] In step ST122, the synthesis unit 62E generates a synthesized image 75F by synthesizing the first image 75D and the second image 75E in accordance with the first weight 104 assigned to the first image 75D in step ST118 and the second weight 106 assigned to the second image 75E in step ST120. That is, the synthesis unit 62E generates a synthesized image 75F (for example, a weighted average image using the first weight 104 and the second weight 106) by synthesizing the pixel values for each pixel between the first image 75D and the second image 75E in accordance with the first weight 104 and the second weight 106. After the processing of step ST122 is executed, the image quality adjustment processing proceeds to step ST124.
[0135] In step ST124, the signal processing unit 62F performs various signal processing (e.g., offset correction processing, white balance correction processing, demosaic processing, color correction processing, gamma correction processing, color space conversion processing, luminance filtering processing, first color difference filtering processing, second color difference filtering processing, resizing processing, and compression processing) on the composite image 75F obtained in step ST22, and outputs the resulting image to a predetermined output destination (e.g., image memory 46) as a processed image 75B. After the processing of step ST124 is executed, the image quality adjustment processing proceeds to step ST126.
[0136] In step ST126, the signal processing unit 62F determines whether or not a condition for terminating the image quality adjustment process (hereinafter referred to as the "termination condition") has been satisfied. An example of the termination condition is that an instruction to terminate the image quality adjustment process has been accepted by the acceptance device 76. If the termination condition has not been satisfied in step ST126, the determination is negative, and the image quality adjustment process proceeds to step ST100. If the termination condition has been satisfied in step ST126, the determination is positive, and the image quality adjustment process terminates.
[0137] As described above, the imaging device 10 obtains the first image 75D by processing the RAW image for inference 75A2 using the AI method with the trained NN 82. The imaging device 10 also obtains the second image 75E without processing the RAW image for inference 75A2 using the AI method. Here, due to a characteristic of the trained NN 82, there is a risk that when noise contained in the RAW image 75A is removed, the fine structure may also be removed. Meanwhile, the second image 75E still contains the fine structure that was removed from the RAW image for inference 75A2 by the trained NN 82. Therefore, the imaging device 10 generates a composite image 75F by combining the first image 75D and the second image 75E. This allows for both suppressing the excess or deficiency of noise contained in the image and suppressing the excess or deficiency of sharpness of the fine structure of the subject captured in the image, compared to when the image is processed only using the AI method with the trained NN 82. Therefore, with this configuration, it is possible to obtain an image with adjusted image quality compared to when the image is processed only by an AI method using the trained NN82.
[0138] Furthermore, in the imaging device 10, noise is adjusted by combining a first image 75D obtained by performing AI noise adjustment processing on the RAW image for inference 75A2 with a second image 75E obtained from the RAW image for inference 75A2 without processing it with AI. Therefore, with this configuration, an image can be obtained that has undergone only AI noise adjustment processing, i.e., an image in which both excess noise and loss of fine structure are suppressed compared to the first image 75D.
[0139] Furthermore, in the imaging device 10, noise is adjusted by combining a first image 75D obtained by performing AI noise adjustment processing on the RAW image for inference 75A2 with a second image 75E obtained by performing non-AI noise adjustment processing on the RAW image for inference 75A2. Therefore, with this configuration, an image that has been subjected to only AI noise adjustment processing, that is, an image in which both excess noise and loss of fine structure are suppressed compared to the first image 75D, can be obtained.
[0140] Furthermore, in imaging device 10, first weight 104 is assigned to first image 75D, and second weight 106 is assigned to second image 75E. Then, first image 75D and second image 75E are composited according to first weight 104 assigned to first image 75D and second weight assigned to second image 75E. Therefore, with this configuration, it is possible to obtain composite image 75F, an image in which the degree of influence of first image 75D and the degree of influence of second image 75E on image quality have been adjusted.
[0141] Furthermore, imaging device 10 combines first image 75D and second image 75E by performing a weighted averaging using first weight 104 and second weight 106. Therefore, with this configuration, it is easier to combine first image 75D and second image 75E and adjust the degree of influence that first image 75D and second image 75E have on the image quality of combined image 75F, compared to a case in which first image 75D and second image 75E are combined and then the degree of influence that first image 75D and second image 75E have on the image quality of the combined image are adjusted.
[0142] Furthermore, in the imaging device 10, the first weight 104 and the second weight 106 are changed in accordance with the related information 102. Therefore, according to this configuration, it is possible to suppress degradation of image quality caused by the related information 102, compared to when a fixed weight determined solely on the basis of information completely unrelated to the related information 102 is used.
[0143] Furthermore, in the imaging device 10, the first weight 104 and the second weight 106 are changed in accordance with the sensitivity related information 102A included in the related information 102. Therefore, with this configuration, it is possible to suppress degradation in image quality caused by the sensitivity of the image sensor 20, compared to when a fixed weight is used that is determined solely based on information that is completely unrelated to the sensitivity of the image sensor 20 used in imaging to obtain the inference-use RAW image 75A2.
[0144] In this embodiment, the weighting formula 108 is used to calculate the first weight 104 from a value indicating the sensitivity of the image sensor 20, but the technology of the present disclosure is not limited to this, and a weighting formula may be used to calculate the second weight 106 from the weighting formula 108. In this case, the first weight 104 is calculated from the second weight 106.
[0145] Furthermore, in this embodiment, the weight calculation formula 108 is exemplified, but the technology of the present disclosure is not limited to this, and a weight derivation table in which a value indicating the sensitivity of the image sensor 20 is associated with the first weight 104 or the second weight 106 may be used.
[0146] [First Modification] The trained NN 82 has the property that it is more difficult to distinguish between noise and fine structure in bright image regions than in dark image regions. This property becomes more pronounced as the layer structure of the trained NN 82 is simplified. In this case, as an example, as shown in FIG. 10, it is preferable that the related information 102 includes brightness-related information 102B related to the brightness of the inference-use RAW image 75A2, and that the first weight 104 and the second weight 106 based on the brightness-related information 102B are derived by the weight derivation unit 62C.
[0147] An example of the brightness-related information 102B is pixel statistics of at least a portion of the inference-use RAW image 75A2. The pixel statistics may be, for example, an average pixel value.
[0148] 10, the RAW image for inference 75A2 is divided into multiple divided areas 75A2a, and the related information 102 includes the average pixel value for each divided area 75A2a. The average pixel value refers to, for example, the average pixel value of all pixels included in the divided area 75A2a. The average pixel value is calculated by the CPU 62, for example, each time the RAW image for inference 75A2 is generated.
[0149] The weight calculation formula 110 is stored in the NVM64. The weight derivation unit 62C acquires the weight calculation formula 110 from the NVM64, and calculates the first weight 104 and the second weight 106 using the acquired weight calculation formula 110.
[0150] The weight calculation formula 110 is a formula with the pixel average value as the independent variable and the first weight 104 as the dependent variable. The first weight 104 is changed according to the pixel average value. The correlation between the pixel average value indicated by the weight calculation formula 110 and the first weight 104 is, for example, the first weight 104 less than the threshold th1 of the pixel average value is a fixed value of "w1". Also, the first weight 104 exceeding the threshold th2 (>th1) of the pixel average value is a fixed value of "w2(<w1)". In the range of not less than the threshold th1 and not more than the threshold th2, the first weight 104 decreases as the pixel average value increases. In the example shown in FIG. 10, the first weight changes only between the threshold th1 and the threshold th2, but this is merely an example, and the weight calculation formula 110 may be any formula defined such that the first weight 104 changes according to the pixel average value regardless of the thresholds th1 and th2.
[0151] Also, as the image area is brighter, it becomes more difficult to distinguish noise from fine structures, so it is preferable that the first weight 104 decreases as the pixel average value increases. This is because it is to suppress the degree to which pixels whose discrimination between noise and fine structures is not clear affect the composite image 75F. On the other hand, since the second weight 106 is "1 - w", it increases as the first weight 104 decreases. That is, as the first weight 104 decreases, the degree to which the second image 75E affects the composite image 75F becomes greater than the degree to which the first image 75D affects the composite image 75F.
[0152] 11, the first image 75D is divided into a plurality of divided areas 75D1, and the second image 75E is also divided into a plurality of divided areas 75E1. The positions of the plurality of divided areas 75D1 in the first image 75D correspond to the positions of the plurality of divided areas 75A2a in the RAW image for inference 75A2, and the positions of the plurality of divided areas 75E1 in the second image 75E also correspond to the positions of the plurality of divided areas 75A2a in the RAW image for inference 75A2.
[0153] The weighting unit 62D assigns to each divided area 75D1 the first weight 104 calculated by the weight derivation unit 62C for the divided area 75A2a whose position corresponds to that of the divided area 75D1. The weighting unit 62D also assigns to each divided area 75E1 the second weight 106 calculated by the weight derivation unit 62C for the divided area 75A2a whose position corresponds to that of the divided area 75E1.
[0154] The composition unit 62E generates composite image 75F by combining divided areas 75D1 and 75E1, which correspond to each other in position, according to first weight 104 and second weight 106. Combining divided areas 75D1 and 75E1 according to first weight 104 and second weight 106 is realized, as in the above embodiment, by weighted averaging using, for example, first weight 104 and second weight 106, that is, weighted averaging for each pixel between divided areas 75D1 and 75E1.
[0155] As described above, in the present first modified example, the related information 102 includes brightness-related information 102B related to the brightness of the inference use RAW image 75A2, and the first weight 104 and the second weight 106 corresponding to the brightness-related information 102B are derived by the weight derivation unit 62C. Therefore, according to this configuration, it is possible to suppress degradation in image quality caused by the brightness of the inference use RAW image 75A2, compared to when a fixed weight determined solely on the basis of information completely unrelated to the brightness of the inference use RAW image 75A2 is used.
[0156] Furthermore, in this first modified example, the average pixel value of each divided area 75A2a of the RAW image for inference 75A2 is used as the brightness-related information 102B. Therefore, with this configuration, it is possible to suppress degradation in image quality caused by the pixel statistics regarding the RAW image for inference 75A2, compared to when a certain weight determined solely based on information completely unrelated to the pixel statistics regarding the RAW image for inference 75A2 is used.
[0157] Although the first modified example shows an example in which the first weight 104 and the second weight 106 are derived according to the average pixel value for each divided area 75A2a, the technology of the present disclosure is not limited to this, and the first weight 104 and the second weight 106 may be derived according to the average pixel value for each frame of the RAW image for inference 75A2, or the first weight 104 and the second weight 106 may be derived according to the average pixel value of a portion of the RAW image for inference 75A2. Furthermore, the first weight 104 and the second weight 106 may be derived according to the luminance of each pixel of the RAW image for inference 75A2.
[0158] Furthermore, in this first modified example, the weight calculation formula 110 is exemplified, but the technology of the present disclosure is not limited to this, and a weight derivation table in which multiple pixel average values are associated with multiple first weights 104 may also be used.
[0159] Furthermore, in the first modified example, the pixel average value is exemplified, but this is merely an example, and instead of the pixel average value, a pixel median value or a pixel mode value may be used.
[0160] [Second Modification] The trained NN 82 has the property that it is more difficult to distinguish between noise and fine structure in an image region of high-frequency components than in an image region of low-frequency components. This property becomes more pronounced as the layer structure of the trained NN 82 is simplified. In this case, as an example, as shown in FIG. 12, the related information 102 may include spatial frequency information 102C indicating the spatial frequency of the inference-use RAW image 75A2, and the first weight 104 and the second weight 106 may be derived by the weight derivation unit 62C according to the spatial frequency information 102C.
[0161] 12 differs from the example shown in Fig. 10 in that spatial frequency information 102C for each divided area 75A2a is applied instead of the pixel average value for each divided area 75A2a, and in that weighting calculation formula 112 is applied instead of weighting calculation formula 110. The spatial frequency information 102C for each divided area 75A2a is calculated by the CPU 62, for example, each time an inference-use RAW image 75A2 is generated.
[0162] The weighting formula 112 is a formula with the spatial frequency information 102C as the independent variable and the first weight 104 as the dependent variable. The first weight 104 is changed according to the spatial frequency information 102C. Furthermore, since the higher the spatial frequency indicated by the spatial frequency information 102C, it becomes more difficult to distinguish between noise and fine structure. Therefore, it is preferable that the first weight 104 decreases with increasing spatial frequency indicated by the spatial frequency information 102C. This is to reduce the degree to which pixels that are unclear as being classified as noise or fine structure affect the composite image 75F. In contrast, the second weight 106 is "1-w," and therefore increases with decreasing first weight 104. In other words, with decreasing first weight 104, the degree to which the second image 75E affects the composite image 75F becomes greater than the degree to which the first image 75D affects the composite image 75F. The method for generating the composite image 75F is as explained in the first modified example.
[0163] As described above, in the present second modified example, the related information 102 includes spatial frequency information 102C indicating the spatial frequency of the inference use RAW image 75A2, and the first weight 104 and the second weight 106 are derived by the weight derivation unit 62C according to the spatial frequency information 102C. Therefore, with this configuration, it is possible to suppress degradation in image quality caused by the spatial frequency of the inference use RAW image 75A2, compared to when a fixed weight is used that is determined solely on the basis of information that is completely unrelated to the spatial frequency of the inference use RAW image 75A2.
[0164] In the second modified example, an example is shown in which the first weight 104 and the second weight 106 are derived according to the spatial frequency information 102C for each divided area 75A2a, but the technology of the present disclosure is not limited to this, and the first weight 104 and the second weight 106 may be derived according to the spatial frequency information 102C for each frame of the RAW image for inference 75A2, or the first weight 104 and the second weight 106 may be derived according to the spatial frequency information 102C of a portion of the RAW image for inference 75A2.
[0165] Furthermore, in the second modified example, the weighting calculation formula 112 is exemplified, but the technology of the present disclosure is not limited to this, and a weighting derivation table in which a plurality of pieces of spatial frequency information 102C and a plurality of first weights 104 are associated with each other may be used.
[0166] [Third Modification] The CPU 62 may detect a subject appearing in the inference-use RAW image 75A2 based on the inference-use RAW image 75A2, and change the first weight 104 and the second weight 106 in accordance with the detected subject. In this case, as shown in FIG. 13 as an example, a weight derivation table 114 is stored in the NVM 64, and the weight derivation unit 62C reads out the weight derivation table 114 from the NVM 64 and derives the first weight 104 and the second weight 106 by referring to the weight derivation table 114. The weight derivation table 114 is a table in which a plurality of subjects are associated one-to-one with a plurality of first weights 104.
[0167] The weight derivation unit 62C has a subject detection function. The weight derivation unit 62C uses the subject detection function to detect subjects that appear in the inference-use RAW image 75A2. The subject detection may be an AI-based detection or a non-AI-based detection (e.g., detection by template matching).
[0168] The weight derivation unit 62C derives the first weight 104 corresponding to the detected subject from the weight derivation table 114, and calculates the second weight 106 from the derived first weight 104. Since a different first weight 104 is associated with each subject in the weight derivation table 114, the first weight 104 applied to the first image 75D and the second weight 106 applied to the second image 75E are changed depending on the subject detected from the inference-use RAW image 75A2.
[0169] Alternatively, the weighting unit 62D may assign the first weight 104 only to the image region of the entire image region of the first image 75D that indicates the subject detected by the weight derivation unit 62C, and may assign the second weight 106 only to the image region of the entire image region of the second image 75E that indicates the subject detected by the weight derivation unit 62C. Then, a compositing process according to the first weight 104 and the second weight 106 may be performed only on the image region to which the first weight 104 has been assigned and the image region to which the second weight 106 has been assigned. However, this is merely an example, and the first weight 104 may be assigned to the entire image region of the first image 75D, the second weight 106 may be assigned to the entire image region of the second image 75E, and a compositing process according to the first weight 104 and the second weight 106 may be performed on the entire image region of the first image 75D and the entire image region of the second image 75E.
[0170] In this way, in the third modified example, a subject reflected in the inference RAW image 75A2 is detected, and the first weight 104 and the second weight 106 are changed according to the detected subject. Therefore, with this configuration, it is possible to suppress degradation in image quality caused by a subject reflected in the inference RAW image 75A2, compared to when a fixed weight determined solely on the basis of information completely unrelated to the subject reflected in the inference RAW image 75A2 is used.
[0171] [Fourth Modification] The CPU 62 may detect the parts of the subject that appear in the inference-use RAW image 75A2 based on the inference-use RAW image 75A2, and change the first weight 104 and the second weight 106 in accordance with the detected parts. In this case, as shown in FIG. 13 as an example, a weight derivation table 116 is stored in the NVM 64, and the weight derivation unit 62C reads out the weight derivation table 116 from the NVM 64 and derives the first weight 104 and the second weight 106 by referring to the weight derivation table 116. The weight derivation table 116 is a table in which a plurality of parts of the subject are associated one-to-one with a plurality of first weights 104.
[0172] The weight derivation unit 62C has a subject part detection function. By using the subject part detection function, the weight derivation unit 62C detects the part of the subject (for example, a person's face and / or a person's eyes) that appears in the inference-use RAW image 75A2. The detection of the subject part may be an AI-based detection or a non-AI-based detection (for example, detection by template matching).
[0173] The weight derivation unit 62C derives the first weight 104 corresponding to the part of the detected subject from the weight derivation table 116, and calculates the second weight 106 from the derived first weight 104. Since the weight derivation table 114 associates a different first weight 104 with each part of the subject, the first weight 104 applied to the first image 75D and the second weight 106 applied to the second image 75E are changed depending on the part of the subject detected from the inference-use RAW image 75A2.
[0174] Alternatively, the weighting unit 62D may assign the first weight 104 only to image regions of the entire image region of the first image 75D that indicate the part of the subject detected by the weight derivation unit 62C, and may assign the second weight 106 only to image regions of the entire image region of the second image 75E that indicate the part of the subject detected by the weight derivation unit 62C. Then, a compositing process according to the first weight 104 and the second weight 106 may be performed only on the image regions to which the first weight 104 has been assigned and the image regions to which the second weight 106 has been assigned. However, this is merely an example, and the first weight 104 may be assigned to the entire image region of the first image 75D, the second weight 106 may be assigned to the entire image region of the second image 75E, and a compositing process according to the first weight 104 and the second weight 106 may be performed on the entire image region of the first image 75D and the entire image region of the second image 75E.
[0175] In this way, in the fourth modified example, the parts of the subject that appear in the inference RAW image 75A2 are detected, and the first weight 104 and the second weight 106 are changed depending on the detected parts. Therefore, with this configuration, it is possible to suppress degradation in image quality caused by the parts of the subject that appear in the inference RAW image 75A2, compared to when a fixed weight that is determined based only on information that is completely unrelated to the parts of the subject that appear in the inference RAW image 75A2 is used.
[0176] [Fifth Modification] The CPU 62 may change the first weight 104 and the second weight 106 depending on the degree of difference between the feature values of the first image 75D and the feature values of the second image 75E. As an example, as shown in Fig. 14, the weight derivation unit 62C calculates the pixel average value for each divided area 75D1 of the first image 75D as the feature value of the first image 75D, and calculates the pixel average value for each divided area 75E1 of the second image 75E as the feature value of the second image 75E. The weight derivation unit 62C calculates the difference in pixel average values (hereinafter simply referred to as "difference") as the degree of difference between the feature values of the first image 75D and the feature values of the second image 75E for each divided area 75D1 and 75E1 that correspond to each other.
[0177] The weight derivation unit 62C derives the first weight 104 by referring to the weight derivation table 118. The weight derivation table 118 associates a plurality of differences with a plurality of first weights 104 in a one-to-one correspondence. The weight derivation unit 62C derives the first weight 104 corresponding to the calculated difference for each of the divided areas 75D1 and 75E1 from the weight derivation table 118, and calculates the second weight 106 from the derived first weight 104. Because the weight derivation table 118 associates a different first weight 104 for each difference, the first weight 104 applied to the first image 75D and the second weight 106 applied to the second image 75E are changed depending on the difference.
[0178] In this way, in the fifth modified example, the first weight 104 and the second weight 106 are changed according to the degree of difference between the feature values of the first image 75D and the feature values of the second image 75E. Therefore, according to this configuration, it is possible to suppress degradation in image quality caused by the degree of difference between the feature values of the first image 75D and the feature values of the second image 75E, compared to when a fixed weight is used that is determined solely on the basis of information that is completely unrelated to the degree of difference between the feature values of the first image 75D and the feature values of the second image 75E.
[0179] In this fifth variant, an example has been given in which the difference in pixel average values is calculated for each divided area 75D1 and 75E1, but the technology of the present disclosure is not limited to this, and the difference in pixel average values may be calculated for each frame, or the difference in pixel values may be calculated for each pixel.
[0180] Furthermore, in this fifth variant, pixel average values are used as examples of the feature values of the first image 75D and the second image 75E, but the technology of the present disclosure is not limited to this and may also be pixel medians, pixel modes, etc.
[0181] Furthermore, in the fifth modified example, the weight derivation table 118 is exemplified, but the technology of the present disclosure is not limited to this, and an arithmetic expression in which the difference is an independent variable and the first weight 104 is a dependent variable may be used.
[0182] [Sixth Modification] A trained NN 82 may be provided for each imaging scene. In this case, as shown in FIG. 15 as an example, a plurality of trained NNs 82 are stored in the NVM 64. A trained NN 82 in the NVM 64 is created for each imaging scene. An ID 82A is assigned to each trained NN 82. The ID 82A is an identifier that can identify the trained NN 82. The CPU 62 switches the trained NN 82 to be used for each imaging scene, and changes the first weight 104 and the second weight 106 according to the trained NN 82 to be used.
[0183] 15, an NN determination table 120 and an NN weight table 122 are stored in the NVM 64. In the NN determination table 120, a plurality of imaging scenes and a plurality of IDs 82A are associated one-to-one. In the NN weight table 122, a plurality of IDs 82A are associated one-to-one with a plurality of first weights 104.
[0184] As an example, as shown in FIG. 16, the AI processing unit 62A has an imaging scene detection function. The AI processing unit 62A activates the imaging scene detection function to detect a scene captured in the inference-use RAW image 75A2 as an imaging scene. The imaging scene may be detected by an AI method or a non-AI method (e.g., detection by template matching). The imaging scene may be determined according to an instruction received by the receiving device 76.
[0185] The AI processing unit 62A derives an ID 82A corresponding to the detected captured scene from the NN determination table 120, and acquires a trained NN 82 identified from the derived ID 82A from the NVM 64. Then, the AI processing unit 62A inputs the RAW image for inference 75A2 that was the detection target of the captured scene to the trained NN 82, thereby acquiring a first image 75D.
[0186] 17, the weight derivation unit 62C derives the first weight 104 corresponding to the ID 82A of the trained NN 82 used in the AI processing unit 62A from the per-NN weight table 122, and calculates the second weight 106 from the derived first weight 104. Since the per-NN weight table 122 associates a different first weight 104 with each ID 82A, the first weight 104 applied to the first image 75D and the second weight 106 applied to the second image 75E are changed depending on the trained NN 82 used in the AI processing unit 62A.
[0187] In the sixth modified example, a trained NN 82 is provided for each imaging scene, and the trained NN 82 used in the AI processing unit 62A is switched for each imaging scene. The first weight 104 and the second weight 106 are changed depending on the trained NN 82 used in the AI processing unit 62A. Therefore, with this configuration, even if the trained NN 82 is switched for each imaging scene, it is possible to suppress a deterioration in image quality that accompanies switching of the trained NN 82 for each imaging scene, compared to a case where a fixed weight is always used.
[0188] In the sixth modified example, the NN determination table 120 and the NN weight table 122 are separate tables, but they may be combined into one table. In this case, for example, it is sufficient if the table has a one-to-one correspondence between the ID 82A and the first weight 104 for each imaging scene.
[0189] [Seventh Modification] The CPU 62 may normalize the RAW image for inference 75A2 input to the trained NN 82 with respect to predetermined image characteristic parameters. The image characteristic parameters are parameters determined according to the image sensor 20 and imaging conditions used in capturing the image to obtain the RAW image for inference 75A2 input to the trained NN 82. In the seventh modified example, as shown in FIG. 18 as an example, the image characteristic parameters are the number of bits of each pixel (hereinafter also referred to as the "image characteristic bit number") and an offset value related to optical black (hereinafter also referred to as the "OB offset value"). For example, the image characteristic bit number is 14 bits, and the OB offset value is 1024 LSB.
[0190] As an example, as shown in Fig. 18, a learning execution system 124 differs from the learning execution system 84 shown in Fig. 4 in that a learning execution device 126 is used instead of the learning execution device 88. The learning execution device 126 differs from the learning execution device 88 in that it has a normalization processing unit 128.
[0191] The normalization processing unit 128 acquires the training RAW image 75A1 from the storage device 86 and normalizes the acquired training RAW image 75A1 with respect to the image characteristic parameters. For example, the normalization processing unit 128 adjusts the number of image characteristic bits of the training RAW image 75A1 acquired from the storage device 86 to 14 bits and adjusts the OB offset value of the training RAW image 75A1 to 1024 LSB. The normalization processing unit 128 inputs the training RAW image 75A1 normalized with respect to the image characteristic parameters to the NN 90. As a result, the trained NN 90 is generated in the same manner as the example shown in FIG. 4. N82 is generated. The trained NN82 is associated with image feature parameters used for normalization, i.e., 14 bits of the image feature bit count and 1024 LSB of the OB offset value. The 14 bits of the image feature bit count and 1024 LSB of the OB offset value are an example of a "first parameter" according to the technology of the present disclosure. Hereinafter, for convenience of explanation, when it is not necessary to distinguish between the image feature bit count and the OB offset value associated with the trained NN82, they will be referred to as the "first parameter."
[0192] 19, the AI method processing unit 62A has a normalization processing unit 130 and a parameter restoration unit 132. The normalization processing unit 130 normalizes the RAW image for inference 75A2 using a first parameter and a second parameter that is the number of image characteristic bits and the OB offset value of the RAW image for inference 75A2.
[0193] In the seventh modified example, the imaging device 10 is an example of a "first imaging device" and a "second imaging device" according to the technology of the present disclosure. The learning RAW image 75A1 normalized by the normalization processing unit 128 is an example of a "learning image" according to the technology of the present disclosure. The learning RAW image 75A1 is an example of a "first RAW image" according to the technology of the present disclosure. The inference RAW image 75A2 is an example of an "inference image" and a "second RAW image" according to the technology of the present disclosure.
[0194] The normalization processing unit 130 normalizes the inference-use RAW image 75A2 using the following formula (1). In formula (1), “B t " is the number of image feature bits associated with the trained NN82, and "O t " is the OB offset value associated with the trained NN82, and "B i " is the number of image characteristic bits of the inference RAW image 75A2, and "O i " is the OB offset value of the RAW image for inference 75A2, "P0" is the pixel value of the RAW image for inference 75A2, and "P1" is the normalized pixel value of the RAW image for inference 75A2.
[0195]
number
[0196] The normalization processing unit 130 inputs the RAW image for inference 75A2 normalized using Equation (1) to the trained NN 82. When the RAW image for inference 75A2 is input to the trained NN 82, the trained NN 82 outputs a normalized noise-adjusted image 134 as the first image 75D defined by the first parameter.
[0197] The parameter restoration unit 132 acquires the normalized noise-adjusted image 134. Then, the parameter restoration unit 132 adjusts the normalized noise-adjusted image 134 to an image of the second parameters using the first parameter and the second parameter. That is, the parameter restoration unit 132 restores the number of image characteristic bits and the OB offset value before normalization by the normalization processing unit 130 from the number of image characteristic bits and the OB offset value of the normalized noise-adjusted image 134 using the following mathematical formula (2). The normalized noise-adjusted image 134 defined by the second parameter restored according to mathematical formula (2) is used as an image to which the first weight 104 is assigned. In mathematical formula (2), "P2" is the pixel value after the inference RAW image 75A2 is restored to the number of image characteristic bits and the OB offset value before normalization by the normalization processing unit 130.
[0198]
number
[0199] In this way, in the seventh modified example, the RAW image for inference 75A2 input to the trained NN 82 is normalized with respect to the default image characteristic parameters. Therefore, with this configuration, it is possible to suppress degradation in image quality caused by differences in the image characteristic parameters of the RAW image for inference 75A2 input to the trained NN 82, compared to when the RAW image for inference 75A2 that has not been normalized with respect to the image characteristic parameters is input to the trained NN 82.
[0200] Furthermore, in the seventh modified example, when training the NN 90, the training images input to the NN 90 are training RAW images 75A1 whose image characteristic parameters have been normalized by the normalization processing unit 128. Therefore, with this configuration, it is possible to suppress degradation in image quality caused by image characteristic parameters differing for each training RAW image 75A1 input to the NN 90 as a training image, compared to when training RAW images 75A1 whose image characteristic parameters have not been normalized are used as training images for the NN 90.
[0201] Furthermore, in the seventh modified example, the RAW image for inference 75A2, whose image characteristic parameters have been normalized by the normalization processing unit 130, is used as the image for inference input to the trained NN 82. Therefore, with this configuration, it is possible to suppress degradation in image quality caused by differences in the image characteristic parameters of the RAW image for inference 75A2 input to the trained NN 82, compared to when the RAW image for inference 75A2, whose image characteristic parameters have not been normalized, is used as the image for inference of the trained NN 82.
[0202] Furthermore, in the seventh modified example, the image characteristic parameters of the normalized noise-adjusted image 134 output from the trained NN 82 are restored to the second parameters of the RAW image for inference 75A2 before normalization by the normalization processing unit 130. Then, the normalized noise-adjusted image 134 restored to the second parameters is used as the first image 75D to which the first weight 104 is assigned. Therefore, with this configuration, it is possible to suppress degradation in image quality compared to a case in which the image characteristic parameters of the normalized noise-adjusted image 134 are not restored to the second parameters of the RAW image for inference 75A2 before normalization by the normalization processing unit 130.
[0203] In this seventh variant, an example has been described in which both the number of image characteristic bits and the OB offset value of the learning RAW image 75A1 are normalized, but the technology disclosed herein is not limited to this, and either the number of image characteristic bits or the OB offset value of the learning RAW image 75A1 may be normalized.
[0204] Furthermore, although the seventh modified example has been described using an example in which both the number of image characteristic bits and the OB offset value of the RAW image for inference 75A2 are normalized, the technology of the present disclosure is not limited to this, and the number of image characteristic bits or the OB offset value of the RAW image for inference 75A2 may be normalized. Note that, preferably, when the number of image characteristic bits of the RAW image for inference 75A1 is normalized in the learning stage, the number of image characteristic bits of the RAW image for inference 75A2 is normalized, and when the OB offset value of the RAW image for inference 75A1 is normalized in the learning stage, the OB offset value of the RAW image for inference 75A2 is normalized.
[0205] Furthermore, in the seventh modified example, normalization is exemplified, but this is merely an example, and instead of normalization, weights assigned to first image 75D and second image 75E may be changed.
[0206] Furthermore, in the seventh modification, the RAW image for inference 75A2 input to the trained NN 82 is normalized. Therefore, even if a plurality of RAW images for inference 75A2 with different image characteristic parameters are applied to a single trained NN 82, degradation of image quality due to variations in the image characteristic parameters can be suppressed. However, the technology of the present disclosure is not limited to this. For example, a trained NN 82 may be stored in the NVM 64 for each image characteristic parameter. In this case, the trained NN 82 may be used depending on the image characteristic parameters of the RAW image for inference 75A2.
[0207] Furthermore, in the seventh modified example, an example is given in which the learning RAW images 75A1 are normalized by the normalization processing unit 128, but normalization of the learning RAW images 75A1 is not essential. That is, if all learning RAW images 75A1 input to the NN 90 are images with constant image characteristic parameters (for example, 14 bits for the number of image characteristic bits and 1024 LSB for the OB offset value), the normalization processing unit 128 is not necessary.
[0208] [Eighth Modification] The CPU 62 performs signal processing on the first image 75D and the second image 75E according to specified setting values, and the setting values may be different when performing signal processing on the first image 75D and when performing signal processing on the second image 75E. In this case, as shown in FIG. 20 as an example, the CPU 62 further includes a parameter adjustment unit 62G. The parameter adjustment unit 62G sets different luminance filter parameters for the luminance processing unit 62F7 when the signal processing unit 62F performs signal processing on the first image 75D and when the signal processing unit 62F performs signal processing on the second image 75E. The luminance filter parameters are an example of a "setting value" according to the technology of the present disclosure.
[0209] The first image 75D, the second image 75E, and the composite image 75F are selectively input to the signal processing unit 62F. To selectively input the first image 75D, the second image 75E, and the composite image 75F to the signal processing unit 62F, for example, the CPU 62 may change the first weight 104. For example, when the first weight 104 is "0," only the second image 75E of the first image 75D, the second image 75E, and the composite image 75F is input to the signal processing unit 62F. Also, when the first weight 104 is "1," only the first image 75D of the first image 75D, the second image 75E, and the composite image 75F is input to the signal processing unit 62F. Furthermore, when the first weight 104 is greater than "0" and less than "1," only the composite image 75F of the first image 75D, the second image 75E, and the composite image 75F is input to the signal processing unit 62F.
[0210] When the first weight 104 is "0," the parameter adjustment unit 62G sets the luminance filter parameter to a first reference value specialized for adjusting the luminance of the second image 75E. For example, the first reference value is a value that can compensate for the sharpness that has been lost from the second image 75E due to the characteristics of the digital filter 100 (see FIG. 5).
[0211] When the first weight 104 is "1," the parameter adjusting unit 62G sets the brightness filter parameter to a second reference value specialized for adjusting the brightness of the first image 75D. For example, the second reference value is a value capable of compensating for the sharpness lost from the first image 75D due to the characteristics of the trained NN 82 (see FIG. 7).
[0212] If the first weight 104 is greater than "0" and less than "1", the parameter adjustment unit 62G changes the brightness filter parameter according to the first weight 104 and the second weight 106 derived by the weight derivation unit 62C, as described in the above embodiment.
[0213] In this way, in the eighth modified example, the luminance filter parameters are different when performing signal processing on first image 75D and when performing signal processing on second image 75E. Therefore, with this configuration, it is possible to achieve sharpness suitable for first image 75D that has been affected by the AI noise adjustment processing, and sharpness suitable for second image 75E that has not been affected by the AI noise adjustment processing, compared to when filtering with a luminance filter is always performed on the Y signal of first image 75D and the Y signal of second image 75E according to the same luminance filter parameters.
[0214] Furthermore, in the eighth modified example, when first weight 104 is "1" or when first weight 104 is greater than "0" but less than "1," filtering using a luminance filter is performed on the Y signal of first image 75D by luminance processing unit 62F7 as a process to compensate for the sharpness lost by the AI noise adjustment process. Therefore, with this configuration, an image with higher sharpness can be obtained compared to when the process to compensate for the sharpness lost by the AI noise adjustment process is not performed on first image 75D.
[0215] Note that, in the eighth modified example, an example in which the luminance filter parameters are different when signal processing is performed on the first image 75D and when signal processing is performed on the second image 75E has been described, but the technology of the present disclosure is not limited to this, and the parameters used in the offset correction processing, the parameters used in the white balance correction processing, the parameters used in the demosaic processing, the parameters used in the color correction processing, the parameters used in the gamma correction processing, the first chrominance filter parameters, the second chrominance filter parameters, the parameters used in the resizing processing, and / or the parameters used in the compression processing may be different when signal processing is performed on the first image 75D and when signal processing is performed on the second image 75E. Furthermore, if the signal processing unit 62F is provided with a sharpness correction processing unit (not shown) that performs sharpness processing to adjust the sharpness of the image, the parameters used in the sharpness correction processing unit (for example, parameters that can adjust the degree of sharpness enhancement) may be different when signal processing is performed on the first image 75D and when signal processing is performed on the second image 75E.
[0216] [Ninth Variation] The trained NN82 has the property that it is more difficult to distinguish between noise and fine structure in bright image regions than in dark image regions. This property becomes more pronounced as the layer structure of the trained NN82 becomes simpler. If it is more difficult to distinguish between noise and fine structure in bright image regions than in dark image regions, the trained NN82 will distinguish the fine structure as noise and remove it, which is expected to result in an image lacking sharpness as the first image 75D. One possible cause of the lack of sharpness in the first image 75D is a lack of luminance, which forms the fine structure. This is because, although luminance contributes more to the formation of the fine structure than color, it is more likely to be distinguished as noise and removed by the trained NN82.
[0217] Therefore, in the ninth modification, first image 75D and second image 75E to be combined in the combining process are converted into images represented by Y, Cb, and Cr signals, and signal processing is performed on first image 75D and second image 75E so as to weight the Y signal of second image 75E more heavily than the Y signal of first image 75D and to weight the Cb and Cr signals of the first image more heavily than the Cb and Cr signals of second image 75E. Specifically, signal processing is performed on first image 75D and second image 75E so as to make the signal level of the Y signal of second image 75E higher than that of first image 75D and to make the signal levels of the Cb and Cr signals of first image 75D higher than those of second image 75E in accordance with first weight 104 and second weight 106.
[0218] In this case, as shown in FIG. 21 as an example, the CPU 62 has a signal processing unit 62H instead of the synthesis unit 62E and signal processing unit 62F described in the above embodiment. The signal processing unit 62H has a first image processing unit 62H1, a second image processing unit 62H2, a synthesis processing unit 62H3, a resizing processing unit 62H4, and a compression processing unit 62H5. The first image processing unit 62H1 acquires a first image 75D from the AI processing unit 62A and performs signal processing on the first image 75D. The second image processing unit 62H2 acquires a second image 75E from the non-AI processing unit 62B and performs signal processing on the second image 75E. The synthesis processing unit 62H3 performs synthesis processing similar to the synthesis unit 62E described above. That is, the synthesis processing unit 62H3 generates the above-mentioned synthesized image 75F by synthesizing the first image 75D, which has been signal-processed by the first image processing unit 62H1, and the second image 75E, which has been signal-processed by the second image processing unit 62H2. The resizing unit 62H4 performs the above-described resizing process on the composite image 75F generated by the composition processing unit 62H3. The compression processing unit 62H5 performs the above-described compression process on the composite image 75F that has been resized by the resizing processing unit 62H4. By performing the compression process, the processed image 75B (see FIGS. 2, 8, and 20) is obtained as described above.
[0219] 22, the first image processing unit 62H1 includes an offset correction unit 62H1a having a function similar to that of the offset correction unit 62F1 described above, a white balance correction unit 62H1b having a function similar to that of the white balance correction unit 62F2 described above, a demosaic processing unit 62H1c having a function similar to that of the demosaic processing unit 62F3 described above, a color correction unit 62H1d having a function similar to that of the color correction unit 62F4 described above, a gamma correction unit 62H1e having a function similar to that of the gamma correction unit 62F5 described above, a color space conversion unit 62H1f having a function similar to that of the color space conversion unit 62F6, and a first image weighting unit 62i. The first image weighting unit 62i includes a luminance processing unit 62H1g having a function similar to that of the luminance processing unit 62F7 described above, a color difference processing unit 62H1h having a function similar to that of the color difference processing unit 62F8 described above, and a color difference processing unit 62H1i having a function similar to that of the color difference processing unit 62F9 described above.
[0220] When the first image 75D is input from the AI processing unit 62A to the first image processing unit 62H1 (see Figure 21), offset correction processing, white balance processing, demosaic processing, color correction processing, gamma correction processing, and color space conversion processing are sequentially performed on the first image 75D.
[0221] The luminance processing unit 62H1g filters the Y signal using a luminance filter in accordance with the luminance filter parameters. The first image weighting unit 62i acquires the first weight 104 from the weight derivation unit 62C and sets the acquired first weight 104 for the Y signal output from the luminance processing unit 62H1g. As a result, the first image weighting unit 62i generates a Y signal having a lower signal level than the Y signal of the second image 75E (see FIGS. 23 and 24).
[0222] The color difference processing unit 62H1h performs filtering on the Cb signal using a first color difference filter in accordance with the first color difference filter parameters.
[0223] The color difference processing unit 62H1i performs filtering on the Cr signal using a second color difference filter in accordance with the second color difference filter parameters.
[0224] The first image weighting unit 62i acquires the second weight 106 from the weight derivation unit 62C and sets the acquired second weight 106 to the Cb signal output from the color difference processing unit 62H1h and the Cr signal output from the color difference processing unit 62H1i. As a result, the first image weighting unit 62i generates a Cb signal having a higher signal level than the Cb signal of the second image 75E (see FIGS. 23 and 24), and generates a Cr signal having a higher signal level than the Cr signal of the second image 75E (see FIGS. 23 and 24).
[0225] 23, the second image processing unit 62H2 includes an offset correction unit 62H2a having a function similar to that of the offset correction unit 62F1, a white balance correction unit 62H2b having a function similar to that of the white balance correction unit 62F2, a demosaic processing unit 62H2c having a function similar to that of the demosaic processing unit 62F3, a color correction unit 62H2d having a function similar to that of the color correction unit 62F4, a gamma correction unit 62H2e having a function similar to that of the gamma correction unit 62F5, a color space conversion unit 62H2f having a function similar to that of the color space conversion unit 62F6, and a second image weighting unit 62j. The first image weighting unit 62j includes a luminance processing unit 62H2g having a function similar to that of the luminance processing unit 62F7, a color difference processing unit 62H2h having a function similar to that of the color difference processing unit 62F8, and a color difference processing unit 62H2i having a function similar to that of the color difference processing unit 62F9.
[0226] When the second image 75E is input from the non-AI processing unit 62B to the second image processing unit 62H2 (see Figure 21), offset correction processing, white balance processing, demosaic processing, color correction processing, gamma correction processing, and color space conversion processing are sequentially performed on the second image 75E.
[0227] The luminance processing unit 62H2g filters the Y signal using a luminance filter in accordance with the luminance filter parameters. The second image weighting unit 62j acquires the first weight 104 from the weight derivation unit 62C and sets the acquired first weight 104 for the Y signal output from the luminance processing unit 62H2g. As a result, the second image weighting unit 62j generates a Y signal having a higher signal level than the Y signal of the first image 75D (see FIGS. 22 and 24).
[0228] The color difference processing unit 62H2h performs filtering on the Cb signal using a second color difference filter in accordance with the second color difference filter parameters.
[0229] The color difference processing unit 62H2i performs filtering on the Cr signal using a second color difference filter in accordance with the second color difference filter parameters.
[0230] The second image weighting unit 62j acquires the second weight 106 from the weight derivation unit 62C and sets the acquired second weight 106 to the Cb signal output from the color difference processing unit 62H2h and the Cr signal output from the color difference processing unit 62H2i. As a result, the second image weighting unit 62j generates a Cb signal having a lower signal level than the Cb signal of the first image 75D (see FIGS. 22 and 24), and generates a Cr signal having a lower signal level than the Cr signal of the first image 75D (see FIGS. 22 and 24).
[0231] 24, the synthesis processor 62H3 acquires Y, Cb, and Cr signals from the first image weighting unit 62i as a first image 75D, and acquires Y, Cb, and Cr signals from the second image weighting unit 62j as a second image 75E. The synthesis processor 62H3 then synthesizes the first image 75D represented by the Y, Cb, and Cr signals with the second image 75E represented by the Y, Cb, and Cr signals to generate a synthesized image 75F represented by the Y, Cb, and Cr signals. The resizing processor 62H4 performs the above-described resizing process on the synthesized image 75F generated by the synthesis processor 62H3. The compression processor 62H5 performs the above-described compression process on the resized synthesized image 75F.
[0232] As described above, in the present ninth modification, signal processing is performed on the first image 75D and the second image 75E so that the signal level of the Y signal is made higher in the second image 75E than in the first image 75D, and the signal levels of the Cb and Cr signals are made higher in the first image 75D than in the second image 75E. This makes it possible to both suppress insufficient removal of noise contained in the images and suppress insufficient sharpness in the images, compared to when signal processing is performed on the first image 75D and the second image 75E so that the signal level of the Y signal is made lower in the second image 75E than in the first image 75D, and the signal levels of the Cb and Cr signals are made lower in the first image 75D than in the second image 75E.
[0233] In the ninth modification, the signal level of the Y signal is set higher in the second image 75E than in the first image 75D, and the signal levels of the Cb and Cr signals are set lower than in the second image 75E. Although the above embodiment shows an example in which signal processing is performed on the first image 75D and the second image 75E to increase the signal level of the first image 75D, the technology of the present disclosure is not limited to this. For example, of a first process that reduces the signal level of the Y signal in the second image 75E compared to the first image 75D, and a second process that reduces the signal levels of the Cb signal and the Cr signal in the first image 75D compared to the second image 75E, only the first process may be performed.
[0234] Furthermore, in the ninth modification, an example was described in which the Y signal, Cb signal, and Cr signal obtained from the first image weighting unit 62i are used as the first image 75D, but the technology of the present disclosure is not limited to this. For example, the first image 75D to be combined in the combination process may be an image represented by the Cb signal and Cr signal obtained by performing AI noise adjustment processing on the inference RAW image 75A2. In this case, for example, the weight for the signal output from the luminance processing unit 62H1g may be set to "0." Therefore, with this configuration, noise caused by luminance can be suppressed more effectively than when the Y signal is used as the first image 75D.
[0235] Furthermore, in the ninth modification, an example was described in which the Y signal, Cb signal, and Cr signal obtained from the second image weighting unit 62j are used as the second image 75E. However, the technology of the present disclosure is not limited to this. For example, the second image 75E to be combined in the combination process may be an image represented by a Y signal obtained without performing AI noise adjustment processing on the inference RAW image 75A2. In this case, the weight for the signal output from the color difference processing unit 62H2h may be set to "0," and the weight for the signal output from the color difference processing unit 62H2i may also be set to "0." Therefore, with this configuration, the reduction in sharpness of the fine structure of the combined image 75F obtained by combining the first image 75D and the second image 75E can be suppressed compared to the combined image 75F obtained by combining the first image 75D with an image including Cb and Cr signals as the second image 75E.
[0236] Furthermore, in the ninth modification, an example has been described in which the Y signal, Cb signal, and Cr signal obtained from the first image weighting unit 62i are used as the first image 75D, and the Y signal, Cb signal, and Cr signal obtained from the second image weighting unit 62j are used as the second image 75E. However, the technology of the present disclosure is not limited to this. For example, an image represented by a Cb signal and a Cr signal obtained by performing AI noise adjustment processing on the inference RAW image 75A2 may be used as the first image 75D to be combined in the combination process, and an image represented by a Y signal obtained without performing AI noise adjustment processing on the inference RAW image 75A2 may be used as the second image 75E to be combined in the combination process. In this case, for example, the weight for the signal output from the luminance processing unit 62H1g may be set to "0," the weight for the signal output from the color difference processing unit 62H2h may be set to "0," and the weight for the signal output from the color difference processing unit 62H2i may also be set to "0." Therefore, with this configuration, it is possible to suppress insufficient removal of noise contained in the image and suppress insufficient sharpness of the image, compared to when Y signals, Cb signals, and Cr signals are used for the first image 75D and Y signals, Cb signals, and Cr signals are used for the second image 75E.
[0237] Note that, in the above embodiment (for example, the example shown in FIG. 7), an example was described in which the second weight 106 is assigned to the second image 75E obtained by adjusting noise from the inference-use RAW image 75A2 using a non-AI method, but the technology of the present disclosure is not limited to this. For example, as shown in FIG. 25 as an example, the second weight 106 may be assigned to an image obtained without adjusting noise from the inference-use RAW image 75A2, i.e., the inference-use RAW image 75A2. In this case, the inference-use RAW image 75A2 is an example of a "second image" according to the technology of the present disclosure.
[0238] In this way, when the second weight 106 is assigned to the RAW image for inference 75A2, the composition unit 62E composites the first image 75D and the RAW image for inference 75A2 according to the first weight 104 and the second weight 106. Due to the nature of the trained NN 82, luminance is determined as noise and is therefore excessively removed from the first image 75D, but noise caused by luminance remains in the RAW image for inference 75A2 to which the second weight 106 has been assigned. Therefore, by combining the first image 75D and the RAW image for inference 75A2, it is possible to avoid the loss of fine structures caused by insufficient luminance.
[0239] In the above examples, the image quality adjustment process is performed by the CPU 62 of the image processing engine 12 included in the imaging device 10. However, the technology of the present disclosure is not limited to this example. The device that performs the image quality adjustment process may be provided external to the imaging device 10. In this case, as shown in FIG. 26 , an imaging system 136 may be used. The imaging system 136 includes the imaging device 10 and an external device 138. The external device 138 is, for example, a server. The server is realized, for example, by cloud computing. While cloud computing is used here as an example, this is merely an example. For example, the server may be realized by a mainframe or by network computing such as fog computing, edge computing, or grid computing. Although a server is used here as an example of the external device 138, this is merely an example. Instead of a server, at least one personal computer or the like may be used as the external device 138.
[0240] The external device 138 includes a CPU 140, an NVM 142, a RAM 144, and a communication I / F 146, which are connected to each other via a bus 148. The communication I / F 146 is connected to the imaging device 10 via a network 150. The network 150 is, for example, the Internet. Note that the network 150 is not limited to the Internet, and may be a WAN and / or a LAN such as an intranet.
[0241] The NVM 142 stores an image quality adjustment processing program 80 and a trained NN 82. The CPU 140 executes the image quality adjustment processing program 80 in the RAM 144. The CPU 140 performs the image quality adjustment processing described above in accordance with the image quality adjustment processing program 80 executed on the RAM 144. When performing the image quality adjustment processing, the CPU 140 processes the RAW image for inference 75A2 using the trained NN 82 as described in the above examples. The RAW image for inference 75A2 is transmitted, for example, from the imaging device 10 to the external device 138 via the network 150. The communication I / F 146 of the external device 138 receives the RAW image for inference 75A2. The CPU 140 performs image quality adjustment processing on the RAW image for inference 75A2 received by the communication I / F 146. The CPU 140 generates a composite image 75F by performing the image quality adjustment processing and transmits the generated composite image 75F to the imaging device 10. The imaging device 10 receives the composite image 75 transmitted from the external device 138 via the communication I / F 52 (see FIG. 2).
[0242] In the example shown in FIG. 26, external device 138 is an example of an "information processing device" according to the technology of the present disclosure, CPU 140 is an example of a "processor" according to the technology of the present disclosure, and NVM 142 is an example of a "memory" according to the technology of the present disclosure.
[0243] Furthermore, the image quality adjustment process may be performed in a distributed manner by a plurality of devices including the image capture device 10 and the external device 138 .
[0244] Furthermore, in the above embodiment, CPU 62 is exemplified, but instead of CPU 62 or together with CPU 62, at least one other CPU, at least one GPU, and / or at least one TPU may be used.
[0245] In the above embodiment, an example has been described in which the image quality adjustment processing program 80 is stored in the NVM 62, but the technology of the present disclosure is not limited to this. For example, the image quality adjustment processing program 80 may be stored in a portable non-transitory storage medium such as an SSD or a USB memory. The image quality adjustment processing program 80 stored in the non-transitory storage medium is installed in the image processing engine 12 of the imaging device 10. The CPU 62 executes image quality adjustment processing in accordance with the image quality adjustment processing program 80.
[0246] In addition, the image quality adjustment processing program 80 may be stored in a storage device such as another computer or server device connected to the imaging device 10 via a network, and the image quality adjustment processing program 80 may be downloaded and installed in the image processing engine 12 in response to a request from the imaging device 10.
[0247] It is not necessary to store the entire image quality adjustment processing program 80 in a storage device such as another computer or server device connected to the imaging device 10, or in the NVM 62; only a portion of the image quality adjustment processing program 80 may be stored therein.
[0248] Furthermore, although the imaging device 10 shown in Figures 1 and 2 has a built-in image processing engine 12, the technology of the present disclosure is not limited to this, and for example, the image processing engine 12 may be provided outside the imaging device 10.
[0249] In the above embodiment, the image processing engine 12 is exemplified, but the technology of the present disclosure is not limited to this, and a device including an ASIC, an FPGA, and / or a PLD may be applied instead of the image processing engine 12. Furthermore, instead of the image processing engine 12, a combination of a hardware configuration and a software configuration may be used.
[0250] The hardware resources for executing the image quality adjustment process described in the above embodiments can be various processors, as listed below. Examples of processors include a CPU, which is a general-purpose processor that functions as a hardware resource for executing image quality adjustment processes by executing software, i.e., a program. Examples of processors also include dedicated electrical circuits, such as FPGAs, PLDs, or ASICs, which are processors with circuit configurations designed specifically for executing specific processes. Each processor has built-in or connected memory, and each processor uses the memory to execute the image quality adjustment process.
[0251] The hardware resource that executes the image quality adjustment process may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the image quality adjustment process may be a single processor.
[0252] As an example of configuring a system using a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes the image quality adjustment process. Second, there is a form in which a processor is used that realizes the functions of the entire system, including multiple hardware resources that execute the image quality adjustment process, on a single IC chip, as typified by SoCs. In this way, the image quality adjustment process is realized using one or more of the above-mentioned various processors as hardware resources.
[0253] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The above-described image quality adjustment process is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the process.
[0254] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[0255] In this specification, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0256] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[0257] The following additional notes are provided regarding the above-described embodiments.
[0258] (Appendix 1) a processor; a memory connected to or embedded in the processor; The processor: The captured images are processed using an AI method that uses neural networks. performing a synthesis process of synthesizing a first image obtained by processing the captured image using the AI method and a second image obtained by not processing the captured image using the AI method; At least the first process is performed out of a first process for increasing the weight of the luminance signal of the second image relative to the weight of the luminance signal of the first image and a second process for increasing the weight of the color difference signal of the first image relative to the color difference signal of the second image. Information processing device.
Claims
1. A processor; a memory connected to or embedded in the processor; The processor: An AI noise adjustment process is performed to adjust noise contained in the captured image using an AI method that uses a neural network. The noise is adjusted by performing a synthesis process of synthesizing a first image obtained by performing the AI noise adjustment process on the captured image with a second image that is an image processed without performing the AI noise adjustment process on the captured image. Information processing device.
2. the processor performs a non-AI noise adjustment process that adjusts the noise using a non-AI method that does not use the neural network; The second image is an image obtained by adjusting the noise in the captured image through the non-AI noise adjustment process. The information processing device according to claim 1 .
3. The second image is an image obtained without adjusting the noise in the captured image.
3. The information processing device according to claim 1.
4. The processor: assigning weights to the first image and the second image; The first image and the second image are combined according to the weights. The information processing device according to any one of claims 1 to 3.
5. The weights are classified into a first weight assigned to the first image and a second weight assigned to the second image, The processor combines the first image and the second image by performing a weighted average using the first weight and the second weight. The information processing device according to claim 4 .
6. The processor changes the weights in accordance with related information relating to the captured images.
6. The information processing device according to claim 4 or claim 5.
7. The related information includes sensitivity related information related to the sensitivity of an image sensor used in capturing the captured image. The information processing device according to claim 6 .
8. The related information includes brightness-related information related to the brightness of the captured image.
8. The information processing device according to claim 6 or 7.
9. The brightness-related information is pixel statistics of at least a part of the captured image. The information processing device according to claim 8 .
10. The related information includes spatial frequency information indicating a spatial frequency of the captured image. The information processing device according to any one of claims 6 to 9.
11. The processor: Detecting a subject appearing in the captured image based on the captured image; The weight is changed according to the detected object. The information processing device according to any one of claims 4 to 10.
12. The processor: Detecting a part of a subject appearing in the captured image based on the captured image; The weight is changed according to the detected part. The information processing device according to any one of claims 4 to 11.
13. the neural network is provided for each imaging scene, The processor: switching the neural network for each of the captured scenes; Varying the weights according to the neural network The information processing device according to any one of claims 4 to 12.
14. The processor changes the weights according to the degree of difference between the feature values of the first image and the feature values of the second image. The information processing device according to any one of claims 4 to 13.
15. The processor normalizes the image input to the neural network with respect to image characteristic parameters determined according to the image sensor used in capturing the image and the capturing conditions. The information processing device according to any one of claims 1 to 14.
16. When the neural network is trained, the input to the neural network is The learning image is an image obtained by normalizing a first RAW image obtained by capturing an image using a first imaging device with respect to at least one first parameter selected from the number of bits and the offset value of the first RAW image. The information processing device according to any one of claims 1 to 15.
17. the captured image is an image for inference, the first parameter is associated with the neural network to which the training image is input; When a second RAW image obtained by capturing an image using a second imaging device is input as the inference image to the neural network that has been trained by inputting the training image, the processor normalizes the second RAW image using the first parameter associated with the neural network to which the training image has been input and at least one second parameter selected from the number of bits and an offset value of the second RAW image. The information processing device according to claim 16.
18. the first image is a normalized noise-adjusted image obtained by adjusting the noise in the second RAW image normalized using the first parameter and the second parameter through the AI noise adjustment process using the neural network that has been trained by inputting the learning image, The processor adjusts the normalized noise-adjusted image to an image of the second parameters using the first parameters and the second parameters. The information processing device according to claim 17.
19. the processor performs signal processing on the first image and the second image in accordance with designated setting values; The set value is different between when the signal processing is performed on the first image and when the signal processing is performed on the second image. The information processing device according to any one of claims 1 to 18.
20. The processor performs a process on the first image to compensate for the sharpness lost by the AI noise adjustment process.
20. The information processing device according to claim 1.
21. The first image to be synthesized in the synthesis process is an image represented by a color difference signal obtained by performing the AI noise adjustment process on the captured image.
21. The information processing device according to claim 1.
22. The second image to be synthesized in the synthesis process is an image represented by a luminance signal obtained without performing the AI noise adjustment process on the captured image.
22. The information processing device according to claim 1.
23. the first image to be combined in the combining process is an image represented by a color difference signal obtained by performing the AI noise adjustment process on the captured image, The second image is an image represented by a luminance signal obtained without performing the AI noise adjustment process on the captured image.
23. The information processing device according to claim 1.
24. The second image is an image in which the captured image is processed without being processed by the AI method and in which more noise remains than in the first image.
24. The information processing device according to claim 1.
25. a processor; a memory connected to or embedded in said processor; an image sensor, The processor: performing an AI noise adjustment process to adjust noise contained in the captured image obtained by capturing an image with the image sensor using an AI method using a neural network; The noise is adjusted by performing a synthesis process of synthesizing a first image obtained by performing the AI noise adjustment process on the captured image with a second image that is an image processed without performing the AI noise adjustment process on the captured image. Imaging device.
26. Performing an AI noise adjustment process to adjust noise contained in the captured image obtained by capturing an image using an AI method using a neural network; and The noise is adjusted by performing a synthesis process of synthesizing a first image obtained by performing the AI noise adjustment process on the captured image and a second image that is an image processed without performing the AI noise adjustment process on the captured image. An information processing method including:
27. On the computer, Performing an AI noise adjustment process to adjust noise contained in the captured image obtained by capturing an image using an AI method using a neural network; and A program for executing a process that includes adjusting the noise by performing a synthesis process that combines a first image obtained by performing the AI noise adjustment process on the captured image with a second image that is an image processed on the captured image without performing the AI noise adjustment process.
Citation Information
Patent Citations
Image processing system and medical information processing system
JP2018206382A
Facial image processing apparatus, facial image processing method, and non-transitory computer-readable storage medium
US20180204051A1