Image processing device, imaging device, information processing system, and program

WO2026203795A1PCT designated stage Publication Date: 2026-10-01FUJIFILM CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2026/003182
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-28
Filing Date
2026-01-29
Publication Date
2026-10-01

Smart Images

  • Figure JP2026003182_01102026_PF_FP_ABST
    Figure JP2026003182_01102026_PF_FP_ABST
Patent Text Reader

Abstract

An image processing device according to the present invention comprises a processor. The processor outputs a first image, which is a RAW image, and / or a plurality of first divided images obtained by dividing the first image on the basis of a predetermined unit, and on the basis of the plurality of first divided images, acquires a second image in which noise is reduced more than in the first image.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing apparatus, imaging apparatus, information processing system, and program

[0001] The present disclosure relates to an image processing apparatus, an imaging apparatus, an information processing system, and a program.

[0002] Japanese Unexamined Patent Application Publication No. 2021-179833 discloses a learning data generation apparatus for training a demosaic network. Specifically, a technique for improving demosaicing performance by training a CNN based on a pair of a noisy teacher image and a noisy student image is disclosed.

[0003] Japanese Unexamined Patent Application Publication No. 2021-082211 discloses a technique for improving sharpness by first performing noise reduction on a one-plane RAW image, then performing demosaicing, and further mixing a high-frequency enhanced image extracted from an input image.

[0004] Japanese Unexamined Patent Application Publication No. 2009-153013 discloses an imaging apparatus that generates a color-related signal by calculating a target color signal of an image signal to be processed and another color signal, and performs noise reduction processing on the generated color-related signal in a noise reduction unit.

[0005] One embodiment according to the present disclosure provides an image processing apparatus, an imaging apparatus, an information processing system, and a program that can obtain an image with reduced noise compared to a case where noise reduction processing is performed on a developed image.

[0006] A first aspect according to the present disclosure is an image processing apparatus including a processor, wherein the processor outputs at least one of a first image that is a RAW image and a plurality of first divided images obtained by dividing the first image based on a predetermined unit, and acquires a second image in which noise is reduced compared to the first image based on the plurality of first divided images.

[0007] A second aspect according to the present disclosure is the image processing apparatus according to the first aspect, wherein at least one of the first image and the plurality of first divided images is input to AI, so that the AI generates a plurality of second divided images in which noise of the plurality of first divided images is reduced, and the processor acquires the second image based on the plurality of second divided images.

[0008] A third aspect of this disclosure is an image processing apparatus according to the first or second aspect, wherein the unit is a RAW sequence rule.

[0009] A fourth aspect of this disclosure is an image processing apparatus according to the third aspect, wherein the RAW arrangement rule includes the periodic unit of the color filter.

[0010] A fifth aspect of this disclosure is an image processing apparatus according to the fourth aspect, wherein a plurality of pixels representing a first image are assigned functions different from those of a color filter, and include a plurality of periodically arranged functional pixels, and the RAW arrangement rule includes a periodic unit in which the plurality of functional pixels are arranged.

[0011] A sixth aspect of the present disclosure is an image processing apparatus according to the fifth aspect, wherein different functions include image plane phase difference measurement, infrared measurement, depth measurement, multispectral measurement, and / or polarization measurement.

[0012] A seventh aspect of this disclosure is an image processing apparatus according to the second aspect, wherein the AI ​​is a trained model that has undergone first training to reduce noise in a plurality of first example images obtained by dividing a third image, which is a RAW image, based on units.

[0013] An eighth aspect of the present disclosure is an image processing apparatus according to the seventh aspect, wherein the third image is a first captured image obtained by performing a first imaging and / or an image corresponding to the first captured image, and the first learning is learning using a plurality of first example images and a plurality of first correct answer images obtained by dividing a second captured image obtained by performing a second imaging which has less noise than the first imaging, or an image corresponding to the second captured image, based on units.

[0014] A ninth aspect of the present disclosure is an image processing apparatus according to the eighth aspect, wherein the first learning is learning using a plurality of first example images for each different sensitivity used in the first imaging.

[0015] A tenth aspect of the present disclosure is an image processing apparatus according to the ninth aspect, wherein the third image is an image obtained by mixing a first noise image and a second noise image at a variable mixing ratio, the first noise image includes noise caused by a first sensitivity used in imaging, and the second noise image is used as a first ground truth image and includes noise caused by a second sensitivity lower than the first sensitivity.

[0016] An eleventh aspect of the present disclosure is an image processing device relating to any one of the eighth to tenth aspects, wherein the first learning is learning using a plurality of first example images, first imaging condition information that can identify the first imaging conditions used in imaging to obtain the first captured image, and a plurality of first correct images.

[0017] A twelfth aspect of the present disclosure is an image processing apparatus according to the eleventh aspect, wherein at least one of a first image and a plurality of first segmented images is input to the AI, the AI ​​generates a plurality of second segmented images in which the noise of the plurality of first segmented images has been reduced, the processor acquires a second image based on the plurality of second segmented images, and the plurality of second segmented images are images generated by the AI ​​when at least one of a first image and a plurality of first segmented images and second imaging condition information that can identify second imaging conditions used for imaging to obtain the first image are input to the AI.

[0018] A thirteenth aspect of this disclosure is an image processing apparatus according to the twelfth aspect, wherein the first imaging condition and the second imaging condition include ISO sensitivity, shutter speed, and / or F-number.

[0019] A fourteenth aspect of the present disclosure is an image processing device according to the twelfth or thirteenth aspect, wherein the first imaging condition information is information in which the first imaging condition is represented as a one-hot vector, and the second imaging condition information is information in which the second imaging condition is represented as a one-hot vector.

[0020] The 15th aspect of this disclosure is an image processing device relating to any one of the 7th to 14th aspects, wherein the AI ​​is a trained model requiring a convolution operation, and the kernel size used in the convolution operation is a size predetermined according to the units as a size that can maintain the color periodicity of the third image in the convolution operation.

[0021] A sixteenth aspect of the present disclosure is an image processing device relating to any one of the seventh to fifteenth aspects, wherein a plurality of first example images are a first image group obtained by averaging a plurality of first single-channel images obtained by converting a third image into a single channel for each unit, and these first single-channel images are averaged among the same colors within the unit, or a second image group including a plurality of first single-channel images and at least one first additional image, wherein the first additional image is an image obtained by averaging the first single-channel images of the same colors within the unit among the plurality of first single-channel images.

[0022] A 17th aspect of the present disclosure is an image processing apparatus according to the 16th aspect, wherein a plurality of first segmented images are a third image group obtained by averaging a plurality of second single-channel images obtained by single-channelizing each unit of the first image, and the same colors within the unit, or a fourth image group including a plurality of second single-channel images and at least one second additional image, and the second additional image is an image obtained by averaging the second single-channel images of the same colors within the unit of the plurality of second single-channel images.

[0023] The eighteenth aspect of the present disclosure is an image processing apparatus according to the seventeenth aspect, wherein at least one of a first image and a plurality of first segmented images is input to the AI, the AI ​​generates a plurality of second segmented images in which the noise of the plurality of first segmented images has been reduced, the second image is an image obtained by mixing the first image and a third single-channel image, and the third single-channel image is an image in which the plurality of second segmented images have been converted to single channels, and the resolution after averaging between the same colors has been restored to the resolution before averaging between the same colors by upsampling or super-resolution.

[0024] A 19th aspect of the present disclosure is an image processing device according to any one of the seventh to eighteenth aspects, wherein the AI ​​is a trained model that has undergone a second training to reduce noise in a third image.

[0025] A 20th aspect of the present disclosure is an image processing apparatus according to the 19th aspect, wherein the third image is a third captured image obtained by performing a third imaging and / or an image corresponding to the third captured image, and the second learning is learning using the third image and a second ground truth image which is a fourth captured image obtained by performing a fourth imaging which has less noise than the third imaging, or an image corresponding to the fourth captured image.

[0026] A 21st aspect of the present disclosure is an image processing device according to a 20th aspect, wherein the image corresponding to the third captured image is a single image obtained by an AI in the learning stage when a plurality of first example images are input to the AI, and a plurality of inferred images obtained by the AI ​​are converted into single channels based on units.

[0027] A 22nd aspect of the present disclosure is an image processing device according to the 7th aspect, wherein the AI ​​is a trained model that has undergone a first training to reduce noise in a plurality of first example images obtained by dividing a third image, which is a RAW image, based on units, and a second training to reduce noise in the third image or an image corresponding to the third image, the third image being a first captured image obtained by performing a first imaging and / or an image corresponding to the first captured image, the first training being training using a first difference degree which is the degree of difference between a plurality of first example images and a plurality of first ground truth images obtained by dividing a second captured image or an image corresponding to the second captured image, obtained by performing a second imaging which has less noise than the first imaging, based on units, the second training being training using a second difference degree which is the degree of difference between the third image or an image corresponding to the third image and a third ground truth image which is a second captured image or an image corresponding to the second captured image, and the first difference degree and the second difference degree are evaluated by different indices.

[0028] A 23rd aspect of this disclosure is an image processing apparatus according to the 22nd aspect, wherein a first difference is evaluated by a first indicator, a second difference is evaluated by a second indicator, and in the AI ​​learning process, the first indicator has a higher effect of reducing image noise than the second indicator, and the second indicator has a higher effect of preserving the image structure of the image than the first indicator.

[0029] A 24th aspect of this disclosure is an image processing apparatus according to the 23rd aspect, wherein the first indicator is the mean absolute error of a zone different from the volume zone in the histogram of the mean absolute error of each pixel of a plurality of first example images and a plurality of first correct answer images, or the mean square error of a zone different from the volume zone in the histogram of the mean square error of each pixel of a plurality of first example images and a plurality of first correct answer images.

[0030] A 25th aspect of the present disclosure is an image processing apparatus according to any one of the 7th to 24th aspects, wherein an image representing frequencies in the Nyquist limit region is used as one of the third types of images.

[0031] A 26th aspect of the present disclosure is an image processing apparatus according to any one of the first to 25th aspects, wherein a processor acquires unit identification information that allows for the identification of units, and a plurality of first divided images are obtained by dividing the first image according to the units identified by the unit identification information.

[0032] A 27th aspect of this disclosure is an image processing apparatus according to the 26th aspect, wherein the first image is an image obtained by imaging a device, and the unit identification information is information included in metadata obtained in connection with imaging to obtain the first image by the imaging device.

[0033] A 28th aspect of the present disclosure is an image processing apparatus according to the second aspect, wherein the second image is a single image in which a plurality of second segmented images have been converted into a single channel based on a unit.

[0034] A 29th aspect of the present disclosure is an imaging device comprising an image processing device according to any one of the first to 28th aspects, and an image sensor for capturing an image to obtain a first image.

[0035] A 30th aspect of the present disclosure is an information processing system comprising an image processing device relating to any one of the first to 28th aspects, and an information processing device that performs processing based on a second image.

[0036] A 31st aspect of this disclosure is a program for causing a computer to perform a process that includes outputting at least one of a first image which is a RAW image and a plurality of first divided images obtained by dividing the first image based on predetermined units, and obtaining a second image which has less noise than the first image based on the plurality of first divided images.

[0037] This is a schematic diagram showing an example of the configuration of an imaging device. This is a conceptual diagram showing an example of a state in which noise is present in a RAW image obtained by imaging the imaging device. This is a conceptual diagram showing an example of the processing content when noise reduction processing using image processing AI is performed by the processor of the imaging device. This is a schematic diagram showing an example of the configuration of a learning execution device. This is a conceptual diagram showing an example of the configuration of a model. This is a conceptual diagram showing an example of the processing content of the multi-channel layer during the learning stage. This is a conceptual diagram showing an example of the optimization processing performed by the processor of the learning execution device. This is a conceptual diagram showing an example of the processing content of the single-channel layer during the learning stage. This is a conceptual diagram showing an example of the processing content of the multi-channel layer included in the image processing AI. This is a conceptual diagram showing an example of the processing content of the CNN included in the image processing AI. This is a conceptual diagram showing an example of the processing content of the single-channel layer included in the image processing AI. This is a conceptual diagram showing an example of a state in which development processing is performed on a noise-reduced image. This is a flowchart showing an example of the flow of control processing. This is a flowchart showing an example of the flow of noise reduction processing. This is a conceptual diagram showing an example of optimization processing according to the second embodiment. This is a conceptual diagram showing an example of a state in which the volume zone of the error histogram and the error zone adopted in the optimization processing are distinguished. This is a conceptual diagram showing an example of edge components when the example image and the correct answer image are similar, and an example of edge components when the example image and the correct answer image are not similar. This is a conceptual diagram showing an example of how to create example images with different noise levels. This is a conceptual diagram showing an example of noise reduction processing according to the fourth embodiment. This is a conceptual diagram showing an example of the machine learning content performed on the model to obtain the image processing AI shown in Figure 19. This is a conceptual diagram showing an example of the processing content of the multi-channel layer according to the fifth embodiment. This is a conceptual diagram showing an example of the processing content by the multi-channel layer included in the image processing AI according to the fifth embodiment. This is a flowchart showing an example of the noise reduction processing flow according to the fifth embodiment. This is a conceptual diagram showing an example of the processing content of the multi-channel layer according to the sixth embodiment. This is a conceptual diagram showing an example of the machine learning content performed on the model to obtain the image processing AI according to the sixth embodiment. This is a conceptual diagram showing an example of the processing content by the multi-channel layer included in the image processing AI according to the sixth embodiment.This is a conceptual diagram showing an example of the processing content of the multi-channel layer according to the seventh embodiment. This is a conceptual diagram showing an example of the processing content of the multi-channel layer according to the eighth embodiment. This is a schematic configuration diagram showing an example of the configuration of an information processing system.

[0038] Hereinafter, an example of an image processing apparatus, imaging apparatus, image processing system, and program related to this disclosure will be described with reference to the attached drawings.

[0039] First, let's explain the terminology used in the following explanation.

[0040] CPU stands for "Central Processing Unit". GPU stands for "Graphics Processing Unit". GPGPU stands for "General-Purpose computing on Graphics Processing Units". APU stands for "Accelerated Processing Unit". TPU stands for "Tensor Processing Unit". NPU stands for "Neural Processing Unit". DSP stands for "Digital Signal Processor". RAM stands for "Random Access Memory". DRAM stands for "Dynamic Random Access Memory". NVM stands for "Non-volatile memory". ROM stands for "Read Only Memory". EEPROM stands for "Electrically Erasable Programmable Read Only Memory". MRAM stands for "Magnetoresistive Random Access Memory". ReRAM stands for "Resistive Random Access Memory". FRAM (registered trademark) stands for "Ferroelectric Random Access Memory". ASIC stands for "Application Specific Integrated Circuit". FPGA stands for "Field-Programmable Gate Array". CD-ROM stands for "Compact Disc Read Only Memory". DVD-ROM stands for "Digital Versatile Disc Read Only Memory". SSD stands for "Solid State Drive". HDD stands for "Hard Disk Drive". USB stands for "Universal Serial Bus".UI stands for "User Interface". I / F stands for "Interface". AI stands for "Artificial Intelligence". LAN stands for "Local Area Network". WAN stands for "Wide Area Network". 5G stands for "5th Generation Mobile Communication System". CG stands for "Computer Graphics". Exif stands for "Exchangeable Image File Format". CMOS stands for "Complementary Metal-Oxide-Semiconductor". CCD stands for "Charge Coupled Device". JPEG stands for "Joint Photographic Experts Group". BMP stands for "Bitmap". TIFF stands for "Tagged Image File Format". JPEG XR stands for "Joint Photographic Experts Group Extended Range". MPEG stands for "Moving Picture Experts Group". AVI stands for "Audio Video Interleave". CNN stands for "Convolutional Neural Network". MAE stands for "Mean Absolute Error". MSE stands for "Mean Squared Error". SSIM stands for "Structural Similarity Index Measure". FIR stands for "Finite Impulse Response". FFT stands for "Fast Fourier Transform".

[0041] In the following description, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic device, or may be a combination of a plurality of arithmetic devices. Further, the processor may be one type of arithmetic device, or may be a combination of a plurality of types of arithmetic devices. Examples of arithmetic devices include CPU, GPU, GPGPU, NPU, APU, TPU, and DSP.

[0042] In the following description, the labeled RAM is a volatile memory that temporarily stores information and is used as work memory by the processor. An example of RAM includes DRAM.

[0043] In the following description, the labeled NVM is a non-volatile memory that retains stored information even when the power is turned off, and is used for storing programs and data. Examples of NVM include flash memory, EEPROM, ROM, MRAM, ReRAM, and FRAM.

[0044] In the present specification, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Further, in the present specification, when three or more matters are expressed by connecting with "and / or", the same concept as "A and / or B" is applied.

[0045] [First Embodiment] FIG. 1 is a schematic configuration diagram showing an example of the overall configuration of an imaging apparatus. As shown in FIG. 1, the imaging apparatus 10 is an apparatus for imaging a subject, and includes an imaging apparatus main body 12 and an interchangeable lens 14. The imaging apparatus 10 is an example of the "imaging apparatus" according to the present disclosure.

[0046] The interchangeable lens 14 is interchangeably attached to the imaging apparatus main body 12. The interchangeable lens 14 is provided with a focus ring 14A. The focus ring 14A is operated by a user of the imaging apparatus 10 (hereinafter simply referred to as the "user") when the user manually adjusts the focus on a subject by the imaging apparatus 10 in a state where the interchangeable lens 14 is attached to the imaging apparatus main body 12. In the example shown in FIG. 1, an interchangeable-lens digital camera is illustrated as an example of the imaging apparatus 10, but this is merely an example; the imaging apparatus may be a fixed-lens digital camera, or may be a digital camera incorporated in various electronic devices such as a smart device, a wearable terminal, an endoscope system, a cell observation apparatus, an ophthalmic observation apparatus, or a surgical microscope.

[0047] The imaging apparatus main body 12 is provided with an image sensor 16. The image sensor 16 is an example of the "image sensor" according to the present disclosure. The image sensor 16 is a CMOS image sensor. The image sensor 16 images an imaging range including at least one subject. When the interchangeable lens 14 is attached to the imaging apparatus main body 12, subject light representing the subject passes through the interchangeable lens 14, forms an image on the image sensor 16, and is photoelectrically converted.

[0048] In the present embodiment, a CMOS image sensor is exemplified as the image sensor 16, but the present disclosure is not limited thereto. For example, the present disclosure is still applicable even if the image sensor 16 is another type of image sensor such as a CCD image sensor.

[0049] The image sensor 16 includes a photoelectric conversion element 18 and an A / D converter 20. The photoelectric conversion element 18 has a light-receiving surface 18A. The photoelectric conversion element 18 is disposed in the imaging apparatus main body 12 such that the center of the light-receiving surface 18A coincides with the optical axis OA (see also FIG. 1). The photoelectric conversion element 18 has a plurality of photosensitive pixels arranged in a matrix, and the light-receiving surface 18A is formed by the plurality of photosensitive pixels. Each photosensitive pixel has a microlens (not shown).

[0050] Furthermore, each of the multiple photosensitive pixels has a color filter (not shown) of the three primary colors of light, namely red (hereinafter also referred to as "R"), green (hereinafter also referred to as "G"), or blue (hereinafter also referred to as "B"), arranged in a predetermined pattern arrangement. In this first embodiment, the X-Trans® arrangement is used as an example of the predetermined pattern arrangement. Note that the X-Trans arrangement is merely an example, and this disclosure is valid even if the predetermined pattern arrangement is another type of pattern arrangement such as a Bayer arrangement, G-stripe R / G checkerboard, or honeycomb arrangement.

[0051] Each photosensitive pixel is a physical pixel having a photodiode (not shown), which converts received light into photoelectric signals and outputs an electrical signal corresponding to the amount of received light to the A / D converter 20 as analog image data indicating the subject light. For the sake of explanation, below, a photosensitive pixel having a microlens and an R color filter will be referred to as an R pixel, a photosensitive pixel having a microlens and a G color filter will be referred to as a G pixel, and a photosensitive pixel having a microlens and a B color filter will be referred to as a B pixel. For the sake of explanation, below, the electrical signal output from the R pixel of the photosensitive pixel will be referred to as the "R signal," the electrical signal output from the G pixel of the photosensitive pixel will be referred to as the "G signal," and the electrical signal output from the B pixel of the photosensitive pixel will be referred to as the "B signal." For the sake of explanation, below, the R signal, G signal, and B signal will also be referred to as the "RGB color signals."

[0052] The A / D converter 20 reads analog image data from the photoelectric conversion element 18 in one-frame units and horizontal line units using an exposure sequential readout method. The analog image data includes RGB color signals. The A / D converter 20 generates a RAW image 22 by digitizing the analog image data. That is, the RAW image 22 is an image in which red pixels, green pixels, and blue pixels are arranged in a mosaic pattern. The RAW image 22 is an example of the "first image" according to this disclosure.

[0053] For the sake of explanation, in the following, the R pixels, G pixels, and B pixels that make up the RAW image 22 will also be referred to as "R pixels," "G pixels," and "B pixels." Here, R pixels, G pixels, and B pixels are given as examples, but these are merely examples. If colors other than R, G, and B (i.e., colors other than the primary colors) are also regularly arranged together with R, G, and B in the color filter, the RAW image 22 will also include pixels of colors other than R pixels, G pixels, and B pixels, but even in such cases, this disclosure is valid.

[0054] The imaging device body 12 includes an image sensor 16, a processing unit 24, a photoelectric conversion element driver 26, an image memory 28, a UI system device 30, an external I / F 32, a communication I / F 34, and an input / output interface 36. The processing unit 24 includes a system controller 38 and an image processing engine 40. The processing unit 24 is an example of the "image processing device" and "computer" according to this disclosure.

[0055] The input / output interface 36 is connected to an A / D converter 20, a photoelectric conversion element driver 26, an image memory 28, a UI system device 30, an external I / F 32, a communication I / F 34, a system controller 38, and an image processing engine 40.

[0056] The system controller 38 includes a processor 42, an NVM 44, and RAM 46. The processor 42, NVM 44, and RAM 46 are connected via a bus 48, which is connected to an input / output interface 36.

[0057] The NVM 44 is a computer-readable non-temporary storage medium that stores various parameters and programs. These programs include a control program 50. The processor 42 controls the entire imaging device 10 by reading the control program 50 from the NVM 44 and executing it on the RAM 46. For example, the processor 42 controls the image sensor 16, the photoelectric conversion element driver 26, the image memory 28, the UI system 30, the external I / F 32, the communication I / F 34, and the image processing engine 40.

[0058] A photoelectric conversion element driver 26 is connected to the photoelectric conversion element 18. The photoelectric conversion element driver 26 supplies imaging timing signals, which define the timing of imaging performed by the photoelectric conversion element 18, to the photoelectric conversion element 18 according to instructions from the processor 42. The photoelectric conversion element 18 performs reset, exposure, and output of electrical signals according to the imaging timing signals supplied by the photoelectric conversion element driver 26. Examples of imaging timing signals include a vertical synchronization signal and a horizontal synchronization signal.

[0059] The photoelectric conversion element 18, under the control of the photoelectric conversion element driver 26, photoelectrically converts the subject light received by the light-receiving surface 18A and outputs an electrical signal corresponding to the amount of subject light as analog image data representing the subject light to the A / D converter 20. Then, as described above, the A / D converter 20 digitizes the analog image data to generate a RAW image 22. The processor 42 acquires the RAW image 22 from the A / D converter 20 and outputs the acquired RAW image 22 to the image processing engine 40.

[0060] The image processing engine 40 operates under the control of the system controller 38. The image processing engine 40 includes a processor 52, an NVM 54, and RAM 56. The processor 42 of the system controller 38 and the processor 52 of the image processing engine 40 are examples of "processors" according to this disclosure.

[0061] The processor 52, NVM 54, and RAM 56 are connected via a bus 58, which is connected to an input / output interface 36.

[0062] The NVM 54 is a computer-readable non-temporary storage medium that stores various parameters and programs different from those stored in the NVM 44 of the system controller 38. The processor 52 reads the necessary programs from the NVM 54 and executes them in the RAM 56. The processor 52 performs various image processing according to the programs executed on the RAM 56.

[0063] For example, the processor 52 generates an image file 116 by performing image processing, including development, on the RAW image 22 input from the processor 42 of the system controller 38. Here, development refers to color space conversion processing, luminance filtering processing, color difference processing, resizing processing, and compression processing.

[0064] Color space conversion processing refers to the process of converting the color space of an RGB image that has undergone gamma correction processing from the RGB color space to the YCbCr color space. Luminance filtering processing refers to the process of filtering the luminance signal (so-called Y signal) using a luminance filter (not shown in the illustration). Chromatic difference processing refers to the process of filtering the Cb signal and Cr signal to reduce high-frequency noise. Resizing processing refers to the process of adjusting the luminance chromatic difference signal to match the size of the image indicated by the luminance chromatic difference signal to a size specified by the user, etc. Compression processing refers to the process of compressing the luminance chromatic difference signal according to a predetermined compression method. Examples of predetermined compression methods (i.e., image file formats) include JPEG, BMP, TIFF, JPEG XR, MPEG, or AVI.

[0065] Image processing also includes image quality adjustment for the RAW image 22. Image quality adjustment for the RAW image 22 is achieved through processes such as tone correction (for example, a process that corrects the tone of the RGB image according to the gamma value), gain correction, and noise reduction.

[0066] Image processing, including development, is performed to generate an image file 116, which is then stored in the image memory 28 by the processor 52.

[0067] The UI system device 30 is equipped with a display 30A, and the processor 42 displays various information on the display 30A. The UI system device 30 is also equipped with a reception device 30B. The reception device 30B is equipped with a touch panel and multiple hard keys, etc. The processing unit 24 operates according to the various instructions received by the reception device 30B.

[0068] The external I / F 32 is responsible for the exchange of various types of information between the imaging device 10 and devices located outside of it (hereinafter also referred to as "external devices"). An example of the external I / F 32 is a USB interface. External devices such as smart devices, personal computers, servers, USB memory, memory cards, and / or printers (not shown) can be directly or indirectly connected to the USB interface.

[0069] The communication interface 34 is connected to a network (not shown). The communication interface 34 is responsible for the exchange of information between communication devices (not shown), such as servers, on the network and the processing unit 24.

[0070] The communication interface 34 includes a communication processor and an antenna, and manages communication between multiple computers over a network. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G, Wi-Fi®, or Bluetooth®. The communication interface 34 transmits information requested by the system controller 38 to communication devices (e.g., servers, personal computers, and / or smart devices) over the network. The communication interface 34 also receives information transmitted from communication devices and outputs the received information to the system controller 38 via the input / output interface 36.

[0071] Incidentally, as shown in Figure 2 as an example, the RAW image 22 generated by imaging by the imaging device 10 contains noise. It is generally known that increasing the ISO sensitivity of the imaging device 10 increases the amount of noise in the RAW image 22.

[0072] Because the RAW image 22 has different colors linked in the planar direction of one channel, it is difficult to distinguish between noise components and edge components compared to a three-channel image such as a developed JPEG or BMP. Therefore, conventional known techniques have a lower noise reduction effect on the RAW image 22 than on three-channel images such as JPEG or BMP, in terms of reducing the amount of noise while preserving the image structure. Furthermore, when having AI determine whether something is noise or an edge component, the more complex the color filter array, the more difficult it becomes to train the AI ​​to make this determination.

[0073] Therefore, in light of these circumstances, in the image processing engine 40 according to this first embodiment, as an example shown in Figure 3, noise reduction processing is performed by the processor 52. The NVM 54 stores a noise reduction processing program 66. The processor 52 reads the noise reduction processing program 66 from the NVM 54 and executes the read noise reduction processing program 66 on the RAM 56. Noise reduction processing is achieved by the processor 52 executing the noise reduction processing program 66.

[0074] The NVM 54 stores the image processing AI 68. The image processing AI 68 is used by the processor 52 in noise reduction processing. For example, the processor 52 uses the image processing AI 68 to obtain a noise-reduced image 70. That is, the processor 52 inputs the RAW image 22 containing noise to the image processing AI 68, causing the image processing AI 68 to generate a noise-reduced image 70. The noise-reduced image 70 is a RAW image with reduced noise compared to the RAW image 22. In other words, the noise-reduced image 70 is an image obtained by reducing the noise in the RAW image 22.

[0075] The control program 50 (see Figure 1) and the noise reduction processing program 66 are examples of the "programs" related to this disclosure. The image processing AI 68 is an example of the "AI" related to this disclosure. The noise-reduced image 70 is an example of the "second image" related to this disclosure.

[0076] Next, an example of the first machine learning process performed to obtain the image processing AI 68 will be explained with reference to Figures 4 to 8.

[0077] As an example, as shown in Figure 4, the first machine learning process for obtaining the image processing AI 68 is performed by the learning execution device 72. The learning execution device 72 is implemented by a computer having a processor 74, an NVM 76, and RAM (not shown), etc. Since the hardware configuration of the learning execution device 72 is basically the same as the hardware configurations of the system controller 38 and the image processing engine 40, a description of the hardware configuration of the learning execution device 72 is omitted here.

[0078] The NVM76 stores multiple training data sets 78. The training data set 78 consists of example images 80 and correct answer images 82. Both example images 80 and correct answer images 82 are RAW images.

[0079] Example image 80 is an image corresponding to the first captured image obtained by performing the first imaging. The first captured image has the same X-Trans array as the RAW image 22 and is a RAW image with the same resolution as the RAW image 22. Since the first captured image has an X-Trans array on a single plane, it can be said to be a single-channel image in terms of data format.

[0080] An example of an image corresponding to the first captured image is an image that simulates the first captured image. For example, an image that simulates the first captured image is generated by a conventionally known generative AI. Multiple example images 80 included in multiple training data 78 are images for each different sensitivity used in the first image capture. For example, different sensitivities refer to ISO 200, ISO 6400, and ISO 12800, etc.

[0081] Furthermore, among the multiple example images 80 included in the multiple training data 78, there is at least one image that represents high-frequency components that cannot be fully covered by real-world images alone, i.e., frequencies in the Nyquist limit region. The at least one image that represents frequencies in the Nyquist limit region is a virtual image generated by computer graphics technology.

[0082] Here, in order to facilitate understanding of this disclosure, an example is given in which an image corresponding to the first captured image is used as example image 80. However, this disclosure is not limited to this, and the first captured image may be used as example image 80. Furthermore, the multiple example images 80 included in the multiple training data 78 may contain a mixture of the first captured image and images corresponding to the first captured image.

[0083] On the other hand, the ground truth image 82 is the image corresponding to the second image obtained by performing a second image capture, which has less noise than the first image capture. The second image capture has the same X-Trans array as the RAW image 22 and is a RAW image with the same resolution as the RAW image 22. Since the second image capture has an X-Trans array on a single plane, it can be said to be a single-channel image in terms of data format.

[0084] An example of an image corresponding to the second captured image is an image that simulates the second captured image. For example, an image that simulates the second captured image can be generated by a conventionally known generative AI.

[0085] Here, in order to facilitate understanding of this disclosure, an example is given in which the image corresponding to the second captured image is used as the ground truth image 82. However, this disclosure is not limited to this, and the second captured image may also be used as the ground truth image 82. Furthermore, the multiple ground truth images 82 included in the multiple training data 78 may contain a mixture of the second captured image and images corresponding to the second captured image.

[0086] In this first embodiment, the first machine learning is an example of the "first learning" as described in this disclosure. Also, in this first embodiment, the first captured image is an example of the "first captured image" and "third captured image" as described in this disclosure. Also, in this first embodiment, the second captured image is an example of the "second captured image" and "fourth captured image" as described in this disclosure. Also, in this first embodiment, the first imaging is an example of the "first imaging" and "third imaging" as described in this disclosure. Also, in this first embodiment, the second imaging is an example of the "second imaging" and "fourth imaging" as described in this disclosure. Also, in this first embodiment, the example image 80 is an example of the "third image," "first captured image," "image corresponding to the first captured image," and "image corresponding to the third captured image" as described in this disclosure. Also, in this first embodiment, the correct answer image 82 is an example of the "fourth image," "image corresponding to the second captured image," "image corresponding to the fourth captured image," and "second correct answer image" as described in this disclosure.

[0087] Assuming that multiple training data sets 78 configured in this manner are stored in the NVM 76, the processor 74 in the learning execution device 72 acquires the training data 78 from the NVM 76. The processor 74 then performs a first machine learning operation using the training data 78.

[0088] In this case, for example, the processor 74 generates an image processing AI 68 by optimizing the model 84 using backpropagation based on multiple training data 78.

[0089] For example, Model 84 is a model that includes a neural network. Specifically, as shown in Figure 5 as an example, Model 84 is a model that includes a multi-channel layer 86, a CNN 88, and a single-channel layer 90.

[0090] As an example, as shown in Figure 6, the multi-channel layer 86 is a layer that multi-channels the example image 80 and the correct answer image 82 by dividing them into predetermined units. Here, the predetermined unit refers to, for example, a RAW array rule. The RAW array rule can also be described as a unit determined based on periodic features assigned to multiple pixels representing a RAW image. An example of a RAW array rule is a rule that includes the periodic unit of a color filter applied to multiple pixels. A rule that includes the periodic unit of a color filter refers to, for example, a repeating pattern of 6x6 units in the case of an X-Trans array according to this first embodiment, a repeating pattern of 2x2 units in the case of a Bayer array, and if functional pixels such as phase difference pixels or infrared pixels are periodically mixed in the image sensor 16, it refers to an array rule that includes the period in which the functional pixels are mixed.

[0091] The example image 80 and the correct answer image 82 are both single-channel images with an X-Trans array. Therefore, the multi-channel layer 86 creates multiple divided images 80A by dividing the example image 80 according to the repeating units of the X-Trans array, and also creates multiple divided images 82A by dividing the correct answer image 82 according to the repeating units of the X-Trans array. In other words, in the case of an X-Trans array, multiple divided images 80A and 82A are created by setting 36 pixels of a 6x6 block 92, which is the smallest block unit of the X-Trans array, as one set. Each of the multiple divided images 80A can be said to be a single-channel image in which the example image 80 has been converted into a single-channel image according to the repeating units of the X-Trans array. Similarly, each of the multiple divided images 82A can be said to be a single-channel image in which the correct answer image 82 has been converted into a single-channel image according to the repeating units of the X-Trans array. In the example shown in Figure 6, a multi-channel image 94 is shown in which multiple segmented images 80A are bundled in the channel direction, which is along the channel axis, and a multi-channel image 96 is shown in which multiple segmented images 82A are bundled in the channel direction.

[0092] To explain the multi-channel layer 86 in more detail, the multi-channel layer 86 rearranges the multiple pixels P1 that make up the example image 80 from the spatial direction to the channel direction within each block 92. Similarly, the multi-channel layer 86 rearranges the multiple pixels P2 that make up the correct image 82 from the spatial direction to the channel direction within each block 92. Specifically, calculations such as SpaceToDepth or PixelSuffle are performed so that the example image 80 and the correct image 82 (height H, width V, 1 channel) are rearranged into multi-channel images 94 and 96 (height H / 6, width V / 6, 36 channels). In this way, the multi-channel layer 86 separates the color pixels mosaicked in the X-Trans array and performs preprocessing to make them easier for the subsequent CNN 88 to handle.

[0093] As an example, as shown in Figure 7, the multi-channel layer 86 inputs the multi-channel image 94 to the CNN 88. The CNN 88 has multiple stages of convolutional layers, activation layers, and pooling layers. When the multi-channel image 94 is input to the CNN 88, it performs inference by executing a convolution operation.

[0094] The initial kernel size used in the convolution operation is a size predetermined according to the repeating unit of the X-Trans array, which is a size that can maintain the color periodicity of the example image 80 in the convolution operation. The initial kernel size refers to the kernel size used for the resolution transformation of the first layer used in the convolution operation during the initial training stage of CNN88. An example of an initial kernel size is a size that is a multiple of 3. In this case, at the stage of resolution transformation of the second and subsequent layers used in the convolution operation, that is, when the number of pixels in the convolution operation no longer depends on the X-Trans array, the kernel size is changed to a size that is a multiple of 2. Note that although the kernel size used in the convolution operation is explained here, the same can be said for the prieling size.

[0095] CNN88 outputs an inference result 100 generated by performing a convolution operation. Here, the processor 74 performs an optimization process. In the optimization process, first, the error 102 between the inference result 100 and the multi-channel image 96 is calculated. The error 102 is a value evaluated by MAE.

[0096] In the optimization process, several adjustment values ​​104 are calculated to minimize the error 102. Then, several learning parameters within CNN88 are adjusted by these adjustment values ​​104. Here, the several learning parameters refer to the weights and biases included in CNN88.

[0097] A series of processes—inputting the multi-channel image 94 into the CNN 88, calculating the error 102, calculating multiple adjustment values ​​104, and adjusting multiple learning parameters within the CNN 88—are repeatedly performed using all the multi-channel images 94 and 96 obtained by the multi-channelization layer 86. This adjusts the learning parameters so that noise in the multi-channel image 94 is effectively removed, thereby optimizing the CNN 88. The image processing AI 68 is obtained by optimizing the CNN 88 in this way.

[0098] For example, when a multi-channel image 94 is input to the optimized CNN 88, the CNN 88 captures and corrects the noise characteristics corresponding to each pixel P1 in block 92 (see Figure 6), thereby reducing the noise. The inference result 100 obtained by inputting the multi-channel image 94 to the optimized CNN 88 is a multi-channel image with (height H / 6, width V / 6, 36 channels).

[0099] Therefore, as shown in Figure 8 as an example, the single-channelization layer 90 performs an inverse transformation process such as SpaceToDepth or PixelSuffl on the inference result 100 to convert the multi-channel image (height H / 6, width V / 6, 36 channels) back into a single-channel image 106 (height H, width V, 1 channel). Specifically, the single-channelization layer 90 expands the pixels P1 that were separated in units of each block 92 (see Figure 6) in the inference result 100 again in the spatial direction and restores them to the original two-dimensional array of the RAW image (i.e., the X-Trans array) to generate a single-channel image 106, which is a single RAW image.

[0100] The final output obtained by the single-channel layer 90 (height H, width V, 1 channel) is a single-channel image 106 that has the same resolution mosaic arrangement as the original RAW image, i.e., the example image 80 (see Figures 4 and 6), but with the noise reduction effect of CNN88 applied. As a result, even when development processing and / or blending processing with the original RAW image are performed in a later stage, a RAW image with reduced noise can be obtained compared to the original RAW image that has not been processed using the optimized CNN88.

[0101] In this first embodiment, the multi-channel image 94 is an example of the "multiple first example images" related to this disclosure. Also, in this first embodiment, the multi-channel image 96 is an example of the "multiple first correct answer images" related to this disclosure. The error 102 is an example of the "first degree of difference" related to this disclosure, and MAE is an example of the "first index" related to this disclosure. Here, MAE is used as an example, but the error 102 may be a value evaluated by MSE. Also, here, the error 102 is used as an example, but it may be a percentage, and this disclosure is valid as long as the degree of difference is the degree to which the inference result 100 and the multi-channel image 96 differ.

[0102] Next, an example of the noise reduction process using the image processing AI 86 generated by optimizing model 84 through the first machine learning method described above will be explained with reference to Figures 9 to 12.

[0103] As an example, as shown in Figure 9, the processor 52 inputs the RAW image 22 to the multi-channelization layer 86 of the image processing AI 68. Since the RAW image 22 is a single-channel image with an X-Trans array applied to a single plane, the multi-channelization layer 86 creates multiple divided images 22A by dividing the RAW image 22 according to the repeating units of the X-Trans array. That is, the multi-channelization layer 86 creates multiple divided images 22A by setting 36 pixels of block 92 as one set. Each of the multiple divided images 22A can be said to be a single-channel image in which the RAW image 22 has been converted into a single channel for each repeating unit of the X-Trans array. In the example shown in Figure 9, a multi-channel image 110 is shown in which multiple divided images 22A are bundled in the channel direction.

[0104] More specifically, the multi-channel layer 86 rearranges the multiple pixels P3 that make up the RAW image 22 from the space direction to the channel direction within each block 92. Specifically, calculations such as SpaceToDepth or PixelSuffl are performed to rearrange the RAW image 22 (height H, width V, 1 channel) into a multi-channel image 110 (height H / 6, width V / 6, 36 channels).

[0105] In this embodiment, the repeating unit of the X-Trans sequence is an example of the "predetermined unit," "RAW sequence rule," and "periodic unit of the color filter" as described in this disclosure. Also, in this first embodiment, the multiple divided images 22A are an example of the "multiple first divided images" as described in this disclosure.

[0106] As an example, as shown in Figure 10, the multi-channel layer 86 inputs the multi-channel image 110 to the CNN 88. In response, the CNN 88 generates multiple segmented images 22B in which the noise of the multiple segmented images 22A has been reduced. That is, the CNN 88 reduces noise by capturing and correcting the noise characteristics of the multiple segmented images 22A according to each pixel P3 (see Figure 9) in the block 92. In the example shown in Figure 10, a multi-channel image 112 is shown in which multiple segmented images 22B are bundled in the channel direction.

[0107] As an example, as shown in Figure 11, a multi-channel image 112 is input to the single-channel conversion layer 90. The single-channel conversion layer 90 converts the input multi-channel image 112 into a single channel based on the X-Trans array, thereby generating a noise-reduced image 70, which is a single image with reduced noise compared to the RAW image 22. In other words, the single-channel conversion layer 90 performs an inverse transformation process such as SpaceToDepth or PixelSuffl on the input multi-channel image 112, converting a multi-channel image (height H / 6, width V / 6, 36 channels) back into a single RAW image (height H, width V, 1 channel). Specifically, the single-channel layer 90 unfolds the pixels that were separated in units of each block 92 (see Figure 9) in the multi-channel image 112 in the spatial direction, and restores them to the original two-dimensional array (i.e., X-Trans array) of the RAW image. This generates a noise-reduced image 70 that has the same resolution mosaic array as the original RAW image, but with the noise reduction effect of CNN 88 applied. In the image processing AI 68, the noise-reduced image 70 is the final output (height H, width V, 1 channel). In this embodiment, the noise-reduced image 70 is an example of the "second image" according to this disclosure.

[0108] As an example, as shown in Figure 12, the processor 52 acquires a noise-reduced image 70 from the image processing AI 68. The processor 52 then generates an image file 116 by developing the noise-reduced image 70. The image file 116 contains a developed image 116A obtained by developing the noise-reduced image 70, and metadata 116B related to the developed image 116A. The metadata 116B is obtained in conjunction with imaging by the imaging device 10 to obtain the RAW image 22 that will serve as the basis for the developed image 116A. An example of metadata 116B is data in Exif format.

[0109] The metadata 116B includes unit identification information 116B1 that can identify a predetermined unit when the RAW image 22 that forms the basis of the developed image 116A is divided into predetermined units (here, as an example, the repeating unit of the X-Trans array), the time when the developed image 116A was completed, the time when the image was taken to obtain the developed image 116A, the location where the image was taken, the name of the person who took the image, and various image processing parameters that can be used for image processing, etc. (for example, shutter speed, ISO sensitivity, and F-number).

[0110] In this first embodiment, the multiple divided images 22B are an example of the "multiple second divided images" relating to this disclosure. Also, in this first embodiment, the unit identification information 116B1 is an example of the "unit identification information" relating to this disclosure. Also, in this first embodiment, the metadata 116B is an example of the "metadata" relating to this disclosure.

[0111] Next, an example of the operation of the part of the imaging device 10 related to this disclosure will be described with reference to Figures 13 and 14.

[0112] Figure 13 shows an example of the flow of control processing performed by the processor 42 of the system controller 38. The control processing shown in Figure 13 is realized by the execution of the control program 50 by the processor 42. Figure 14 shows an example of the flow of noise reduction processing performed by the processor 52 of the image processing engine 40.

[0113] The control process shown in Figure 13 and the noise reduction process shown in Figure 14 are examples of the "processing that includes outputting at least one of a first image which is a RAW image and a plurality of first divided images obtained by dividing the first image based on predetermined units, and obtaining a second image in which noise is reduced compared to the first image based on the plurality of first divided images" according to this disclosure.

[0114] In the control process shown in Figure 13, first, in step ST10, the processor 42 determines whether the conditions for acquiring a RAW image 22 have been met. An example of a condition for acquiring a RAW image 22 is that the image sensor 16 has completed capturing one frame. If the conditions for acquiring a RAW image 22 are not met in step ST10, the determination is denied, and the control process proceeds to step ST16. If the conditions for acquiring a RAW image 22 are met in step ST10, the determination is affirmed, and the control process proceeds to step ST12.

[0115] In step ST12, the processor 42 acquires a RAW image 22 from the image sensor 16. After the processing in step ST12 is completed, the control process moves on to step ST14.

[0116] In step ST14, the processor 42 outputs the RAW image 22 acquired in step ST12 to the image processing engine 40. After the processing in step ST14 is completed, the control process moves on to step ST16.

[0117] In step ST16, the processor 42 determines whether the conditions for terminating the control process have been met. One example of a condition for terminating the control process is that an instruction to terminate the control process has been received by the receiving device 30B. If the conditions for terminating the control process are not met in step ST16, the determination is denied, and the control process proceeds to step ST10. If the conditions for terminating the control process are met in step ST16, the determination is affirmed, and the control process terminates.

[0118] In the noise reduction process shown in Figure 14, first, in step ST20, the processor 52 determines whether the RAW image 22 output from the processor 42 as a result of the execution of the process in step ST14, which is included in the control process, has been input to the image processing engine 40. If the RAW image 22 has not been input to the image processing engine 40 in step ST20, the determination is denied, and the noise reduction process proceeds to step ST26. If the RAW image 22 has been input to the image processing engine 40 in step ST20, the determination is affirmed, and the noise reduction process proceeds to step ST22.

[0119] In step ST22, the processor 52 inputs the RAW image 22 to the image processing AI 68. The image processing AI 68 first multi-channels the RAW image 22 using the multi-channelization layer 86, generating a multi-channel image 110. Next, the multi-channel image 110 is input to the CNN 88. The CNN 88 generates a multi-channel image 112 with reduced noise compared to the multi-channel image 110. Then, the multi-channel image 112 is input to the single-channelization layer 90. This results in a noise-reduced image 70, which has the same resolution as the RAW image 22 but with the noise reduction effect applied by the CNN 88, generated by single-channelizing the multi-channel image 112 based on the repeating units of the X-Trans array. After the processing in step ST22 is completed, the noise reduction process proceeds to step ST24.

[0120] In step ST24, the processor 52 acquires the noise-reduced image 70 from the image processing AI 68 and generates an image file 116 by developing the noise-reduced image 70. The processor 52 then outputs the image file 116 to a default output destination. An example of a default output destination is the image memory 28. The image file 116 in the image memory 28 is read by the system controller 38 and output to an external device connected to the external I / F 32 (e.g., a USB memory and / or memory card, etc.) or transmitted to a communication device (e.g., a server, personal computer, and / or smart device, etc.) that is communicatively connected to the communication I / F 34. After the processing in step ST24 is executed, the noise reduction process moves on to step ST26.

[0121] In step ST26, the processor 52 determines whether the conditions for terminating the noise reduction process have been met. One example of a condition for terminating the noise reduction process is that an instruction to terminate the noise reduction process has been received by the receiving device 30B. If the conditions for terminating the noise reduction process are not met in step ST26, the determination is denied, and the noise reduction process proceeds to step ST20. If the conditions for terminating the noise reduction process are met in step ST26, the determination is affirmed, and the noise reduction process terminates.

[0122] As explained above, in the imaging device 10, the RAW image 22 is output to the image processing engine 40 by the processor 42 of the system controller 38 (see Figure 1). Then, the processor 52 of the image processing engine 40 acquires a noise-reduced image 70, which is an image with reduced noise compared to the RAW image 22, based on the multi-channel image 110 (see Figures 9 to 12). In this way, at the RAW stage before development, a lot of raw information necessary for noise reduction can be utilized compared to when noise reduction is performed on an image after development, so an image with reduced noise can be obtained compared to when noise reduction processing is performed on an image after development.

[0123] Furthermore, the multi-channel image 110 allows for pinpoint analysis of noise in the R channel only, noise in the G channel only, and noise in the B channel only. Therefore, it becomes easier to distinguish between noise components and edge components compared to a single-channel image. In other words, using the multi-channel image 110 allows for the separation and effective utilization of color information compared to treating the RAW image 22 as a single channel. Consequently, noise can be reduced while maintaining fine image structure and color tones, compared to reducing noise in a single-channel RAW image 22.

[0124] Furthermore, in the imaging device 10, the acquisition of the noise-reduced image 70 by the processor 52 is achieved when the RAW image 22 is input to the image processing AI 68, which generates a multi-channel image 110, and then generates the noise-reduced image 70 based on the multi-channel image 110. The CNN 88 in the image processing AI 68 has learned the complex color and noise characteristics in the multi-channel image, which were difficult to achieve with conventional simple linear filters or fixed algorithms. That is, because the CNN 88 in the image processing AI 68 includes nonlinear activation such as convolutional layers and activation layers (e.g., ReLU), it can realize noise reduction characteristics and complex effects obtained through learning that cannot be expressed by linear filters or fixed algorithms. Therefore, by using the image processing AI 68, the imaging device 10 can generate a noise-reduced image 70 that effectively suppresses noise while maintaining fine image structure and color, compared to when using conventional simple linear filters or fixed algorithms.

[0125] Furthermore, the multi-channel image 110 is obtained by bundling multiple divided images 22A, which are obtained by dividing the RAW image 22 according to the repeating units of the X-Trans array, in the channel direction. The noise-reduced image 70 is then generated based on the multi-channel image 110. Therefore, compared to the case where the RAW image 22 is divided independently of the repeating units of the X-Trans array, color shifts and / or false colors are less likely to occur at the periodic boundaries of the X-Trans array, so noise can be reduced without disrupting the color balance. As a result, an image with reduced noise can be obtained while maintaining the color reproducibility of the image. Note that the same effect can be obtained even when the RAW image 22 is divided according to the periodic units of a color filter with an array different from the X-Trans array, such as a Bayer array.

[0126] Furthermore, each of the example images 80 includes at least one image that represents frequencies in the Nyquist limit region. By using images containing the Nyquist limit region for training, it is possible to achieve noise reduction while maintaining a high-resolution image structure compared to using only training data that does not consider the Nyquist limit region.

[0127] Furthermore, the image processing AI 68 is a pre-trained model that has undergone first-stage machine learning to reduce noise in multiple segmented images 80A obtained by dividing the example image 80, which is a RAW image, into segments of the X-Trans array. Thus, since the image processing AI 68 has already undergone first-stage machine learning and possesses the ability to reduce noise by dividing RAW images into segments of the X-Trans array, it achieves higher accuracy in noise reduction at the RAW stage compared to AIs that have not undergone such training. As a result, image quality degradation is suppressed even after subsequent development processing, compared to when the pre-trained model is used as is. Consequently, it is possible to obtain images with reduced noise compared to when the pre-trained model is used as is.

[0128] Furthermore, the first machine learning performed to obtain the image processing AI 68 is a learning process using multiple segmented images 80A obtained by dividing the example image 80 into repeating units of the X-Trans array, and multiple segmented images 82A obtained by dividing the ground truth image 82 into repeating units of the X-Trans array. The example image 80 corresponds to the first captured image obtained by performing the first imaging, and the ground truth image 82 corresponds to the second captured image obtained by performing the second imaging, which has less noise than the first imaging. Therefore, the image processing AI 68 can acquire the ability to reduce the amount of noise contained in the example image 80 to the amount of noise contained in the ground truth image 82.

[0129] Furthermore, the first machine learning process performed to obtain the image processing AI 68 uses multiple segmented images 80A obtained by dividing each of the multiple example images 80 according to the repeating units of the X-Trans array. The multiple example images 80 include multiple images obtained by performing the first imaging at different sensitivities. Therefore, the image processing AI 68 can acquire the ability to distinguish the noise characteristics for each different sensitivity. As a result, the image processing AI 68 can optimally remove noise according to a wide range of imaging conditions compared to when it is trained on only a single sensitivity. Consequently, the image processing AI 68 can obtain a stable noise reduction effect even in imaging environments with various sensitivities compared to when it is trained on only a single sensitivity.

[0130] Furthermore, the image processing AI 68 is a pre-trained model that requires convolution operations. The initial kernel size used in the convolution operation is a size predetermined according to the repeating unit of the X-Trans array, so as to be able to maintain the color periodicity of the example image 80, i.e., the 6x6 repeating unit of the X-Trans array, during the convolution operation. Therefore, compared to a case where the initial kernel size is a fixed size determined independently of the color periodicity of the example image 80, color distortion due to the convolution operation can be suppressed. This makes it possible to achieve convolution along the period of the RAW array, and noise can be effectively reduced while maintaining color reproduction. In addition, since the initial kernel size is defined as a size that can maintain the color periodicity of the example image 80, i.e., the 6x6 repeating unit of the X-Trans array, the first machine learning can be expected to converge more stably and quickly.

[0131] Furthermore, the image processing AI 68 generates a noise-reduced image 70 by converting the multi-channel image 112 into a single-channel image based on the repeating units of the X-Trans array. In other words, by converting the multi-channel image 112 into a single channel image, the color information of each divided image 22B is consolidated into a single-color plane. This makes it possible to handle the same-color plane consistently compared to when the image is left as a multi-channel image, making the noise reduction results easier to handle.

[0132] In the first embodiment described above, an example was given in which the RAW image 22 is multi-channelized by the multi-channelization layer 86 of the image processing AI 68, but this is merely one example. For example, the RAW image 22 may be multi-channelized outside of the image processing AI 68. In this case, for example, the RAW image 22 may be multi-channelized by calculations such as SpaceToDepth or PixelSuffle performed by the processor 42 of the system controller 38, or the same processing may be performed by a processor different from the processor 42 of the system controller 38.

[0133] In the first embodiment described above, an example was given in which a multi-channel image 112 is converted to a single channel by the single-channel conversion layer 90 of the image processing AI 68, but this is merely one example. For example, the multi-channel image 112 may be converted to a single channel outside of the image processing AI 68. In this case, for example, the RAW image 22 may be converted to a single channel by an inverse transformation process such as SpaceToDepth or PixelSuffl performed by the processor 42 of the system controller 38, or the same process may be performed by a processor different from the processor 42 of the system controller 38.

[0134] In the first embodiment described above, an example was given in which a noise-reduced image 70 is generated by image processing AI 68, but this disclosure is not limited thereto. For example, an image equivalent to the noise-reduced image 70 may be generated by performing non-AI processing on a multi-channel image 110. An example of non-AI processing is processing using a linear filter. The linear filter may be a two-dimensional, channel-dimension filter that includes inter-channel coupling coefficients that take into account the correlation between channels and the intrinsic noise characteristics of each channel. In this case, for example, even if the multi-channel image 110 has 6 x 6 = 36 channels, spatial convolution (i.e., FIR) that takes into account the connections between channels may be performed. Alternatively, filtering processing using a frequency-domain linear filter may be performed on each of the 36 channels, or on the entire 36-dimensional frequency space. In this case, for example, each channel may be converted to the frequency domain by FFT, high frequencies may be cut with a low-pass filter, and then converted back to space by inverse FFT.

[0135] In the first embodiment described above, an example configuration was given in which the example image 80 and the correct answer image 82 are stored in the NVM 76. However, this is merely one example, and the NVM 76 only needs to store multi-channel images 94 and / or 96. In this case, multi-channelization of the RAW image 22 by a multi-channelization layer 86, etc., becomes unnecessary.

[0136] In the first embodiment described above, an example was given in which a RAW image 22 is output from processor 42 to processor 52, but the disclosure is not limited thereto. For example, if a processor is mounted on the image sensor 16, the RAW image 22 may be output directly or indirectly from the processor mounted on the image sensor 16 to processor 52. An example of an image sensor 16 mounted on a processor is a stacked image sensor having a structure in which a pixel layer having a photoelectric conversion element and a signal processing layer having a processor are stacked.

[0137] [Second Embodiment] In the first embodiment described above, a first machine learning method was illustrated in which the CNN 88 is optimized by backpropagation based on the error 102 between the inference result 100 and the multi-channel image 96. However, the disclosure is not limited thereto. For example, in addition to adjusting the learning parameters of the CNN 88 by backpropagation based on the error 102, the learning parameters of the CNN 88 may also be adjusted by backpropagation based on the error between the single-channel image of the inference result 100 and the ground truth image 82, which is a single-channel image.

[0138] In this case, for example, in addition to adjusting the learning parameters of CNN 88 by backpropagation based on the error 102 described in the first embodiment above, a second machine learning process is performed in which CNN 88 is optimized by backpropagation based on the error 118, as shown in Figure 15. In the second machine learning process, the learning parameters of CNN 88 are adjusted by backpropagation based on the error 118. As shown in Figure 15, in the second machine learning process, the processor 74 generates a single-channel image 120 by single-channelizing the inference result 100, which can be described as multiple inference images (here, as an example, a 36-channel inference image) obtained by inference by CNN 88. The single-channel image 120 is an image having an X-Trans array with the same resolution as the example image 80 and the correct answer image 82.

[0139] In the second machine learning step, the processor 74 calculates the error 118 between the single-channel image 120 and the ground truth image 82. The error 118 is evaluated using SSIM, which is a different metric from MAE used in the evaluation of the error 102. In the learning process of CNN88, MAE is a metric that is more effective than SSIM at reducing image noise, while SSIM is a metric that is more effective at preserving the image structure of the image than MAE.

[0140] The processor 74 calculates multiple adjustment values ​​122 that minimize the error 118. Then, the processor 74 adjusts the learning parameters of the CNN 88 using the multiple adjustment values ​​122. This series of processes—inputting the multi-channel image 94 into the CNN 88, calculating the error 118, calculating the multiple adjustment values ​​122, and adjusting the multiple learning parameters within the CNN 88—is repeated for all the training data 78.

[0141] In this second embodiment, the inference result 100 is an example of the "multiple inference images" related to this disclosure. Also, in this second embodiment, the second machine learning is an example of the "second learning" related to this disclosure. Also, in this second embodiment, the single-channel image 120 is an example of the "image corresponding to the third image" and the "single image in which multiple inference images have been converted into a single channel based on units" related to this disclosure. Also, in this second embodiment, the error 118 is an example of the "second difference" related to this disclosure. Also, in this second embodiment, SSIM is an example of the "second index" related to this disclosure. Here, the error 118 is given as an example, but it may also be a percentage, and this disclosure is valid as long as it is a difference that represents the degree to which the ground truth image 82 and the single-channel image 120 differ.

[0142] As described above, in this second embodiment, error backpropagation using the error 102 is performed not only in the first machine learning but also in the second machine learning, so that the degree of difference between the example image 80 and the ground truth image 82 is learned by the CNN 88. As a result, the image processing AI 68 can be updated to a trained model that can handle imaging conditions different from those of the first machine learning. As a result, the performance in reducing noise by handling an even wider range of ISO sensitivity differences is improved compared to when only the first machine learning is performed.

[0143] Furthermore, in this second embodiment, the multi-channel image 94 is input to the CNN 88 in the learning stage (in other words, the CNN 88 before the optimization process is completed), and the inference results 100, which can be said to be multiple inference images obtained by the CNN 88, are converted into single channels based on the repeating units of the X-Trans array. As a result, the single-channel image 120 obtained by converting the inference results 100 into single channels can be used for machine learning different from the first machine learning (for example, the second machine learning). As a result, it is possible to construct an image processing AI 68 with higher accuracy compared to when a single-channel image 120 is not obtained.

[0144] Furthermore, in this second embodiment, in addition to adjusting the learning parameters of CNN 88 by backpropagation based on error 102, adjustment of the learning parameters of CNN 88 is also performed by backpropagation based on error 118. Therefore, as shown in Figure 15, the image processing AI 68 including the optimized CNN 88 can achieve a higher level of balance between noise reduction and image structure preservation compared to the case where only the learning parameters of CNN 88 are adjusted by backpropagation based on error 102. In addition, since noise reduction is performed in the second machine learning in addition to the first machine learning, the image processing AI 68 can handle a wider variety of noise patterns and conditions compared to an AI that has only performed a single machine learning process. As a result, the image quality improvement performance, including noise reduction at the inference stage of the RAW image 22, can be further enhanced compared to an AI that has only performed a single machine learning process.

[0145] Furthermore, in this second embodiment, since the error 102 is evaluated by MAE, it is effective in significantly reducing noise, and since the error 118 is evaluated by SSIM, it is effective in preserving the image structure. Therefore, as the CNN 88 is trained to remove a large amount of noise with error 102 and preserve the image structure with error 118, a noise-reduced image 70 can be obtained that achieves both noise reduction and suppression of image structure collapse, compared to the case where backpropagation is performed based only on an error evaluated by only a single metric.

[0146] In the second embodiment described above, backpropagation based on the error 118 between the single-channel image 120 and the correct image 82 was illustrated, but the disclosure is not limited thereto. For example, backpropagation based on the error between the example image 80 and the correct image 82 may be performed, and the same effects as in the second embodiment can be obtained by doing so. In this case, the example image 80 is an example of the "third image" according to the disclosure.

[0147] In the second embodiment described above, all errors 102 evaluated by MAE for each pixel are used in the first machine learning, but the disclosure is not limited thereto. For example, as shown in Figure 16, errors 102 from zones other than the volume zone in the histogram of errors 102 for each pixel may be used in the first machine learning. The volume zone may be a fixed zone or a zone that can be changed according to instructions given from an external source.

[0148] In this way, by using the error 102 from a zone different from the volume zone in the histogram of the error 102 for each pixel in the first machine learning, MAE, which evaluates areas with large errors outside the volume zone (i.e., areas with many small errors), can preferentially reduce noise in areas with large errors compared to when all pixels are treated uniformly with MAE. Although MAE is used as an example here, a similar effect can be obtained with MSE.

[0149] [Third Embodiment] As described in each of the embodiments above, when learning noise reduction using the image processing AI 68, it is necessary to use a pair of example images 80 containing noise and a correct image 82 with less noise than the example image 80 for learning. What is important here is the noise characteristics that the ISO sensitivity imparts to the example image 80 and the correct image 82. For example, if an image processing AI 68 that has been trained to remove noise at a relatively high sensitivity such as ISO 12800 is applied to a relatively low sensitivity image such as ISO 3200, there is a high possibility that not only the noise but also the edge components of the signal will be greatly impaired.

[0150] Specifically, if training is performed using images with extreme differences in sensitivity, such as when the ISO sensitivity used for the first image acquisition to obtain example image 80 is ISO 12800 and the ISO sensitivity used for the second image acquisition to obtain the correct image 82 is ISO 200, there will be a significant difference in the image structure of the two images. As a result, during the training process of CNN88, the color tone of the image and the degree of noise reduction for each pixel will not match, and as a result, a phenomenon may occur where the color tone and edge reconstruction are incorrect. This is because loss functions that judge "the loss is small if the average value matches," such as simple MAE or MSE, have a tendency to try to force large-disparity image structures to come closer together.

[0151] In other words, as shown in Figure 17, the system attempts to quickly bridge the large difference between the highly sensitive and noisy example image 80 (for example, an ISO 12800 image) and the low-noise ground truth image 82 (for example, an ISO 200 image), destroying even the edge components. In the example shown in Figure 17, when the example image 80 and the ground truth image 82 are similar, most of the edge components remain after training, whereas when the example image 80 and the ground truth image 82 are not similar, the edge components are significantly attenuated after training.

[0152] Therefore, in this third embodiment, in addition to pairs of example images 80 with high sensitivity noise (for example, an ISO 12800 image) and correct answer images 82 with little noise (for example, an ISO 200 image), intermediate noise levels, such as a pair of an ISO 6400 image and an ISO 200 image, are also added to the training data 78, thereby providing a wider range of noise variation during learning. By providing a wider range of noise, it is possible to suppress abrupt destruction of the image structure due to large differences in ISO sensitivity (for example, edges disappearing), and to proceed with learning more stably.

[0153] One specific method for introducing a range of noise is to obtain each example image 80 by mixing a first noise image 80B and a second noise image 82B at a variable mixing ratio according to the following formula (1), as shown in Figure 18. The first noise image 80B and the second noise image 82B are the example image 80 and the correct image 82 included in one training data 78. The first noise image 80B is an image containing noise caused by the first sensitivity, which is the ISO sensitivity used for the first imaging, and the second noise image 82B is an image containing noise caused by the second sensitivity, which is lower than the first sensitivity. In the following formula (1), N1 is the amount of noise contained in the first noise image 80B, GT is the amount of noise contained in the second noise image 82B, N2 is the amount of noise contained in the example image 80, and α is a variable coefficient of 0.5 ± 1. By changing α, the amount of noise to be included in the example image 80 is determined.

[0154] N2=N1×α+GT×(1-α)・・・(1)

[0155] In this way, by mixing the first noise image 80B and the second noise image 82B at a variable mixing ratio according to the following formula (1) to obtain each example image 80, intermediate noise level images can be artificially generated and used to train the CNN 88. As a result, multiple training data 78 covering a wide range of noise levels can be generated without performing the first imaging while gradually changing the ISO sensitivity. Furthermore, compared to the case where the first imaging is not performed at multiple ISO sensitivities, the training data can be enriched, resulting in an image processing AI 68 that can handle rapid noise changes (in other words, a rapid gap from high sensitivity to low sensitivity). Therefore, compared to the case where multiple example images 80 with only a constant amount of noise are mixed, noise reduction can be achieved without significantly distorting the image structure within the image.

[0156] In this third embodiment, the first noise image 80B is an example of the "first noise image" according to the present disclosure. Also, in this third embodiment, the second noise image 82B is an example of the "second noise image" according to the present disclosure.

[0157] [Fourth Embodiment] In the above embodiments, an image processing AI 68 was exemplified, but the disclosure is not limited thereto. For example, as shown in Figure 19, an image processing AI 124 may be applied instead of the image processing AI 68. The image processing AI 124 differs from the image processing AI 68 in that, in addition to the RAW image 22, imaging condition information 126 is also input. The imaging condition information 126 is information that can identify the imaging conditions used in imaging to obtain the RAW image 22. Furthermore, the imaging condition information 126 is information in which the imaging conditions used in imaging to obtain the RAW image 22 are expressed as one-hot vectors. The imaging condition used in imaging to obtain the RAW image 22 is the ISO sensitivity. Here, the imaging condition used in imaging to obtain the RAW image 22 is an example of the "second imaging condition" according to the disclosure, and the imaging condition information 126 is an example of the "second imaging condition information" according to the disclosure.

[0158] Figure 20 shows a modified example of the first machine learning performed on model 128 to obtain the image processing AI 124. Model 128 differs from model 84 described in the above embodiment in that it has a CNN 130 instead of a CNN 88, a fully connected layer 132, and a fusion layer 134. While the first machine learning described in the first embodiment is performed on model 84 using multiple training data 78, the first machine learning according to this fourth embodiment is performed on model 128 using multiple training data 78 and example imaging condition information 136. An example of example imaging condition information 136 is information that simulates imaging condition information 126.

[0159] The example imaging condition information 136 is information that can identify the imaging conditions used in the first imaging. Furthermore, the example imaging condition information 136 is information in which the imaging conditions used in the first imaging are represented as one-hot vectors. The imaging condition used in the first imaging is the ISO sensitivity. Here, the example imaging condition information 136 is an example of the "first imaging condition information" relating to this disclosure, and the imaging conditions used in the first imaging are an example of the "first imaging conditions" relating to this disclosure.

[0160] CNN 130 has a pre-convolutional layer group 130A and a post-convolutional layer group 130B. The pre-convolutional layer group 130A has the same structure as CNN 88 described in the first embodiment above. Therefore, the pre-convolutional layer group 130A generates the inference result 100 described in the first embodiment above. The pre-convolutional layer group 130A then outputs the inference result 100 to the fusion layer 134.

[0161] The fully connected layer 132 receives example imaging condition information 136 as input. The fully connected layer 132 maps the vector of example imaging condition information 136 to another vector, the imaging condition vector 138, by performing a linear transformation using matrices and biases on the vector of example imaging condition information 136. For example, in the case of mapping from a 10-dimensional vector to a 5-dimensional vector, the fully connected layer 132 generates a 5-dimensional imaging condition vector 138 by performing a linear transformation using a 5x10 matrix and a 5-dimensional bias on the example imaging condition information 136, which is represented as a 10-dimensional vector.

[0162] The fusion layer 134 broadcasts the imaging condition vector 138 to the inference result 100 in the spatial direction and concatenates it in the channel direction, thereby fusing the inference result 100 and the imaging condition vector 138, and outputs the fusion result 140 to the subsequent convolutional layer group 130B.

[0163] The subsequent convolutional layer group 130B performs convolution operations on the fusion result 140 using multiple layers. This allows the CNN 130 to learn noise reduction processing (for example, "strengthen noise reduction when ISO sensitivity is high") in accordance with the differences in imaging conditions used in the first imaging. The subsequent convolutional layer group 130B outputs the inference result 142 generated by the convolution operations.

[0164] Here, the processor 74 executes the optimization process according to this fourth embodiment. In the optimization process according to this fourth embodiment, first, the error 144 between the inference result 142 and the multi-channel image 96 is calculated. The error 144 is a value evaluated by MAE.

[0165] In the optimization process according to this fourth embodiment, a plurality of adjustment values ​​146 that minimize the error 144 are calculated. Then, a plurality of learning parameters in the CNN 130 and a plurality of learning parameters in the fully connected layer 132 (for example, matrices and biases used in linear transformations) are adjusted by the plurality of adjustment values ​​146.

[0166] A series of processes—inputting the multi-channel image 94 into the CNN 130, inputting the example imaging condition information 136 into the fully connected layer 132, calculating the error 144, calculating multiple adjustment values ​​146, adjusting multiple learning parameters within the CNN 130, and adjusting multiple learning parameters within the fully connected layer 132—are repeatedly performed using all the multi-channel images 94 and 96 obtained by the multi-channelization layer 86. As a result, multiple learning parameters are adjusted so that noise contained in the multi-channel image 94 is removed considering the imaging conditions, and the model 128 is optimized. By optimizing the model 128 in this way, the image processing AI 124 is obtained.

[0167] As described above, the first machine learning according to this fourth embodiment is machine learning that uses a plurality of training data 78 and example imaging condition information 136. Therefore, compared to the case where machine learning is performed using only the training data 78, the model 128 can learn noise reduction that takes imaging conditions into account. As a result, compared to the case where machine learning is performed using only the training data 78, the image processing AI 124 can acquire the optimal degree of noise reduction according to different imaging conditions.

[0168] Furthermore, in this fourth embodiment, when the RAW image 22 and imaging condition information 126 are input to the image processing AI 124, the image processing AI 124 performs noise reduction processing on the RAW image 22 according to the imaging conditions specified by the imaging condition information 126. As a result, the image processing AI 124 can generate a multi-channel image 112 (see Figures 10 and 11) with more appropriate control of the noise level and color reproduction compared to when no imaging conditions are provided. As a result, noise reduction with suppressed color degradation and image structure collapse can be achieved.

[0169] Furthermore, in this fourth embodiment, both the imaging condition information 126 and the example imaging condition information 136 are represented as one-hot vectors. Therefore, by inputting the imaging conditions as one-hot vectors into the model 128, the model 128 can learn each imaging condition clearly without forcing a numerical continuity relationship between them. Also, by inputting the imaging conditions as one-hot vectors into the image processing AI 124, the image processing AI 124 can infer each imaging condition clearly without forcing a numerical continuity relationship between them. As a result, compared to the case where the imaging conditions are simply input numerically into the image processing AI 124, it is possible to stably handle different imaging environments and improve the accuracy of noise reduction.

[0170] In the fourth embodiment described above, ISO sensitivity was given as an example of imaging conditions used for obtaining the RAW image 22. However, this is merely an example, and the imaging conditions used for obtaining the RAW image 22 may include ISO sensitivity, shutter speed, and / or F-number.

[0171] In the fourth embodiment described above, information represented by a one-hot vector was used as an example, but this is merely one example. Binary 0s and 1s may also be used, or values ​​with a range such as 0.0 to 0.1 may be used as one data point.

[0172] In the fourth embodiment described above, the case in which the first machine learning is performed was explained, but in addition to the first machine learning, the second machine learning described in the second embodiment may also be performed.

[0173] [Fifth Embodiment] In the above embodiments, a first machine learning method using a multi-channel image 94 obtained by bundling a plurality of segmented images 80A in the channel direction has been illustrated, but the disclosure is not limited thereto. For example, a three-channel image 148 shown in Figure 21 may be used for the first machine learning method. In the example shown in Figure 21, the multi-channelization layer 86 generates a plurality of segmented images 80A1 by averaging the plurality of segmented images 80A among the same colors in the X-Trans array, and generates a three-channel image 148 by bundling the plurality of segmented images 80A1 in the channel direction. The three-channel image 148 can also be described as a multi-channel image obtained by modifying a plurality of segmented images 80A based on a predetermined unit (for example, a repeating unit of the X-Trans array).

[0174] Here, the multiple segmented images 80A are examples of the "multiple first single-channel images" related to this disclosure. The three-channel image 148 is also an example of the "multiple first example images" and "first image group" related to this disclosure.

[0175] As an example, as shown in Figure 22, the multi-channelization layer 86 of the image processing AI 68, which is obtained by optimizing the model 84 by using the 3-channel image 148 in the first machine learning, generates a multi-channel image 110 in the same manner as in the first embodiment when a RAW image 22 is input. The multi-channelization layer 86 then generates multiple divided images 22A1 by averaging all the divided images 22A included in the multi-channel image 110 among the same colors in the X-Trans array, and generates a 3-channel image 150 by bundling the multiple divided images 22A1 in the channel direction. Furthermore, in this fifth embodiment, a noise-reduced image 70 is generated based on the 3-channel image 150 in the same manner as in the first embodiment where a noise-reduced image 70 is generated based on the multi-channel image 110.

[0176] Here, the multiple segmented images 22A are an example of the "multiple second single-channel images" related to this disclosure. The multiple segmented images 80A1 are an example of the "multiple first segmented images" and "third image group" related to this disclosure. The noise-reduced image 70 generated based on the three-channel image 150 is an example of the "third single-channel image" related to this disclosure.

[0177] The resolution of the noise-reduced image 70 generated based on the multi-channel image 110 is the same as the resolution of the multi-channel image 110. Therefore, the noise-reduced image 70 is upsampled by the processor 52 to the resolution before averaging between identical colors in the X-Trans array. The upsampled image obtained by upsampling the noise-reduced image 70 is an image with the same resolution as the RAW image 22, so the processor 52 can blend the upsampled image and the RAW image 22. One example of blending is averaging the signals between pixels. The upsampled image is an example of the "third single-channel image" according to this disclosure.

[0178] Figure 23 shows a flowchart illustrating an example of the noise reduction process according to this fifth embodiment. The flowchart in Figure 23 differs from the flowchart in Figure 14 in that it includes the processes of step ST100 and step ST102 instead of the process of step ST24.

[0179] In step ST100, which is included in the noise reduction process shown in Figure 23, the processor 52 acquires the noise-reduced image 70 and generates an upsampled image by upsampling it to return to the resolution before averaging between the same colors in the X-Trans array. After the processing of step ST100 is executed, the noise reduction process moves on to step ST102.

[0180] In step ST102, the processor 52 blends the upsampled image and the RAW image 22. After the processing in step ST102 is completed, the noise reduction process proceeds to step ST26.

[0181] As described above, in this fifth embodiment, since multiple segmented images 80A1 obtained by averaging all segmented images 22A included in the multi-channel image 110 among the same colors in the X-Trans array are used in the first machine learning, an image processing AI 68 with enhanced noise resistance can be constructed compared to the case where the multi-channel image 110 is used directly in the first machine learning without averaging. As a result, the image processing AI 68 can be made less susceptible to the influence of noise included in the RAW image 22 (for example, random pixel variations and disturbances). Furthermore, even if the RAW image 22 contains a lot of noise, the noise-reduced image 70 can maintain its original image structure and color reproduction.

[0182] Furthermore, in this fifth embodiment, a 3-channel image 150 obtained by averaging multiple segmented images 22A among those of the same color is used during inference, so extraneous noise is mitigated in advance and then further reduced by the image processing AI 68.

[0183] Furthermore, in this fifth embodiment, the RAW image 22 is multi-channelized and averaged to reduce its resolution, then noise reduction is performed by the CNN 88, and then it is upsampled to return to the original resolution. This suppresses noise and also allows for the interpolation of edges, etc. Finally, by blending the RAW image 22 and the upsampled image, an image with reduced noise can be obtained without compromising resolution and texture compared to the case where the images are not blended.

[0184] In the fifth embodiment described above, an example was given in which the first machine learning using the 3-channel image 148 is performed on model 84. However, this is merely one example, and the first machine learning using the 3-channel image 148 may also be performed on model 128 as shown in Figure 20. Furthermore, the 3-channel image 148 may also be used in the second machine learning described in the second embodiment described above.

[0185] In the fifth embodiment described above, an example was given in which the noise reduction image 70 is upsampled. However, the disclosure is not limited thereto, and the noise reduction image 70 may be restored by super-resolution to the resolution before averaging between the same colors (i.e., the same resolution as the RAW image 22).

[0186] In the fifth embodiment described above, an example of the morphology of the X-Trans sequence was given, but this is merely one example. Even if the RAW sequence rules are different from the repeating units of the X-Trans sequence, the same effect can be obtained by processing in the same manner as in the fifth embodiment.

[0187] [Sixth Embodiment] In the fifth embodiment described above, an example was given in which a 3-channel image 148 is used for the first machine learning. However, the disclosure is not limited thereto, and for example, a 39-channel image 152 shown in Figure 24 may be used for the first machine learning.

[0188] In the example shown in Figure 24, the multi-channel layer 86 generates three divided images 80A1 by averaging multiple divided images 80A among the same colors in the X-Trans array, and generates a 39-channel image 152 by bundling the multiple divided images 80A (i.e., 36 divided images 80A) and the three divided images 80A1 in the channel direction. The 39-channel image 152 can also be described as a multi-channel image obtained by modifying multiple divided images 80A based on a predetermined unit (for example, the repeating unit of the X-Trans array).

[0189] In this sixth embodiment, the multiple segmented images 80A are an example of the "multiple first single-channel images" according to the disclosure. Also in this sixth embodiment, the 39-channel image 152 is an example of the "multiple first example images" and "second image group" according to the disclosure. Also in this sixth embodiment, the segmented image 80A1 is an example of the "first additional image" according to the disclosure.

[0190] By the way, the number of output channels of the CNN 130 shown in Figure 20 is 36, which does not match the number of channels of the 39-channel image 152. Therefore, in this sixth embodiment, in order to match the number of channels, as an example, as shown in Figure 25, the first machine learning is performed on model 154 instead of model 128 shown in Figure 20, so that the image processing AI 156 shown in Figure 26 is obtained as an example.

[0191] Model 154 differs from Model 128 in that it has CNN 158 instead of CNN 130. CNN 158 differs from CNN 130 in that it has a subsequent convolutional layer group 158A instead of a subsequent convolutional layer group 130B. The subsequent convolutional layer group 158A differs from the subsequent convolutional layer group 130B in that it has a channel matching layer 158A1 and outputs an inference result 160.

[0192] The 39-channel image 152 generated by the multi-channel layer 86 is subjected to convolution by the preceding convolution layer group 130A, and the inference result 100 of a 39-dimensional vector is output to the fusion layer 134. The fusion layer 134 broadcasts the imaging condition vector 138 to the 39-dimensional vector inference result 100 in the spatial direction and concatenates it in the channel direction, thereby fusing the inference result 100 and the imaging condition vector 138, and outputs the fused 39-dimensional vector result 140 to the subsequent convolution layer group 130B.

[0193] The subsequent convolutional layer group 130B has multiple convolutional layers and a channel matching layer 158A1 arranged in sequence. The channel matching layer 158A1 is a convolutional layer that uses a kernel with a spatial size of 1x1, and adjusts the 39-dimensional vector to a 36-dimensional vector by performing a linear transformation on the 39-dimensional vector of a single pixel using a 36x39 matrix and a 36-dimensional bias. As a result, the subsequent convolutional layer group 158A outputs a 36-channel inference result 160.

[0194] As an example, as shown in Figure 26, the RAW image 22 is input to the multi-channelization layer 86 of the image processing AI 156 obtained by optimizing the model 154 (see Figure 25) through a first machine learning process, similar to the embodiments described above. The multi-channelization layer 86 converts the RAW image 22 into a multi-channel image 110, similar to the embodiments described above. The multi-channelization layer 86 then generates three segmented images 22A1 by averaging the multi-channel image 110 among the same colors in the X-Trans array, and generates a 39-channel image 162 by bundling the multiple segmented images 22A (i.e., 36 segmented images 22A) and the three segmented images 22A1 in the channel direction.

[0195] The 39-channel image 162 is noise-reduced by the CNN 158 to generate a 36-channel multi-channel image 112. Subsequently, various processes such as the generation of a noise-reduced image 70 are performed, similar to the embodiments described above. For example, in this sixth embodiment, a noise-reduced image 70 is generated in the same manner as in the first embodiment, where a noise-reduced image 70 is generated based on the multi-channel image 112.

[0196] In this sixth embodiment, the multiple segmented images 22A are an example of the "multiple second single-channel images" according to the disclosure. Also in this sixth embodiment, the 39-channel image 162 is an example of the "multiple first segmented images" and "fourth image group" according to the disclosure. Also in this sixth embodiment, the segmented image 22A1 is an example of the "second additional image" according to the disclosure.

[0197] As explained above, in this sixth embodiment, since the first machine learning is performed using the 39-channel image 152, a different learning example is added that cannot be obtained from the multiple segmented images 22A alone. That is, the three segmented images 80A1 (i.e., averaged RGB images) added to the multiple segmented images 80A (i.e., 36 segmented images 80A) included in the 39-channel image 152 behave as information that smooths out local noise, providing the model 154 with spatially and statistically stable information, so that the model 154 can learn to distinguish between noise and edge components more clearly. This makes it possible to achieve both smoothing of noise components and preservation of the original image structure. Therefore, compared to the case where the first machine learning is performed without the addition of a different learning example, the noise reduction effect and image structure preservation effect of the image processing AI 156 can be enhanced. As a result, compared to the case where the first machine learning is performed without the addition of a different learning example, the image processing AI 124 is able to generate a high-quality noise-reduced image 70 that achieves both noise removal and image structure preservation.

[0198] Furthermore, in this sixth embodiment, since the 39-channel image 162 is used during inference, a noise-reduced image 70 with significantly reduced noise can be obtained compared to the case where only the multiple segmented images 22A are used during inference.

[0199] In the sixth embodiment described above, an example of the morphology of the X-Trans sequence was given, but this is merely an example. Even if the RAW sequence rules are different from the repeating units of the X-Trans sequence, the present disclosure can be established by processing them in the same manner as in the sixth embodiment.

[0200] [Seventh Embodiment] In the above embodiments, examples were given in which the multi-channelization layer 86 directly multi-channelized the example image 80 to generate a multi-channel image 94, but the disclosure is not limited thereto. For example, as shown in Figure 27, with respect to a block 92 of the example image 80, the multi-channelization layer 86 may combine the same sequence into two sets in order to reduce the resolution by half, average each set, and then multi-channelize between the same colors.

[0201] In the example shown in Figure 27, block 92 is divided into four 3x3 subblocks 92A. In this case, the top left 3x3 subblock 92A and the bottom right 3x3 subblock 92A have the same color arrangement, and the top right 3x3 subblock 92A and the bottom left 3x3 subblock 92A have the same color arrangement.

[0202] Therefore, in order to reduce the resolution by half, corresponding pixels in the upper left 3x3 subblock 92A and the lower right 3x3 subblock 92A are matched, and pixels of the same color are grouped together and averaged. Similarly, corresponding pixels in the upper right 3x3 subblock 92A and the lower left 3x3 subblock 92A are matched, and pixels of the same color are grouped together and averaged.

[0203] As a result, the 6x6 block 92 is reduced to a 3x3 block. Then, in the same manner as the multi-channelization described in the first embodiment above, the reduced 3x3 block is multi-channelized, thereby generating a multi-channel image 164 in which multiple divided images 164A are bundled in the channel direction. The multiple divided images 164A may be homochromatically averaged in the same manner as in the example shown in Figure 21, or homochromatically averaged 3-channel images may be added to the multiple divided images 164A in the same manner as in the example shown in Figure 24. The restoration of the reduced-resolution image to its original resolution is achieved by upsampling or super-resolution of the multiple divided images 164A.

[0204] Thus, in this seventh embodiment, the 6x6 block 92 is divided into 3x3 subblocks 92A, and each is averaged, thereby reducing the spatial information of the original high-resolution image to a lower resolution. As a result, for example, the input size of the subsequent convolutional layer group 158A (see Figure 25) is reduced, and the computational load for learning and inference is greatly reduced. Consequently, the processing speed is improved.

[0205] Furthermore, in this seventh embodiment, local noise is smoothed out and noise components are more easily canceled out by averaging subblocks 92A with the same color sequence. As a result, even in areas where noise is mixed in the mosaic state of a single RAW image, a stable signal is obtained through averaging, allowing the image processing AI 124 to distinguish between noise and edge components more accurately.

[0206] Furthermore, in this seventh embodiment, the multi-channel configuration rearranges each subblock 92A in the channel direction, so that each channel contains the same type of pixel (e.g., R pixel, G pixel, and B pixel). This allows models 84, 128, or 154 to learn different noise and signal characteristics for each channel. As a result, the noise reduction performance and color reproduction of the image processing AI 68, 124, or 156 are improved.

[0207] Furthermore, in this seventh embodiment, by using a low-resolution multi-channel image 164 obtained by dividing and averaging the example image 80, different perspectives (for example, local averaging information and fine local information) are used in combination, allowing models 84, 128, or 154 to learn noise reduction and image reconstruction from a more multifaceted perspective.

[0208] Furthermore, in this seventh embodiment, by applying upsampling or super-resolution in a later stage, it is possible to generate an image with noise removed while restoring the features extracted at low resolution back to their original resolution, thus ultimately obtaining a high-quality image.

[0209] In the seventh embodiment described above, an example of the morphology of the X-Trans sequence was given. However, this is merely one example, and even if the RAW sequence rules differ from those of the X-Trans sequence repeating unit, the same effect can be obtained by processing in the same manner as in the seventh embodiment.

[0210] [Eighth Embodiment] In this eighth embodiment, as an example, as shown in Figure 28, an example of a configuration in which phase difference pixels PD1 and PD2 for measuring image plane phase difference are arranged in predetermined periodic units in a RAW image 166 obtained from an image sensor 16 employing an X-Trans array. Here, the RAW image 166 is assumed to be a RAW image for training (for example, an example image) and a RAW image for inference (for example, an image corresponding to the RAW image 22 described in the above embodiment).

[0211] The RAW image 166 has a periodic unit of 6 x 6 = 36 pixels per block. In addition, the bottom row of the 6 x 6 block 168 contained in the RAW image 166 has phase difference pixels PD1 and PD2 arranged alternately.

[0212] The multi-channel layer 86 generates 36 divided images 170A by dividing the RAW image 22 using a method such as SpaceToDepth or PixelSuffle, in the same manner as in the first embodiment described above. The multi-channel image 170 is obtained by arranging the 36 divided images 170A along the channel direction.

[0213] The multi-channel layer 86 generates segmented images 170A for channels 1 to 6 by placing each pixel of the top row (i.e., the first row) of each block 168 into channels 1 to 6. Similarly, the multi-channel layer 86 generates segmented images 170A for channels 7 to 30 by placing each pixel of the second to fifth rows of each block 168 into channels 7 to 30. Furthermore, the multi-channel layer 86 generates segmented images 170A for channels 31 to 36 by placing the pixels of the sixth row of each block 168 into channels 31 to 36. Channels 31, 33, and 35 correspond to the values ​​of phase difference pixels PD1, while channels 32, 34, and 36 correspond to the values ​​of phase difference pixels PD2.

[0214] The multi-channel image 170 obtained in this way can be processed in the same way as in each of the above embodiments to obtain the same effects as in each of the above embodiments.

[0215] As a result of rearranging each pixel of the RAW image 166 in this way, each channel will contain information only from pixels of the same type, allowing models 84, 128, or 154 to individually learn different noise and signal characteristics for each channel.

[0216] Furthermore, in this eighth embodiment, since the phase difference pixels PD1 and PD2 are arranged in predetermined periodic units, the RAW image 166 can be divided considering the presence of phase difference pixels PD1 and PD2, which differ from the normal RGB color arrangement. As a result, noise is reduced in a way that does not impair the characteristics of image plane phase difference measurement, compared to when the phase difference pixels PD1 and PD2 are processed without following predetermined periodic units. In other words, since the phase difference pixels PD1 and PD2 have different light-receiving characteristics and noise characteristics, processing them as separate channels allows for more accurate noise reduction without losing information specific to the phase difference pixels PD1 and PD2 during noise reduction.

[0217] In the eighth embodiment described above, phase difference pixels PD1 and PD2 arranged in predetermined periodic units within the RAW image 166 were used as an example, but the present invention is not limited thereto. For example, if a training image or inference image is provided with a function different from that of a color filter and includes a plurality of periodically arranged functional pixels, then by performing the same processing as in the eighth embodiment, noise can be reduced in a manner that does not impair the inherent characteristics of the plurality of functional pixels, compared to when the functional pixels are processed without following predetermined periodic units. Here, examples of functions different from those of a color filter include the image plane phase difference measurement, infrared measurement, depth measurement, multispectral measurement, and / or polarization measurement described above.

[0218] [Ninth Embodiment] In the above embodiments, examples were given in which noise reduction processing is performed by a processing device 24 included in the imaging device 10. However, the technology of this disclosure is not limited thereto, and the device that performs noise reduction processing may be provided outside the imaging device 10. In this case, as an example, an information processing system 172 may be used, as shown in Figure 29.

[0219] The information processing system 172 includes an imaging device 10 and an information processing device 174. For example, the information processing device 174 is a server. The server is implemented, for example, by cloud computing. Here, cloud computing is given as an example, but this is merely one example, and for example, the server may be implemented by a mainframe, or by network computing such as fog computing, edge computing, or grid computing. Here, a server is given as an example of the information processing device 174, but this is merely one example, and instead of a server, at least one personal computer or the like may be used as the information processing device 174.

[0220] The information processing device 174 includes a processor 176, an NVM 178, a RAM 180, and a communication interface 182. The processor 176, NVM 178, RAM 180, and communication interface 182 are connected to a bus 184. The communication interface 182 is connected to the imaging device 10 via a network 186. The network 186 is, for example, the internet. Note that the network 186 is not limited to the internet, but may also be a WAN and / or a LAN such as an intranet.

[0221] The NVM 178 stores a noise reduction processing program 66 and an image processing AI 68. The processor 176 executes the noise reduction processing program 66 on the RAM 180. The processor 176 performs the noise reduction processing described above according to the noise reduction processing program 66 executed on the RAM 180. The image processing AI 68 is used by the processor 176 in the noise reduction processing.

[0222] The processor 42 of the imaging device 10 (see Figure 1) generates processing request information 188 that requests the information processing device 174 to perform control processing and noise reduction processing, and transmits the generated processing request information 188 to the information processing device 174 via the network 186. The processing request information 188 includes the RAW image 22 and metadata 116B.

[0223] In the information processing device 174, processing request information 188 is received via the communication interface 182. The processor 176 obtains the RAW image 22 and metadata 116B from the processing request information 188. The processor 176 generates a multi-channel image 110 by multi-channelizing the RAW image 22 for each unit identified from the unit identification information 116B1 contained in the metadata 116B. Then, the processor 176 generates a noise-reduced image 70 based on the multi-channel image 110 in the manner described in each of the embodiments above. The processor 176 transmits the noise-reduced image 70 to the imaging device 10 in response to a request from the imaging device 10. In the imaging device 10, the noise-reduced image 70 is received via the communication interface 34, and various processing is performed by the processing device 24.

[0224] It should be noted that this example illustrates how the noise-reduced image 70 is transmitted from the information processing device 174 to the imaging device 10, but this is merely one example. For instance, the processor 176 of the information processing device 174 may generate an image file 116, and the generated image file 116 may be transmitted from the information processing device 174 to the imaging device 10 as the noise-reduced image 70.

[0225] Thus, in the information processing system 172, the information processing device 174 can identify units such as the RAW sequence rules of the RAW image 22 from the unit identification information 116B1 contained in the metadata 116B provided from the imaging device 10. This reduces the effort required and lowers the risk of incorrect settings compared to when a user manually sets the units such as the RAW sequence rules of the RAW image 22. As a result, the information processing device 174 can recognize the units such as the RAW sequence rules and perform noise reduction processing. Furthermore, since the information processing device 174 multi-channels the RAW image 22 for each unit identified from the unit identification information 116B1 contained in the metadata 116B provided from the imaging device 10, it can multi-channel the RAW image 22 according to the correct periodic structure compared to when the RAW image 22 is multi-channeled haphazardly without unit identification information 116B1.

[0226] In the example shown in Figure 29, the information processing system 172 is an example of an "information processing system" related to this disclosure, the information processing device 174 is an example of an "information processing device" related to this disclosure, and the processor 176 is an example of a "processor" related to this disclosure.

[0227] In each of the above embodiments, each process is executed on any computer. Furthermore, any computer may execute these processes using a processor as hardware, a program as software, or a combination thereof. In this case, the processor is configured to work in cooperation with the program to execute the various processes in the above embodiments, and can function as a unit or means in the above embodiments. Also, the execution order of the processes by the processor is not limited to the order described and may be changed as appropriate. Any computer may be a general-purpose computer, a computer designed for a specific purpose, a workstation, or any other system capable of executing the processes.

[0228] A processor may consist of one or more hardware components, and the type of hardware is not limited. For example, a processor may consist of hardware such as a CPU, MPU, FPGA or other programmable logic device, ASIC or other dedicated circuitry for executing specific processes, GPU, and / or NPU. The type of hardware may also be a combination of different types of hardware. When multiple hardware components are configured to execute one or more processes of a processor, the multiple hardware components may reside in physically separate devices or in the same device. Furthermore, in any embodiment, the order of each process performed by the processor is not limited to the order described above and may be changed as appropriate. Hardware is composed of electrical circuits (circuitry) that combine circuit elements such as semiconductor elements.

[0229] Furthermore, the program may be software such as firmware or microcode. Alternatively, the program may be, for example, a group of program modules, each function of which may be implemented by a processor configured to perform its respective function. The program may also be program code and / or multiple code segments stored on one or more non-temporary computer-readable media (e.g., storage). The program may be divided and stored on multiple non-temporary computer-readable media located in physically separate devices. Program code or code segments may represent any combination of procedures, functions, subprograms, routines, subroutines, modules, software packages, classes, instructions, data structures, or program statements. Program code or code segments may be connected to other code segments or hardware circuits by sending and receiving information, data, arguments, parameters, or memory contents.

[0230] Furthermore, while the above embodiment illustrates a configuration in which the control program 50 and the noise reduction processing program 66 are pre-stored in the NVM (i.e., installed), this disclosure is not limited thereto. The control program 50 and / or the noise reduction processing program 66 may be provided in a form stored on a storage medium such as a CD-ROM, DVD-ROM, and / or USB memory. Alternatively, the control program 50 and / or the noise reduction processing program 66 may be provided in a form that can be downloaded from an external device via a network.

[0231] This disclosure covers all program products. Program products include all forms of products for providing programs. For example, program products include programs provided via networks such as the Internet, and non-temporary computer-readable storage media such as CD-ROMs, DVDs, and USB memory sticks on which programs are stored.

[0232] The control and noise reduction processes described above are merely examples. Therefore, it goes without saying that you may remove unnecessary steps, add new steps, or change the processing order, as long as you do not deviate from the main purpose.

[0233] The descriptions and illustrations presented above are detailed explanations of the parts related to this disclosure and are merely examples of this disclosure. For example, the above explanation of the structure, function, operation, and effect is an example of the structure, function, operation, and effect of the parts related to this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace parts of the descriptions and illustrations presented above, as long as you do not deviate from the spirit of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the parts related to this disclosure, explanations of common technical knowledge, etc., that do not require special explanation to enable the implementation of this disclosure have been omitted from the descriptions and illustrations presented above.

[0234] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0235] The following additional information is disclosed regarding the embodiments described above.

[0236] (Note 1) Training data used for AI machine learning, comprising: a plurality of first example images obtained by dividing a third image, which is a RAW image, based on predetermined units; and a plurality of first correct answer images obtained by dividing a second captured image or an image corresponding to the second captured image based on the above units, wherein the third image is a first captured image and / or an image corresponding to the first captured image obtained by performing a first capture, and the plurality of first correct answer images are images obtained by dividing a second captured image or an image corresponding to the second captured image obtained by performing a second capture which has less noise than the first capture, based on the above units.

[0237] (Note 2) The above-mentioned multiple example images are training data as described in Note 1, provided for each different sensitivity used in the first imaging.

[0238] (Note 3) The third image described above is an image obtained by mixing the first noise image and the second noise image at a variable mixing ratio, wherein the first noise image contains noise caused by the first sensitivity used in imaging, and the second noise image is used as the first ground truth image and contains the training data described in Note 2 containing noise caused by the second sensitivity which is lower than the first sensitivity.

[0239] (Note 4) Training data as described in any one of Notes 1 to 3, further comprising first imaging condition information that can identify the first imaging conditions used in imaging to obtain the above-mentioned first image.

[0240] (Note 5) The above-mentioned plurality of first example images are a first group of images obtained by averaging a plurality of first single-channel images obtained by converting the third image into single channels for each unit, and then averaging these single-channel images among the same colors within the unit, or a second group of images including the plurality of first single-channel images and at least one first additional image, wherein the first additional image is an image obtained by averaging the first single-channel images among the plurality of first single-channel images among the same colors within the unit, as described in any one of Notes 1 to 4.

[0241] (Note 6) The above third image is a third image obtained by performing a third imaging and / or an image corresponding to the above third image, and the training data further comprises the above third image and a second ground truth image which is a fourth image obtained by performing a fourth imaging that has less noise than the above third imaging, or an image corresponding to the above fourth image.

[0242] (Note 7) The image corresponding to the third captured image is the training data described in Note 6, which is a single image obtained by the AI ​​when the multiple first example images are input into the AI ​​during the learning stage, and the multiple inferred images obtained by the AI ​​are converted into a single channel based on the above units.

[0243] (Note 8) A method for creating training data used in AI machine learning, wherein the training data comprises a plurality of first example images obtained by dividing a third image, which is a RAW image, based on predetermined units, and a plurality of first correct answer images obtained by dividing a second captured image or an image corresponding to the second captured image based on the above units, and the method includes creating the plurality of first example images and creating the plurality of first correct answer images, wherein the third image is a first captured image and / or an image corresponding to the first captured image obtained by performing a first capture, and the plurality of first correct answer images are images obtained by dividing a second captured image or an image corresponding to the second captured image obtained by performing a second capture which has less noise than the first capture, based on the above units.

[0244] (Note 9) A method for generating AI, comprising: inputting a plurality of first example images obtained by dividing a third image, which is a RAW image, based on predetermined units into a model; the model outputting a plurality of inference images, which are inference results for the plurality of first example images, in response to the input of the plurality of first example images; and optimizing the model based on the comparison result between a plurality of first correct images obtained by dividing an captured image or an image corresponding to the captured image based on the above units and the plurality of inference images.

Claims

1. An image processing device comprising a processor, the processor outputting at least one of a first image which is a RAW image and a plurality of first divided images obtained by dividing the first image based on predetermined units, and acquiring a second image which has less noise than the first image based on the plurality of first divided images.

2. The image processing apparatus according to claim 1, wherein at least one of the first image and the plurality of first divided images is input to the AI, thereby generating a plurality of second divided images in which the noise of the plurality of first divided images is reduced, and the processor acquires the second image based on the plurality of second divided images.

3. The image processing apparatus according to claim 1, wherein the unit is a RAW sequence rule.

4. The image processing apparatus according to claim 3, wherein the RAW sequence rule includes the periodic unit of the color filter.

5. The image processing apparatus according to claim 4, wherein the plurality of pixels representing the first image are assigned functions different from those of the color filter and include a plurality of functional pixels arranged periodically, and the RAW arrangement rule includes a periodic unit in which the plurality of functional pixels are arranged.

6. The image processing apparatus according to claim 5, wherein the different functions include image plane phase difference measurement, infrared measurement, depth measurement, multispectral measurement, and / or polarization measurement.

7. The image processing apparatus according to claim 2, wherein the AI ​​is a trained model that has undergone first learning to reduce noise in a plurality of first example images obtained by dividing a third image, which is a RAW image, based on the units.

8. The image processing apparatus according to claim 7, wherein the third image is a first image obtained by performing a first imaging and / or an image corresponding to the first image, and the first learning is learning using the plurality of first example images and a plurality of first correct images obtained by dividing a second image obtained by performing a second imaging which has less noise than the first imaging, or an image corresponding to the second image, based on the unit.

9. The image processing apparatus according to claim 8, wherein the first learning is learning using the plurality of first example images for each different sensitivity used in the first imaging.

10. The image processing apparatus according to claim 9, wherein the third image is an image obtained by mixing a first noise image and a second noise image at a variable mixing ratio, the first noise image includes noise caused by a first sensitivity used in imaging, and the second noise image is used as the first ground truth image and includes noise caused by a second sensitivity lower than the first sensitivity.

11. The image processing apparatus according to claim 8, wherein the first learning is learning using the plurality of first example images, first imaging condition information that can identify the first imaging conditions used in imaging to obtain the first captured image, and the plurality of first correct images.

12. The image processing apparatus according to claim 11, wherein at least one of the first image and the plurality of first divided images is input to the AI, thereby generating a plurality of second divided images in which the noise of the plurality of first divided images has been reduced, the processor acquires the second image based on the plurality of second divided images, and the plurality of second divided images are images generated by the AI ​​by inputting at least one of the first image and the plurality of first divided images and second imaging condition information that can identify second imaging conditions used for imaging to obtain the first image to the AI.

13. The image processing apparatus according to claim 12, wherein the first imaging condition and the second imaging condition include ISO sensitivity, shutter speed, and / or F-number.

14. The image processing apparatus according to claim 12, wherein the first imaging condition information is information in which the first imaging condition is represented as a one-hot vector, and the second imaging condition information is information in which the second imaging condition is represented as a one-hot vector.

15. The image processing apparatus according to claim 7, wherein the AI ​​is a trained model requiring a convolution operation, and the kernel size used in the convolution operation is a size predetermined according to the unit as a size that can maintain the color periodicity of the third image in the convolution operation.

16. The image processing apparatus according to claim 7, wherein the plurality of first example images are a first image group obtained by averaging a plurality of first single-channel images obtained by converting the third image into a single channel for each unit, and the plurality of first single-channel images obtained by averaging the same-colored first single-channel images within the unit, and the first additional image is an image obtained by averaging the first single-channel images of the same color within the unit from among the plurality of first single-channel images.

17. The image processing apparatus according to claim 16, wherein the plurality of first divided images are a third image group obtained by averaging a plurality of second single-channel images obtained by single-channelizing the first image for each unit, and the second single-channel images obtained by averaging the same-colored images within the unit, or a fourth image group including the plurality of second single-channel images and at least one second additional image, and the second additional image is an image obtained by averaging the second single-channel images of the same color within the unit from among the plurality of second single-channel images.

18. The image processing apparatus according to claim 17, wherein at least one of the first image and the plurality of first divided images is input to the AI, thereby generating a plurality of second divided images in which the noise of the plurality of first divided images is reduced, the second image is an image obtained by mixing the first image and a third single-channel image, and the third single-channel image is an image in which the plurality of second divided images are converted to single channels, and the resolution after averaging among the same colors is returned to the resolution before averaging among the same colors by upsampling or super-resolution.

19. The image processing apparatus according to claim 7, wherein the AI ​​is a trained model that has undergone a second training to reduce noise in the third image.

20. The image processing apparatus according to claim 19, wherein the third image is a third image obtained by performing a third imaging and / or an image corresponding to the third image, and the second learning is learning using the third image and a second ground truth image which is a fourth image obtained by performing a fourth imaging which has less noise than the third imaging, or an image corresponding to the fourth image.

21. The image processing apparatus according to claim 20, wherein the image corresponding to the third captured image is a single image obtained by the AI ​​through inference when the plurality of first example images are input to the AI ​​in the learning stage, and the plurality of inferred images obtained by the AI ​​are converted into a single channel based on the unit.

22. The image processing apparatus according to claim 7, wherein the AI ​​is a trained model that has undergone a first training to reduce noise in a plurality of first example images obtained by dividing a third image, which is a RAW image, based on the units, and a second training to reduce noise in the third image or an image corresponding to the third image, the third image is a first captured image obtained by performing a first imaging and / or an image corresponding to the first captured image, the first training is training using a first difference metric, which is the degree of difference between the plurality of first example images and a plurality of first ground truth images obtained by dividing a second captured image or an image corresponding to the second captured image, obtained by performing a second imaging which has less noise than the first imaging, based on the units, the second training is training using a second difference metric, which is the degree of difference between the third image or an image corresponding to the third image and a third ground truth image which is the second captured image or an image corresponding to the second captured image, and the first difference metric and the second difference metric are evaluated using different indices.

23. The image processing apparatus according to claim 22, wherein the first difference is evaluated by a first indicator, the second difference is evaluated by a second indicator, and in the learning process of the AI, the first indicator has a higher effect of reducing image noise than the second indicator, and the second indicator has a higher effect of preserving the image structure of the image than the first indicator.

24. The image processing apparatus according to claim 23, wherein the first index is the mean absolute error in a zone different from the volume zone of the histogram of the mean absolute errors per pixel of the plurality of first example images and the plurality of first correct answer images, or the mean squared error in a zone different from the volume zone of the histogram of the mean squared errors per pixel of the plurality of first example images and the plurality of first correct answer images.

25. The image processing apparatus according to claim 7, wherein one of the types of the third image is an image in which frequencies in the Nyquist limit region are represented.

26. The image processing apparatus according to claim 1, wherein the processor acquires unit identification information that allows the identification of the unit, and the plurality of first divided images are obtained by dividing the first image according to the unit identified by the unit identification information.

27. The image processing apparatus according to claim 26, wherein the first image is an image obtained by imaging device, and the unit identification information is information included in metadata obtained in connection with imaging to obtain the first image by imaging device.

28. The image processing apparatus according to claim 2, wherein the second image is a single image obtained by converting the plurality of second segmented images into a single channel based on the unit.

29. An imaging apparatus comprising: an image processing apparatus according to any one of claims 1 to 28; and an image sensor for capturing images to obtain the first image.

30. An information processing system comprising: an image processing device according to any one of claims 1 to 28; and an information processing device that performs processing based on the second image.

31. A program for causing a computer to perform a process that includes outputting at least one of a first image which is a RAW image and a plurality of first divided images obtained by dividing the first image based on predetermined units, and obtaining a second image which has less noise than the first image based on the plurality of first divided images.