Image processing method, image processing device, imaging apparatus, and program

The method addresses the high processing load in image estimation by determining a region of interest and processing only that region using a machine learning model, effectively reducing computational demands.

JP2025102012APending Publication Date: 2025-07-08CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023219164
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Conventional image estimation processing using machine learning models applies processing to all regions of an image, including non-attention regions, leading to increased processing load.

Method used

An image processing method that determines a region of interest and processes only a part of the image using a machine learning model, generating a processed image based on the region of interest, thereby reducing processing load.

Benefits of technology

Reduces processing load by focusing image estimation processing only on the region of interest, optimizing computational resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025102012000001_ABST
    Figure 2025102012000001_ABST
Patent Text Reader

Abstract

To provide an image processing method that makes it possible to reduce the processing load in image estimation processing using a machine learning model.SOLUTION: The image processing method includes a step (S104) for determining an area of interest in a first image, a step (S105) for acquiring one or more second images corresponding only to a portion of a plurality of partial images of the first image, a step (S106)for generating a third image corresponding to the second image by inputting the second image into a machine training model, and a step(S110) for generating a fourth image by processing based on the third image and the area of interest. The second image includes at least a portion of the area of interest.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing method, an image processing apparatus, an imaging apparatus, and a program.

Background Art

[0002] Conventionally, in image estimation processing using a machine learning model, a method of reducing the amount of memory used by obtaining a plurality of divided images from an input image and performing image estimation processing for each divided image is known. Patent Document 1 discloses a method of dividing an input image and performing processing to correct a defocused image.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] By the way, there are cases where only a specific region in the input image is desired to be the attention region for image estimation processing. However, in the method disclosed in Patent Document 1, since image estimation processing is performed on all regions including regions other than the attention region, the processing load is large.

[0005] Therefore, an object of the present invention is to provide an image processing method capable of reducing the processing load in image estimation processing using a machine learning model.

Means for Solving the Problems

[0006] As one aspect of the present invention, an image processing method includes a step of determining a region of interest in a first image, a step of obtaining one or more second images corresponding to only a part of a plurality of partial images of the first image, a step of inputting the second image into a machine learning model to generate a third image corresponding to the second image, and a step of generating a fourth image by processing based on the third image and the region of interest, wherein the second image includes at least a part of the region of interest.

[0007] Other objects and features of the present invention will be described in the following embodiments.

Effects of the Invention

[0008] According to the present invention, it is possible to provide an image processing method capable of reducing the processing load in image estimation processing using a machine learning model.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Modes for Carrying Out the Invention

[0010] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In each figure, the same members are denoted by the same reference numerals, and duplicate descriptions are omitted.

[0011] The image processing apparatus of this embodiment performs a sharpening process using a machine learning model based on the optical characteristics of an optical system (imaging optical system) on an input image generated by imaging using the optical system. Here, the optical characteristics indicate the aberration of the optical system or the blur of the input image due to the aberration, and are, for example, a point spread function (PSF) or an optical transfer function (OTF). The optical characteristics may be a modulation transfer function (MTF) which is the amplitude component of the OTF, or a phase transfer function (PTF) which is the phase component of the OTF. By performing the sharpening process based on the optical characteristics, the blur of the input image can be effectively corrected based on the characteristics of the imaging optical system.

[0012] Before describing the configuration of the imaging apparatus of each example, the sharpening process of this embodiment will be described. In this embodiment, the sharpening process is an aberration correction process based on optical characteristics, and is also called an image restoration process, a point image restoration process, or the like. Further, the optical characteristics may include not only the imaging lens but also the characteristics of optical elements such as a low-pass filter and an infrared cut filter, and the influence due to the structure such as the pixel arrangement of the image sensor may be considered.

[0013] To train a machine learning model that performs sharpness enhancement processing (image estimation processing) based on the optical characteristics of an optical system, a training image is generated by adding blur based on the optical transfer function to the original image corresponding to the subject as the correct image. The training image is input to a machine learning model such as a neural network to estimate a sharpened image, and the machine learning model may be optimized so that the difference from the correct image approaches zero. At this time, the correct image and the training image only need to be in a relationship in which the blur due to the optical characteristics can be corrected, and conditions such as the presence or absence of noise, the presence or absence of development processing, and the resolution may be different. In each embodiment, sharpness enhancement processing based on the optical characteristics in the in-focus plane is performed in particular to correct the blur in the in-focus plane.

[0014] Next, the principle of calculating a focus map using a disparity image will be described. The disparity image can be obtained using an imaging unit that guides a plurality of light beams that have passed through different regions of the pupil of one imaging optical system to different light receiving portions (pixels) in one image sensor to perform photoelectric conversion. That is, a disparity image necessary for distance calculation can be obtained with one imaging unit (composed of one optical system and one image sensor).

[0015] FIG. 6 is a diagram showing the relationship between the light receiving portion of the image sensor and the pupil of the imaging optical system in the imaging unit. ML is a microlens, CF is a color filter, and EXP indicates the exit pupil of the imaging optical system. G1 and G2 are light receiving portions (hereinafter referred to as G1 pixel and G2 pixel, respectively), and one G1 pixel and one G2 pixel form a pair with each other. A plurality of pairs of G1 pixels and G2 pixels (pixel pairs) are arranged in the image sensor. The paired G1 pixel and G2 pixel have an approximately conjugate relationship with the exit pupil EXP via a common microlens ML (that is, one microlens ML is provided for each pixel pair). The plurality of G1 pixels arranged in the image sensor are also collectively referred to as a G1 pixel group, and similarly, the plurality of G2 pixels arranged in the image sensor are also collectively referred to as a G2 pixel group.

[0016] FIG. 7 is an explanatory diagram of the relationship between the light-receiving unit of the imaging device and the subject, and schematically shows an imaging system when it is assumed that there is a thin lens at the position of the exit pupil EXP in FIG. 6. Pixel G1 receives the light beam that has passed through region P1 of the exit pupil EXP, and pixel G2 receives the light beam that has passed through region P2 of the exit pupil EXP. OSP is the object point being imaged. It is not necessary for an object (subject) to exist at the object point OSP, and the light beam passing through this point enters pixel G1 or pixel G2 according to the region (position) within the pupil through which it passes. The fact that the light beam passes through different regions within the pupil corresponds to the incident light from the object point OSP being separated by an angle (parallax). That is, among pixels G1 and G2 provided for each microlens ML, the image generated using the output signal from pixel G1 and the image generated using the output signal from pixel G2 become a plurality (here, a pair) of parallax images having parallax with each other.

[0017] In the following description, the light beams passing through different regions within the pupil being received by different light-receiving units (pixels) is also referred to as pupil division. Here, an example of being divided into two in one pupil division direction has been described, but the pupil division direction and the number of divisions are arbitrary. For example, it may be divided in two directions, the horizontal direction and the vertical direction. Also, the pupil division direction and the number of divisions may vary according to the pixel positions of the imaging device. In pupil division, the parallax direction is the displacement direction when the divided pupils are regarded as respective viewpoints. The viewpoint position is defined based on the pupil region through which each light beam passes, and may be, for example, the centroid position of the light beam. Also, the pupil division direction is not limited to the horizontal and vertical directions, and may be an oblique direction.

[0018] Also, in FIGS. 6 and 7, even if the position of the exit pupil EXP is displaced, etc., so that the above-described conjugate relationship is not perfect or regions P1 and P2 partially overlap, the plurality of obtained images can be treated as parallax images.

[0019] The amount of displacement of the subject between the parallax images (parallax displacement amount) can be calculated by identifying the corresponding subject regions between the parallax images. Various methods can be used to identify the same subject region between images. For example, a block matching method using one of the parallax images as a reference image may be employed. Thereby, the parallax displacement amount can be obtained. In each embodiment, the parallax displacement amount is considered with respect to the in-focus plane. The distance from the imaging device to the focus of the imaging optical system (defocus amount) can be calculated based on the parallax displacement amount and the positional displacement amount of the viewpoints (baseline length). Also, the subject distance can be calculated using the focal length of the imaging optical system.

[0020] Note that the methods for obtaining the parallax displacement amount, defocus amount, and subject distance are not limited to the above. For example, the defocus amount may be directly obtained from the parallax images by machine learning. The parallax images may be obtained by a plurality of imaging devices. Note that the parallax displacement amount and the defocus amount are signed quantities that can take positive and negative values. When the parallax displacement amount or the defocus amount is smaller than a predetermined value, it can be regarded as being in focus, and a map indicating the in-focus subject region is called a focus map.

[0021] Hereinafter, each embodiment will be described in detail.

[0022] [Embodiment 1] In this embodiment, an image estimation process is performed in which an image is input to a machine learning model and an image is output. The input image is divided into blocks to generate a plurality of divided images, and the image estimation process is performed by inputting each divided image to the machine learning model. In order to correct the blur of the in-focus plane, a sharpness process is applied to the input image using a machine learning model learned based on the optical characteristics of the in-focus plane. In order to apply the sharpness process only to the in-focus plane, only the necessary regions are input to the machine learning model based on the focus map indicating the in-focus plane.

[0023] However, since the input specifications of the machine learning model are determined by the model's architecture, training data, software implementation, or the implementation of the image processing circuit, it is necessary to divide the data to meet the input specifications. For example, there are specifications such as the position and shape for dividing the input image, the number of overlapping pixels between the divided images, the number of pixels in the divided images, and the range of pixel values. In particular, since the position and shape for dividing the input image are determined as specifications, it is not possible to input the data into the machine learning model according to the area to be processed, and it is not possible to process only the area to be processed.

[0024] Therefore, for each divided region (sub-image) when divided to meet the specifications of the machine learning model, an input divided image (input sub-image) to be input into the machine learning model is obtained based on whether or not the region to be processed is included. Any of the plurality of divided regions can be a region that can be obtained as an input divided image (sub-image) in the image estimation process by the machine learning model. If all the divided regions are obtained as input divided images, the processing load (computational load) is large. According to this embodiment, since there are divided regions that are not obtained as input divided images, the processing load can be reduced.

[0025] The input divided image may also include regions that do not necessarily need to be processed. Therefore, after being input into the machine learning model, the influence of the image estimation process on the region is reduced or removed. Even if data indicating the region to be processed is input into the machine learning model together with the input image, since the machine learning model is not rule-based processing, generally, it may not be learned to suppress the influence of the image estimation process as intended. Therefore, a process for reducing or removing the influence of the image estimation process is performed as post-processing on the output of the machine learning model.

[0026] Next, with reference to FIG. 1, the configuration of the imaging device of this embodiment will be described. Referring to FIG. 1, the imaging device 100 of this embodiment will be described. FIG. 1 is a block diagram showing the configuration of the imaging device 100. An image processing program for performing the edge enhancement process of this embodiment is installed in the imaging device 100. The edge enhancement process of this embodiment is executed by an image processing unit (image processing device) 104 inside the imaging device 100.

[0027] The imaging device 100 includes an optical system (imaging optical system) 101 and an imaging device main body (camera main body). The optical system 101 has a diaphragm 101a and a focus lens 101b, and is integrally configured with the camera main body. However, the present invention is not limited to this, and is also applicable to an imaging device in which the optical system 101 is detachably attached to the camera main body. Further, the optical system 101 may include an optical element having a diffractive surface or an optical element having a reflective surface in addition to an optical element having a refractive surface such as a lens.

[0028] The imaging element 102 has a CCD (Charge Coupled Device) sensor or a CMOS (Complementary Metal Oxide Semiconductor) sensor. The imaging element 102 photoelectrically converts the subject image (optical image formed by the optical system 101) formed through the optical system 101 to generate (output) a captured image (image data). That is, the subject image is converted into an analog signal (electrical signal) by photoelectric conversion by the imaging element 102. The A / D converter 103 converts the analog signal input from the imaging element 102 into a digital signal and outputs it to the image processing unit 104.

[0029] The image processing unit 104 performs predetermined processing on the digital signal and performs the edge enhancement processing of this embodiment. The image processing unit 104 includes an imaging condition acquisition unit 104a, an information acquisition unit 104b, and a processing unit 104c. The imaging condition acquisition unit 104a acquires the imaging conditions of the imaging device 100 from the state detection unit 107. The imaging conditions are the aperture value, the imaging distance (focus position), or the focal length of the zoom lens, etc. The state detection unit 107 may acquire the imaging conditions directly from the system controller 110 or from the optical system control unit 106.

[0030] The machine learning model that performs the edge enhancement processing can perform aberration correction processing to correct the blur caused by the optical characteristics of the in-focus plane by performing learning using the learning data generated based on the optical transfer function. The parameters of the machine learning model are held in the storage unit (storage means) 108. The storage unit 108 is composed of, for example, a ROM (Read Only Memory). The output image processed by the image processing unit 104 is stored in the image recording medium 109 in a predetermined format. On the display unit 105 composed of a liquid crystal monitor or an organic EL display, an image obtained by performing predetermined processing for display on the image subjected to the edge enhancement processing is displayed. However, the image displayed on the display unit 105 is not limited to this, and an image obtained by performing simple processing for high-speed display may be displayed on the display unit 105.

[0031] The system controller 110 controls the imaging device 100. The mechanical drive of the optical system 101 is performed by the optical system control unit 106 based on the instructions of the system controller 110. The optical system control unit 106 controls the aperture diameter of the aperture 101a so as to obtain a predetermined F number. In addition, since the optical system control unit 106 performs focus adjustment according to the subject distance, the position of the focus lens 101b is controlled by an autofocus (AF) mechanism (not shown) or a manual focus mechanism. Note that functions such as the aperture diameter control of the aperture 101a and manual focus do not have to be executed according to the specifications of the imaging device 100.

[0032] Note that optical elements such as a low-pass filter and an infrared cut filter may be arranged between the optical system 101 and the imaging device 102. However, when using an element that affects the optical characteristics of a low-pass filter or the like, consideration may be required at the time of creating the sharpening filter. Regarding the infrared cut filter as well, since it affects each OTF of the RGB channels, which is the integrated value of the optical transfer function (OTF) of the spectral wavelength, particularly the OTF of the R channel, consideration may be required at the time of creating the sharpening filter. Therefore, the sharpening filter may be changed according to the presence or absence of a low-pass filter or an infrared cut filter.

[0033] Note that the image processing unit 104 is composed of an ASIC, and the optical system control unit 106, the state detection unit 107, and the system controller 110 are each composed of a CPU or an MPU. Also, one or more of these image processing unit 104, optical system control unit 106, state detection unit 107, and system controller 110 may be configured to be shared by the same CPU or MPU.

[0034] Next, with reference to FIG. 2, the image processing method performed in the image processing unit 104 of the present embodiment will be described. FIG. 2 is a flowchart showing the image processing method in the present embodiment. The image processing method of the present embodiment is executed based on an instruction from the image processing unit 104. The sharpening process shown in FIG. 2 can be embodied as a program for causing a computer to execute the functions of each step. The same applies to the following flowcharts.

[0035] First, in step S101, the information acquisition unit 104b acquires the image (captured image) captured by the imaging device 100 as an input image (first image). The input image is stored in the storage unit 108. Also, the information acquisition unit 104b may acquire the image stored in the image recording medium 109 as the input image.

[0036] Subsequently, in step S102, the imaging condition acquisition unit 104a acquires the imaging conditions at the time of imaging the input image. The imaging conditions are the focal length of the optical system 101, the aperture value (F value), the imaging distance determined for the in-focus subject, and the like. In the case of an imaging device in which the lens is detachably mounted on the camera body, the imaging conditions further include the lens ID and the camera ID. The imaging conditions may be acquired directly from the imaging device or from information (for example, EXIF information) attached to the input image.

[0037] Subsequently, in S103, the information acquisition unit 104b acquires a machine learning model for performing sharpening processing based on the optical characteristics of the optical system 101 from the storage unit 108. The machine learning model of this embodiment includes processing by a neural network. Here, since the optical characteristics vary depending on the imaging conditions and the angle of view, it is also possible to configure such that information regarding the imaging conditions and the angle of view is also input to the machine learning model, or different machine learning models may be used according to the imaging conditions and the angle of view. The optical characteristics used for the sharpening processing are acquired based on the imaging distance (the distance to the in-focus position), which is one of the imaging conditions in particular. The optical characteristics are the optical characteristics when in-focus on the axis with respect to a subject on a plane perpendicular to the optical axis at a distance corresponding to the imaging distance.

[0038] Subsequently, in step S104, the information acquisition unit (determination means) 104b acquires the region of interest (ROI) to be subjected to the sharpening processing. That is, the information acquisition unit 104b determines the ROI from the captured image. In this embodiment, the ROI is acquired as a two-dimensional map based on the focus map. Specifically, the focus map acquired by the above method is acquired at the same resolution as the input image, and the subject region indicated as being in focus in the focus map is set as the ROI. This is because the sharpening processing is preferably performed in the in-focus region since it is based on the optical characteristics of the in-focus plane. The out-of-focus region is blurred due to defocusing and thus does not need to be sharpened. Also, in the out-of-focus region, the subject distance is different from the distance corresponding to the optical characteristics used in the learning of the machine learning model, and thus it is blurred with different optical characteristics.

[0039] The blur in the out-of-focus area is different from the blur assumed by the machine learning model. Therefore, due to the sharpening process, adverse effects such as overcorrection and the contour enhancement of defocus blur may occur. Therefore, it is better not to apply the sharpening process to the out-of-focus area under shooting conditions where adverse effects occur. Note that the focus map does not have to be the same resolution as the input image. After obtaining the focus map at a low resolution, it may be upscaled to the same resolution as the input image.

[0040] Subsequently, in step S105, the information acquisition unit (first acquisition means) 104b acquires one input divided image (second image) to be input to the machine learning model from a plurality of divided regions obtained by dividing the input image. That is, the information acquisition unit 104b acquires one or more input divided images corresponding to only a part of the plurality of partial images of the captured image. Here, with reference to FIG. 3, the method of acquiring the input divided image will be described in detail. FIG. 3 is a diagram showing the relationship between the region of interest and the input divided image in the present embodiment.

[0041] 201 is the input image, 202 is the two-dimensional map obtained in step S104 of FIG. 2, and the region of interest 205 is represented by the hatched area. 203 are a plurality of divided images obtained by dividing the input image 201 in a block shape, and each divided image is shown as partial regions a1 to a24. The division positions of the divided images 203 are represented by broken lines, and each of the partial regions surrounded by the broken lines corresponds to the acquisition region when each is acquired as a divided image. The division positions and division sizes of the divided images are implemented in the image processing unit 104 as an ASIC in advance. 204 is an image showing the superposition of the region of interest 205 and the division positions of the divided images 203.

[0042] From the image 204, the partial regions a14 and a20 are included inside the target region 205, and the entire partial regions a14 and a20 are subject to the sharpening process. On the other hand, the partial regions a7 - a9, a13, a15, a19, and a21 partially overlap with the target region 205 (having a first region including the target region and a second region not including the target region), and only a part of each partial region is subject to the sharpening process. The machine learning model performs processing by inputting in block units. For this reason, the input targets as the input divided images are the partial regions a7 - a9, a13 - a15, and a19 - a21, and in step S105 of FIG. 2, one of the partial regions a7 - a9, a13 - a15, and a19 - a21 is acquired. That is, it is necessary to acquire, as the input divided image, a partial region included in the target region 205 or a partial region that partially overlaps with the target region 205.

[0043] Note that for the divided regions not used as the input divided images (divided regions not including the target region 205), although the division positions and division sizes are defined, it is not necessary to be acquired as an image. Note that it is only necessary for the target region 205 to be processed using the machine learning model, and it is not necessary for all of the partial regions included in the target region 205 or the partial regions that partially overlap with the target region 205 to be acquired as the input divided image. This is because when partial regions overlap, even if a certain partial region overlapping with the target region is not input to the machine learning model, the entire target region can be input to the machine learning model by other partial regions being input to the machine learning model.

[0044] Subsequently, in step S106 of FIG. 2, the processing unit (first generation means) 104c inputs the input divided image acquired in step S105 to the machine learning model acquired in step S103, and generates a processed divided image (third image corresponding to the input divided image) to which the sharpening process is applied. When a partial region that partially overlaps with the target region is acquired as the input divided image, the processed divided image includes a region not included in the target region.

[0045] Subsequently, in step S107, the information acquisition unit 104b acquires the correction intensity to be applied to the processed divided image from the storage unit 108. In this embodiment, the correction intensity (first intensity) for the pixels included in the target area (the area corresponding to the first area) is 1, and the correction intensity (second intensity) for the pixels not included in the target area (the area corresponding to the second area) is 0. Here, in order to give the correction intensity spatial continuity, the first intensity for the pixels included in the target area may be less than 1, and the second intensity for the pixels not included in the target area may be a value greater than 0. The correction intensity may be gradually decreased toward the non-target area side around the boundary so as to be continuous at the boundary of the target area. In this case, the correction intensity is acquired as a two-dimensional map corresponding to the pixel positions of the processed image. Note that the correction intensity may be acquired as a scalar value according to whether it is a target area or not without acquiring it as a two-dimensional map. In this case, it is not necessary to acquire the correction intensity again when step S107 is executed again.

[0046] Subsequently, in step S108, the weighted average of the processed target image and the input divided image is taken (the processed target image and the input divided image are combined) according to the correction intensity, thereby reflecting the correction intensity of the sharpening process. If the correction intensity is 1, no weighted average is taken, and if the correction intensity is 0, it may be replaced with the input divided image. Note that if the correction intensity is 1 for all the pixels of the processed divided image, the process proceeds to step S109 without any processing.

[0047] Subsequently, in step S109, the processing unit 104c determines whether all the target partial areas of the input divided image have been processed. If not all the processing of the input divided image (partial area) has been completed, the process returns to step S105, the unprocessed partial area is acquired as the input divided image, and steps S106 to S108 are executed in the same manner. On the other hand, if the processing of all the input divided images (partial areas) has been completed, the process proceeds to step S110.

[0048] In step S110, the processing unit (second generation means) 104c arranges the plurality of processed divided images so as to have the positional relationship before division, and generates a combined (synthesized) output image (fourth image). That is, the processing unit 104c generates an output image by performing processing based on the processed divided image and the target region. When the divided regions overlap each other, they may be cut out so as not to overlap, or the overlapping portions may be combined by weighted averaging. Thereby, the entire image processing in this embodiment is completed.

[0049] As described above, in this embodiment, the input divided image has a first region including the target region and a second region not including the target region. In the region corresponding to the first region among the processed divided images, the pixel values are changed by setting the intensity of the processing by the machine learning model to the first intensity. On the other hand, in the region corresponding to the second region among the processed divided images, the pixel values are changed by setting the intensity of the processing to a second intensity smaller than the first intensity. The output image is obtained using the region to which the first intensity is applied and the region to which the second intensity is applied among the processed divided images.

[0050] Note that the processing order of this embodiment may be changed as appropriate. The acquisition of the shooting conditions in step S102 and the acquisition of the machine learning model in step S103 may be executed respectively until the machine learning model is used in step S106. When different machine learning models are used for each input divided image, a plurality of machine learning models may be acquired in step S103, or the machine learning models required sequentially may be acquired.

[0051] The acquisition of the correction intensity in step S107 may be performed before applying it in step S108. In this embodiment, the correction intensity is acquired for each input divided image. However, it may also be acquired for the entire captured image, and the portion corresponding to the input divided image may be extracted and used. The application of the correction intensity in step S108 only needs to be reflected in the output image, and it may also be applied after combining the processed divided images in step S110. At this time, instead of taking the weighted average of the processed target image and the input divided image, the correction intensity can be applied by taking the weighted average of the output image and the input image. That is, in this embodiment, the processed divided image or the output image is obtained by changing the pixel values based on the target area.

[0052] In this embodiment, the information acquisition unit 104b acquires the machine learning model and the correction intensity from the storage unit 108 of the imaging device 100, but it is not limited thereto. For example, in the case of an imaging device in which the optical system 101 is detachably attached to the camera body, the information acquisition unit 104b may acquire the optical characteristics and the correction intensity stored in the storage unit in the lens device including the optical system 101 through communication with the imaging device. In this case, the lens device (optical system 101) has a storage unit (not shown) for storing the optical characteristics and the correction intensity, and a communication unit (not shown) for transmitting the optical characteristics and the correction intensity to the camera body.

[0053] In addition to storing the machine learning model and the correction intensity in the storage unit in the imaging device or the lens device, they may be held in advance on the server. In this case, it is possible to download the machine learning model and the correction intensity by communicating with the imaging device or the lens device as needed.

[0054] In this embodiment, the correction intensity is acquired from the storage unit 108. However, an index value such as a defocus amount or a subject distance may be acquired from the storage unit 108, and the correction intensity may be separately converted from the index value. Thereby, even when the imaging device, the optical system, or the shooting conditions are different, different correction intensities can be obtained for the same index value. Note that the index value is preferably information related to the target area.

[0055] In this embodiment, the correction intensity is applied by taking the weighted average of the processed target image and the input divided image, but the application of the correction intensity is not limited thereto. Various modifications can be made, such as applying a blurring process again to the sharpened image.

[0056] The imaging device may output an input image to an image processing device provided separately from the imaging device, and the image processing device may be configured to perform image processing. In this case, each of the imaging condition information and the information regarding the region of interest can also be transferred from the imaging device to the image processing device directly or indirectly via communication. When transferring via communication, it is possible to appropriately select whether to attach each piece of information to the input image or not.

[0057] In this embodiment, the image estimation process applied to the input image is a sharpening process, but it is not limited thereto, and any other image process may be used as long as it is a process to be applied limited to a predetermined region of interest. For example, the output image may be an image to which defocus blurring conversion or resolution enhancement is applied to the input image.

[0058] Defocus blurring conversion is a process of changing the luminance distribution of a blurred image by controlling the optical characteristics of the defocus region. For example, it is to change the size and shape of the blur, to convert between a uniform blurred image by an ideal optical system and a blurred image with a decreasing light amount toward the peripheral part, or to control the change of the blurred image with respect to the defocus amount.

[0059] When applying defocus blur conversion, since the out-of-focus region becomes the region of interest, it is preferable to set the out-of-focus region as the region of interest based on the focus map. Also, in order to change only the blur of a predetermined defocus amount, the region of interest may be determined based on the defocus map. When applying high-resolution conversion, it is preferable to set only the in-focus region as the region of interest based on the focus map. In regions that do not include defocus regions or other high-frequency components, high-precision high-resolution conversion by a machine learning model has little effect, and known methods such as bicubic interpolation may be used for high-resolution conversion. As described above, the image estimation process is preferably a process that changes the frequency characteristics of the input image.

[0060] The process of changing the frequency characteristics of the input image varies in effect depending on the frequency characteristics of the input image and the optical characteristics of the imaging device that captured the input image. Therefore, by setting the target region according to the effect, the processing load can be efficiently reduced. It is more preferable to include a convolution process, which is a process corresponding to applying a gain to the frequency components, as the process of changing the frequency characteristics of the input image.

[0061] In this embodiment, the region of interest is determined based on the focus map, but it may also be determined based on other distance-related information. For example, it may be a defocus map corresponding to the distance on the image side of the optical system, or a parallax shift amount map corresponding to the defocus map. Furthermore, the distance-related information may be a depth map corresponding to the distance on the subject side. For example, regardless of in-focus or out-of-focus, the region of interest can be determined by selecting a range where these distance information values are constant. Also, in this embodiment, the region of interest may be determined using semantic region segmentation information (such as information indicating a person). Sharpening or high-resolution conversion may be applied only to the main subject such as a person as the region of interest, or a region excluding specific subjects such as an empty region that generally has low contrast may be used as the region of interest.

[0062] In this embodiment, the region of interest may be determined based on crop information (information indicating a crop region). Cropping is a process of cutting out only a predetermined range of an image, which is a process of acquiring only a predetermined range of the output of an imaging sensor as a recorded image during imaging, or a process of acquiring only a predetermined range of an input image during image processing. When performing image estimation processing on the input image before cropping in a state where cropping is known, the processing load can be reduced by setting the crop region as the region of interest.

[0063] In this embodiment, the region of interest may be determined based on the image circle information of the imaging optical system. For example, an image captured using a full-frame fisheye lens may include regions that are not imaged by the optical system, that is, regions outside the image circle. Therefore, for example, the region of interest may be set as only the region within the image circle.

[0064] In this embodiment, the region of interest may be obtained based on optical performance. For example, in sharpness enhancement, a region with particularly low optical performance may be set as the region of interest, or in high-resolution conversion, a region with high optical performance where high-precision estimation processing using a machine learning model is effective may be set as the region of interest. In defocus blur conversion, a blurred image with a blur effect that is not desired by the user may be set as the region of interest. As an example of a case where the blur effect is not desired, there may be a case where the luminance increases at the peripheral part, but the luminance distribution of the blurred image is determined based on optical characteristics. That is, only the imaging conditions and the angular field region of the optical characteristics that result in such a blurred image may be set as the region of interest. Also, the region of interest may be based on a plurality of the above-mentioned information. For example, only the in-focus region within the crop region may be set as the region of interest.

[0065] In this embodiment, the region of interest may be determined based on the content of the input image. By determining the region of interest based on the content of the input image, it is possible to efficiently reduce the processing load based on information that has not been determined in advance, such as before imaging or before image processing. Further, by using an image processing method for setting the region of interest based on the content of the input image, it is possible to perform image estimation processing with a low processing load without the need to implement so as to be able to perform image segmentation corresponding to the input image in software or an image processing circuit. Therefore, the implementation cost can be reduced. Here, the content refers to the imaged subject.

[0066] In this embodiment, the region of interest may be determined based on settings (changeable image processing settings or imaging settings) that can be changed during image processing or imaging. For example, it is based on a crop region or other region specified by the user as the target of image processing. It may also be based on the focus information at the time of shooting. These are not specific to the device, but are based on post-attached conditions such as user specifications and shooting methods. By obtaining the region of interest based on changeable settings, it is possible to efficiently reduce the processing load corresponding to these settings.

[0067] The divided regions are preferably divided regardless of the content of the input image, such as by equally dividing the input image. That is, it is preferable that the divided regions are not divided by designating the region of the subject of interest. As described above, the implementation cost can be reduced in software or an image processing circuit.

[0068] The divided regions are preferably divided coarser than the resolution of one pixel of the region of interest. That is, the region in the first image corresponding to one pixel in the information (for example, the focus map) used to determine the region of interest is smaller than the partial image. By increasing the size of each of the partial images input to the machine learning model, the processing speed can be improved. Further, by setting the region of interest finer than the divided regions, post-processing corresponding to the fine structure of the subject is possible even when each divided region is large, and the output image after image processing can be made of high quality.

[0069] The intensity of the processing by the machine learning model for the processed segmented image or the output image may be changed in pixel values so as to be smaller for the regions not included in the attention region than for the regions included in the attention region. That is, the processed segmented image or the output image may be obtained by changing the pixel values so that the intensity of the processing by the machine learning model for the regions not included in the attention region is smaller than the intensity of the processing by the machine learning model for the regions included in the attention region. As described above, by reducing the correction intensity outside the attention region, it is possible to perform image estimation processing that emphasizes the attention region. Note that the correction intensity in the attention region only needs to be statistically larger than the correction intensity in the non-attention region, and there may be a region in the attention region where the correction intensity is smaller than that in the non-attention region.

[0070] In this embodiment, a sharpening process is applied to the captured image, but the target of the image estimation process may be a map having two-dimensional pixel values that can be represented as an image. For example, a depth map, CG, or the like may be used. The divided regions are preferably regions that are uniformly divided. Thereby, the input data input to the machine learning model can be made to have a uniform size.

[0071] In this embodiment, the divided regions are divided without overlap, but they may be divided so that the divided regions overlap each other. Also in this embodiment, the image is a two-dimensional image. In order to input to a machine learning model involving convolution or the like, the divided regions are preferably rectangular. As the dividing method, it is preferably divided into blocks in the two-dimensional directions of horizontal and vertical. Also, this embodiment can be applied to a moving image. In a moving image, one attention region may be commonly used for frames with different times, or different attention regions may be used for each frame.

[0072] [Embodiment 2] Next, the imaging device in Embodiment 2 of the present invention will be described. In this embodiment, a resolution enhancement process is performed using the attention region set based on the crop region. Note that in this embodiment, the block diagram of the imaging device 100 is the same as that in Embodiment 1.

[0073] Referring to FIG. 4, the edge enhancement process in this embodiment will be described. FIG. 4 is a flowchart regarding the edge enhancement process (image estimation process) in this embodiment. Steps S201 and S202 in FIG. 4 are the same as steps S101 and S102 in FIG. 2 described in Embodiment 1, respectively. In step S203, different from step S103, the information acquisition unit 104b acquires a machine learning model for performing a resolution enhancement process.

[0074] Subsequently, in step S204, the information acquisition unit 104b acquires a region of interest to be subjected to the resolution enhancement process. In this embodiment, in addition to the resolution enhancement process, an image crop process is performed. Since only the region included in the region remaining after the crop process (crop region) is required as the output image, the crop region is set as the region of interest.

[0075] In this embodiment, the crop process and the resolution enhancement process are applied to the image stored in the storage unit 108 after imaging, but it is not limited thereto. For example, a predetermined region may be specified as the crop region at the time of shooting, and the resolution enhancement process may be applied simultaneously with imaging. The crop region does not necessarily have to be acquired as a two-dimensional map, and only the coordinates of the four corners of the rectangular region may be acquired. The crop region may be specified by the user, automatically set based on the specified zoom ratio, or automatically set based on the subject.

[0076] Subsequently, in step S205, the information acquisition unit 104b acquires an input divided image based on the region of interest, in the same manner as in step S105. FIG. 5 is a diagram showing the relationship between the region of interest and the input divided image in this embodiment.

[0077] In FIG. 5, 301 is an input image, and the hatched portion in image 302 is the region of interest 305 obtained in step S204. 303 is a divided image, and the division positions and partial regions of the divided image 303 are indicated by broken lines. 304 is an image obtained by superimposing the region of interest 305 and the divided image 303. What needs to be obtained as the input divided image is the partial region included in the region of interest 305 or the partial region that partially overlaps with the region of interest 305, that is, partial regions a1 to a3, a7 to a9, a13 to a15, and a19 to 21. One of these partial regions to which the image estimation process is not applied is obtained as the input divided image.

[0078] Subsequently, in step S206 of FIG. 4, the processing unit 104c inputs the input divided image obtained in step S205 into the machine learning model obtained in step S203 to obtain a processed divided image to which the resolution enhancement process is applied. When a partial region that partially overlaps with the region of interest is obtained as the input divided image, the processed divided image includes a region not included in the region of interest.

[0079] Subsequently, in step S207, the processing unit 104c determines whether all the processing of the partial regions targeted by the input divided image has been completed. If not all the processing of the targeted partial regions has been completed, it returns to step S205, obtains the unprocessed partial region as the input divided image, and executes step S206 in the same manner. On the other hand, if all the processing of the targeted partial regions has been completed, it proceeds to step S208.

[0080] In step S208, the processing unit 104c performs cropping (changing the number of pixels) on the processed divided image obtained by processing the partial region that partially overlaps with the region of interest. Specifically, for partial regions a1 to a3, a7, a9, a13, a15, and a19 to 21, the regions not included in the hatched portion (region of interest 305) in the image 304 shown in FIG. 5 are deleted. For the partial regions included inside the region of interest 305, there is no need to perform cropping processing because they do not include regions that are not the region of interest.

[0081] Subsequently, in step S209, the processing unit 104c arranges the plurality of processed and divided images so as to have the positional relationship before division, and obtains a combined output image. As a result, an output image in which only the target region in the input image has been upsampled is obtained, and the overall image processing in this embodiment is completed. Note that the order of steps S208 and S209 may be reversed, or only the target region may be cut out from the output image after combination. Steps S208 and S209 can also be processed simultaneously. That is, the output image may be generated by referring only to the pixels included in the target region among the processed and divided images. At this time, the output image is an image based only on the region included in the target region among the processed and divided images. Thus, in this embodiment, the processed and divided image or the output image is obtained by changing the number of pixels based on the target region.

[0082] (Other embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or apparatus via a network or a storage medium, and causing one or more processors in a computer of the system or apparatus to read and execute the program. Further, it can also be realized by a circuit (for example, ASIC) that realizes one or more functions. The image processing apparatus in the present invention may be any apparatus having the image processing function of the present invention, and can be realized in the form of an imaging apparatus or a PC.

[0083] According to each embodiment, it is possible to provide an image processing method, an image processing apparatus, an imaging apparatus, and a program that can reduce the processing load in image estimation processing using a machine learning model.

[0084] The disclosure of each embodiment includes the following configurations and methods. (Method 1) A step of determining a target region in a first image, A step of obtaining one or more second images corresponding to only a part of a plurality of partial images of the first image, A step of inputting the second image into a machine learning model and generating a third image corresponding to the second image, A step of generating a fourth image by a process based on the third image and the region of interest, The second image is an image processing method characterized by including at least a part of the region of interest. (Method 2) The region of interest is determined based on at least one of information regarding the distance corresponding to the first image, semantic region division information, crop information, image circle information, or optical performance, according to the image processing method described in Method 1. (Method 3) The region of interest is determined based on at least one of the content of the first image, changeable image processing settings, or changeable imaging settings, according to the image processing method described in Method 1. (Method 4) The information regarding the distance is a focus map, according to the image processing method described in Method 2. (Method 5) The plurality of partial images are obtained by equally dividing the first image, according to the image processing method described in any one of Methods 1 to 4. (Method 6) The fourth image is an image based only on the region included in the region of interest in the third image, according to the image processing method described in any one of Methods 1 to 5. (Method 7) The third image or the fourth image is generated by processing such that the intensity of the processing by the machine learning model for the region not included in the region of interest is smaller than the intensity of the processing by the machine learning model for the region included in the region of interest, according to the image processing method described in any one of Methods 1 to 5. (Method 8) The second image has a first region including the region of interest and a second region not including the region of interest, For the region corresponding to the first region in the third image, the intensity of the processing by the machine learning model is set to a first intensity, For the region corresponding to the second region in the third image, the intensity of the processing is set to a second intensity smaller than the first intensity. The fourth image is generated by using the region set to the first intensity and the region set to the second intensity among the third images, and is the image processing method according to any one of Methods 1 to 5. (Method 9) The fourth image is an image to which at least one of sharpness enhancement, high-resolution conversion, or defocus blur conversion is applied to the first image, and is the image processing method according to any one of Methods 1 to 8. (Configuration 1) Determining means for determining a region of interest in the first image, First acquisition means for acquiring one or more second images corresponding to only a part of a plurality of partial images of the first image, First generation means for inputting the second image into a machine learning model and generating a third image corresponding to the second image, Second generation means for generating a fourth image by processing based on the third image and the region of interest, and The second image includes at least a part of the region of interest, and is an image processing apparatus. (Configuration 2) An imaging apparatus having the image processing apparatus according to Configuration 1 and an imaging element. (Configuration 3) A program for causing a computer to execute the image processing method according to any one of Methods 1 to 9.

[0085] Although the preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of the gist thereof.

Explanation of Signs

[0086] 104 Image processing unit (image processing apparatus) 104b Information acquisition unit (determining means, first acquisition means) 104c Processing unit (first generation means, second generation means)

Claims

1. A step of determining a region of interest in the first image; A step of obtaining one or more second images corresponding to only a part of a plurality of partial images of the first image; A step of inputting the second image into a machine learning model and generating a third image corresponding to the second image; A step of generating a fourth image by a process based on the third image and the region of interest, and The second image includes at least a part of the region of interest. An image processing method characterized by this.

2. The region of interest is determined based on at least one of information regarding the distance corresponding to the first image, semantic region division information, crop information, image circle information, or optical performance. The image processing method according to Claim 1, characterized by this.

3. The region of interest is determined based on at least one of the content of the first image, changeable image processing settings, or changeable imaging settings. The image processing method according to Claim 1, characterized by this.

4. The information regarding the distance is a focus map. The image processing method according to Claim 2, characterized by this.

5. The plurality of partial images are obtained by equally dividing the first image. The image processing method according to Claim 1, characterized by this.

6. The fourth image is an image based on only the region included in the region of interest in the third image. The image processing method according to any one of Claims 1 to 5, characterized by this.

7. The third image or the fourth image is generated by processing such that the intensity of the processing by the machine learning model for the region not included in the region of interest is smaller than the intensity of the processing by the machine learning model for the region included in the region of interest. The image processing method according to any one of Claims 1 to 5, characterized by this.

8. The second image has a first region including the region of interest and a second region not including the region of interest. For the region corresponding to the first region in the third image, the intensity of the processing by the machine learning model is set to a first intensity. For the region corresponding to the second region in the third image, the intensity of the processing is set to a second intensity smaller than the first intensity. The fourth image is generated using the region set to the first intensity and the region set to the second intensity in the third image. The image processing method according to any one of Claims 1 to 5, characterized by this.

9. The image processing method according to any one of claims 1 to 5, wherein the fourth image is an image to which at least one of sharpening, high-resolution conversion, or defocus blur conversion is applied with respect to the first image.

10. Determining means for determining a region of interest in the first image; First acquisition means for acquiring one or more second images corresponding to only a part of a plurality of partial images of the first image; First generation means for inputting the second image into a machine learning model and generating a third image corresponding to the second image; Second generation means for generating a fourth image by processing based on the third image and the region of interest, The image processing apparatus, wherein the second image includes at least a part of the region of interest.

11. An imaging apparatus, comprising: the image processing apparatus according to claim 10; and an imaging element.

12. A program for causing a computer to execute the image processing method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Image processing method, image processing apparatus, image capturing apparatus, program, and storage medium

    JP2019212139A