Image processing method, image processing apparatus, imaging apparatus, and storage medium

By determining a specific area in the image processing, acquiring the corresponding partial images and processing it using a machine learning model, the problem of large amount of memory usage in image processing in the prior art is solved, and efficient image processing and resource management are realized.

CN120219694APending Publication Date: 2025-06-27CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411934999.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-26
Filing Date
2024-12-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively reduce memory usage in image processing, especially when only a specific area in the input image is desired to be processed.

Method used

By determining a specific area in the image, obtaining partial images corresponding to the area, processing these partial images using a machine learning model, and generating the final image based on the original image and the processed image.

Benefits of technology

This method can effectively reduce the computational load and storage requirements of image processing, and only process specific areas required, thereby improving processing efficiency and reducing resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219694A_ABST
    Figure CN120219694A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method, an image processing apparatus, an imaging apparatus, and a storage medium. The image processing method includes determining a specific region in a first image, acquiring at least one second image corresponding to some of a plurality of partial images of the first image, generating a third image corresponding to the second image by inputting the second image into a machine learning model, and generating a fourth image based on the first image and the third image, and the second image including at least a portion of the specific region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an image processing method, an image processing apparatus, a imaging apparatus, and a storage medium. Background Art

[0002] Conventionally, a method of reducing the memory usage by obtaining a plurality of segmented images from an input image and performing image estimation processing on each of the segmented images in image estimation processing using a machine learning model has been known. Japanese Patent Application Laid-Open No. 2019-212139 discloses a method of correcting a defocused image by segmenting an input image and performing processing.

[0003] In some cases, it is only desired to set a specific region in the input image as a specific region for image estimation processing. Summary of the Invention

[0004] An image processing method according to an aspect of the present disclosure includes determining a specific region in a first image; obtaining at least one second image corresponding to some of a plurality of partial images of the first image; generating a third image corresponding to the second image by inputting the second image into a machine learning model; and generating a fourth image based on the first image and the third image, wherein the second image includes at least a part of the specific region.

[0005] Other features of the present invention will become apparent from the following description of exemplary embodiments with reference to the accompanying drawings. Brief Description of the Drawings

[0006] Figure 1 is a block diagram of the imaging apparatus in each example.

[0007] Figure 2 is a flowchart showing the image processing in Example 1.

[0008] Figure 3 shows the relationship between the specific region and each input segmented image in Example 1.

[0009] Figure 4 is a flowchart showing the image processing in Example 2.

[0010] Figure 5 shows the relationship between the specific region and each input segmented image in Example 2.

[0011] Figure 6 illustrates the relationship between the light receiving unit of the image sensor and the pupil of the imaging optical system in each example.

[0012] Figure 7 illustrates the relationship between the light receiving unit of the image sensor and the subject in each example. Detailed Description of the Embodiments

[0013] In the following, the term "unit" may refer to a software environment, a hardware environment, or a combination of a software and a hardware environment. In a software environment, the term "unit" refers to a functionality, an application, a software module, a function, a routine, a set of instructions, or a program that can be executed by a programmable processor such as a microprocessor, a central processing unit (CPU), or a specially designed programmable device, or a controller. The memory contains instructions or programs that, when executed by the CPU, cause the CPU to perform operations corresponding to the unit or function. In a hardware environment, the term "unit" refers to a hardware element, a circuit, a component, a physical structure, a system, a module, or a subsystem. According to a specific embodiment, the term "unit" may include mechanical, optical, or electrical components, or any combination thereof. The term "unit" may include active (e.g., transistors) or passive (e.g., capacitors) components. The term "unit" may include semiconductor devices having a substrate and other material layers having various conductive concentrations. It may include a CPU or a programmable processor that can execute a program stored in the memory to perform a specified function. The term "unit" may include logic elements (e.g., AND, OR) implemented by transistor circuits or any other switching circuits. In a combination of a software and a hardware environment, the term "unit" or "circuit" refers to any combination of the software and the hardware environments as described above. In addition, the terms "element", "component", "part", or "device" may also refer to a "circuit" integrated or not integrated with a packaging material.

[0014] Referring now to the drawings, a detailed description of embodiments according to the present disclosure will be given. Corresponding elements in the respective figures will be denoted by the same reference numerals, and their repeated description will be omitted.

[0015] The image processing apparatus of the present embodiment performs a sharpening process on an input image captured using an optical system using a machine learning model based on the optical characteristics of the optical system (imaging optical system). The optical characteristics refer to the aberration of the optical system or the blur of the input image caused by the aberration, for example, a point spread function (PSF) or an optical transfer function (OTF). The optical characteristics may be, for example, a modulation transfer function (MTF) as an amplitude component of the OTF or a phase transfer function (PTF) as a phase component of the OTF. By performing the sharpening process based on the optical characteristics, the blur of the input image based on the characteristics of the imaging optical system can be effectively corrected.

[0016] Along with the description related to the configuration of the imaging device in each example, the sharpening process of this embodiment will be described below. In this embodiment, the sharpening process is an aberration correction process based on optical characteristics, and is also referred to as an image restoration process or a point image restoration process, for example. The optical characteristics may include an imaging lens and optical elements such as a low-pass filter or an infrared cut-off filter, and the influence due to the structure of the pixel array of the image sensor, for example, may be considered.

[0017] To train a machine learning model that performs a sharpening process (image estimation process) based on the optical characteristics of an optical system, the original image is used as the correct image, and a training image is generated by applying blur based on the optical transfer function to the original image corresponding to the subject. The machine learning model such as a neural network can be optimized by inputting the training image thereto to estimate the sharpened image and minimize the difference from the correct image. In this case, the correct image and the training image only need to have a relationship in which the blur caused by the optical characteristics can be corrected, and the presence of noise, the presence of development processing, and the resolution conditions may be different. In each example, in particular, the sharpening process is performed based on the optical characteristics at the focal plane to correct the blur on the focal plane.

[0018] Next, the principle of calculating the focus map using the disparity image will be described. The disparity image can be acquired by using an imaging unit that guides a plurality of light beams that have respectively passed through different regions in the pupil of one imaging optical system to different light receiving units (pixels) in one image sensor for photoelectric conversion. In other words, the disparity image required for distance calculation can be acquired by one imaging unit (including one optical system and one image sensor).

[0019] Figure 6 The relationship between the light receiving unit of the image sensor and the pupil of the imaging optical system in the imaging unit is shown. Reference numeral ML denotes a microlens, and reference numeral CF denotes a color filter. Reference numeral EXP denotes the exit pupil of the imaging optical system. Reference numerals G1 and G2 denote light receiving units (hereinafter each referred to as pixel G1 or pixel G2), and one pixel G1 is paired with one pixel G2. In the image sensor, a plurality of pairs of pixel G1 and pixel G2 (pixel pairs) are arranged. Each pair of pixel G1 and pixel G2 has an approximately conjugate relationship with the exit pupil EXP through a shared microlens ML (provided for each pixel pair). The plurality of pixel G1 arranged in the image sensor are also collectively referred to as pixel group G1, and the plurality of sub-pixels G2 arranged in the image sensor are also collectively referred to as pixel group G2.

[0020] Figure 7 The relationship between the light receiving unit of the image sensor and the subject is described, schematically showing the case where a thin lens is assumed to be arranged at Figure 6An imaging system in the case of the position of the exit pupil EXP. Each pixel G1 receives a light beam that has passed through the area P1 of the exit pupil EXP, and each pixel G2 receives a light beam that has passed through the area P2 of the exit pupil EXP. The reference numeral OSP is the object point being imaged. There does not need to be any physical object (subject) at the object point OSP, and the light beam that has passed through this point is incident on the pixel G1 or the pixel G2 according to the area (position) in the pupil through which the light beam has passed. The light beams passing through different areas in the pupil correspond to the separation of the incident light from the object point OSP based on the angle (parallax). Therefore, in the pixels G1 and G2 provided for each microlens ML, the images generated using the output signals from the pixel G1 and the images generated using the output signals from the pixel G2 become a plurality of (in this example, a pair of) parallax images with parallax therebetween.

[0021] In the following description, the reception of light beams that have passed through different areas in the pupil by different light-receiving units (pixels) is also referred to as pupil division. In this example, pupil division in one pupil division direction is described, but the number of pupil division directions and the number of divisions are optional. For example, division can be performed in two directions, the horizontal direction and the vertical direction. Alternatively, the number of pupil division directions and the number of divisions can vary according to the pixel positions in the image sensor. The parallax direction in pupil division is the position offset direction when each divided pupil is regarded as a viewpoint. The viewpoint position is defined based on the pupil area through which each light beam has passed and can be, for example, the centroid position of the light beam. The pupil division direction is not limited to the horizontal direction and the vertical direction and can also be an inclined direction.

[0022] In Figure 6 and Figure 7 , even when the above conjugate relationship becomes incomplete due to, for example, a position shift of the exit pupil EXP, or even when the areas P1 and P2 partially overlap each other, the multiple acquired images can be used as parallax images.

[0023] By specifying the corresponding subject area in the parallax images, the position offset amount (parallax offset amount) of the subject between the parallax images can be calculated. Various methods can be used to specify the same subject area in the images. For example, a block matching method can be used, which uses one of the parallax images as a reference image. Therefore, the parallax offset amount can be obtained. In each example, the parallax offset amount is considered with respect to the focal plane. The distance (defocus amount) from the image sensor to the focal point of the imaging optical system can be calculated based on the parallax offset amount and the position offset amount (baseline length) of the viewpoints. In addition, the subject distance can be calculated by using the focal length of the imaging optical system.

[0024] The acquisition of the parallax shift amount, defocus amount, and subject distance is not limited to the above methods. For example, the defocus amount can be directly obtained from the parallax image by machine learning. The parallax image can be acquired by a plurality of imaging devices. The parallax shift amount and the defocus amount are signed quantities that can take positive or negative values. In the case where the parallax shift amount or the defocus amount is less than a predetermined value, it can be considered in focus, and a map indicating the in-focus subject area is called a focus map.

[0025] Each example will be described in detail below.

[0026] Example 1

[0027] In this example, an image estimation process of outputting an image by inputting an image into a machine learning model is performed. The input image is divided into blocks to generate a plurality of divided images, and the image estimation process is performed by inputting each divided image into the machine learning model. In order to correct the blur on the focal plane, a sharpening process is applied to the input image by using a machine learning model trained based on the optical characteristics of the focal plane. Since the sharpening process is applied only at the focal plane, only the necessary region is input into the machine learning model based on the focus map indicating the focal plane.

[0028] However, the input specifications of the machine learning model are determined by the model architecture, learning data, software implementation, or image processing circuit implementation, so partitioning is required to meet the input specifications. For example, the specifications include the position and shape of the division of the input image, the number of overlapping pixels of the divided image, and the number and pixel value range of the pixels of the divided image. In particular, since the position and shape of the division of the input image are determined as specifications, the accuracy of processing only the processing target region can be improved by inputting it into the machine learning model according to the processing target region.

[0029] Therefore, in the case of performing division to meet the specifications of the machine learning model, any divided image (input divided image) to be input into the machine learning model is acquired from among the divided regions (partial images) based on whether the divided image includes the region to be the processing target. Each divided region is a region that can be acquired as an input divided image (partial image) in the image estimation process using the machine learning model. When all the divided regions are acquired as input divided images, the processing load (computational load) is large. According to this example, since some divided regions are not acquired as input divided images, the processing load can be reduced.

[0030] The input segmented image may include regions that are not targets for processing. Therefore, after being input into the machine learning model, the influence of the image estimation process on the regions is reduced or eliminated. Even when data representing the regions to be processed as targets is input into the machine learning model together with the input image, generally, since the machine learning model is not rule-based processing, the machine learning model can be trained to reduce the influence of the image estimation process as expected. Therefore, as post-processing, a process of reducing or eliminating the influence of the image estimation process is performed on the output of the machine learning model.

[0031] The following will refer to Figure 1 Describe the structure of the imaging device of this example. The following will refer to Figure 1 Describe the imaging device 100 of this embodiment. Figure 1 is a block diagram showing the structure of the imaging device 100. An image processing program for performing the sharpening process of this example is installed on the imaging device 100. The sharpening process of this embodiment is executed by the image processing unit (image processing device) 104 in the imaging device 100.

[0032] The imaging device 100 includes an optical system (imaging optical system) 101 and an imaging device body (camera body). The optical system 101 includes an aperture 101a and a focusing lens 101b and is integrated with the camera body. However, the present disclosure is not limited thereto, but also applies to an imaging device in which the optical system 101 is detachably mounted on the camera body. In addition to optical elements having refractive surfaces (such as lenses), the optical system 101 may also include optical elements having diffractive surfaces, optical elements having reflective surfaces, and the like.

[0033] The image sensor 102 includes a charge-coupled device (CCD) sensor or a complementary metal oxide semiconductor (CMOS) sensor. The image sensor 102 generates (outputs) a captured image (image data) by photoelectrically converting the subject image (optical image formed by the optical system 102) formed by the optical system 101. Specifically, the subject image is converted into an analog signal (electrical signal) by the photoelectric conversion of the image sensor 102. The A / D converter 103 converts the analog signal input from the image sensor 102 into a digital signal and outputs the digital signal to the image processing unit 104.

[0034] The image processing unit 104 performs a predetermined process on the digital signal and also performs the sharpening process of this example. The image processing unit 104 includes an imaging condition acquisition unit 104a, an information acquisition unit 104b, and a processing unit 104c. The imaging condition acquisition unit 104a acquires the imaging conditions of the imaging device 100 from the state detector 107. The imaging conditions include, for example, the aperture value (F-number), the shooting distance (focus position), and the focal length of the zoom lens. The state detector 107 can acquire the imaging conditions directly from the system controller 110 or can acquire the imaging conditions from the optical system control unit 106.

[0035] The machine learning model that performs the sharpening process can be trained by using learning data generated based on the optical transfer function to perform an aberration correction process that corrects the blur caused by the optical characteristics of the focal plane. The parameters of the machine learning model are stored in the memory 108. The memory 108 is constituted by, for example, a read-only memory (ROM). The output image processed by the image processing unit 104 is stored in the image recording medium 109 in a predetermined format. The display unit 105 is constituted by a liquid crystal monitor or an organic EL display and displays an image obtained by performing a predetermined display process on the sharpened image. However, the image displayed on the display unit 105 is not limited thereto, and an image that has been simplified for high-speed display can be displayed on the display unit 105.

[0036] The system controller 110 controls the imaging device 100. The optical system control unit 106 performs mechanical driving of the optical system 101 based on an instruction from the system controller 110. The optical system control unit 106 controls the opening diameter of the aperture 101a to achieve a predetermined F-number. In order to focus according to the subject distance, the optical system control unit 106 controls the position of the focusing lens 101b by using an autofocus (AF) mechanism or a manual focusing mechanism (not shown). Depending on the specifications of the imaging device 100, it is not always necessary to perform functions such as controlling the opening diameter of the aperture 101a and manual focusing.

[0037] Optical elements such as a low-pass filter or an infrared cut-off filter may be arranged between the optical system 101 and the image sensor 102, but when using an element that affects the optical characteristics (such as a low-pass filter), it may be necessary to consider this when producing the sharpening filter. Regarding the infrared cut-off filter, it affects the OTF of each RGB channel, that is, the integrated value of the spectral wavelength optical transfer function (OTF), especially the OTF of the R channel, so it may be necessary to consider this when producing the sharpening filter. Therefore, the sharpening filter can be changed according to the presence of the low-pass filter or the infrared cut-off filter.

[0038] The image processing unit 104 is composed of an ASIC, and the optical system control unit 106, the state detector 107, and the system controller 110 are each composed of a CPU or an MPU. One or more of the image processing unit 104, the optical system control unit 106, the state detector 107, and the system controller 110 may be composed of the same CPU or MPU.

[0039] Next, reference will be made to Figure 2 describe the image processing method performed by the image processing unit 104 of this example. Figure 2 is a flowchart showing the image processing method of this example. The image processing method of this example is executed based on an instruction from the image processing unit 104. Figure 2 The sharpening process shown in can be embodied as a computer program for causing a computer to execute each step. The same applies to the following flowcharts.

[0040] First, in step S101, the information acquisition unit 104b acquires an image (captured image) captured by the imaging device 100 as an input image (first image). The input image is stored in the memory 108. Alternatively, the information acquisition unit 104b may acquire an image stored in the image recording medium 109 as an input image.

[0041] Subsequently, in step S102, the imaging condition acquisition unit 104a acquires the imaging conditions at the time of capturing the input image. The imaging conditions include the focal length of the optical system 101, the aperture value (F-number), and the imaging distance determined for focusing the subject. In the case of an imaging device in which the lens is detachably mounted on the camera body, the imaging conditions also include the lens ID and the camera ID. The imaging conditions can be acquired directly from the imaging device or from information associated with the input image (such as EXIF information).

[0042] Subsequently, in S103, the information acquisition unit 104b acquires a machine learning model for sharpening processing from the memory 108 based on the optical characteristics of the optical system 101. The machine learning model of this example includes processing by a neural network. Since the optical characteristics vary depending on the imaging conditions and the viewing angle, information about the imaging conditions and the viewing angle can be additionally input to the machine learning model, and different machine learning models can be used according to the imaging conditions and the viewing angle. In particular, the optical characteristics for sharpening processing are acquired based on the imaging distance (distance to the focus position) as the imaging condition. The optical characteristics are the optical characteristics when the subject separated by the imaging distance and located on a plane orthogonal to the optical axis is focused on the axis.

[0043] Subsequently, in step S104, the information acquisition unit (determination unit) 104b acquires a specific area (area of interest) that is the target of the sharpening process. In other words, the information acquisition unit 104b determines a specific area from the captured image. In this example, a specific area is acquired as a two-dimensional map based on the focus map. Specifically, the focus map obtained by the above method is acquired at the same resolution as the input image, and the subject area indicated as being in focus on the focus map is set as the specific area. This is because, since the sharpening process is based on the optical characteristics of the focal plane, the sharpening process can be performed in the focused area. The out-of-focus area is blurred due to defocus, so sharpening is not required. In the out-of-focus area, the subject distance is different from the distance corresponding to the optical characteristics used in the machine learning model training, resulting in blurring with different optical characteristics.

[0044] The blurring in the out-of-focus area and the blurring assumed by the machine learning model are different from each other. Therefore, the sharpening process may cause adverse effects such as overcorrection and emphasizing the defocus blur around the contour. Therefore, under the image capture conditions where such adverse effects occur, it is best not to apply the sharpening process to the out-of-focus area. The focus map does not need to have the same resolution as the input image. After being acquired at a low resolution, the focus map can be enlarged to the same resolution as the input image.

[0045] Subsequently, in step S105, the information acquisition unit (first acquisition unit) 104b acquires one input segmentation image (second image) to be input to the machine learning model from among the multiple segmentation regions obtained by segmenting the input image. Specifically, the information acquisition unit 104b acquires at least one input segmentation image (one or more input segmentation images) corresponding to only a part (or some) of the multiple partial images of the captured image. The method of acquiring the input segmentation image will be described in detail below with reference to Figure 3 The method of acquiring the input segmentation image will be described in detail. Figure 3 The relationship between the specific area and each input segmentation image in this example is shown.

[0046] Reference numeral 201 denotes the input image, reference numeral 202 denotes the two-dimensional map acquired in step S104 in Figure 2 and the specific area 205 is shaded. Reference numeral 203 denotes the multiple segmentation images obtained by dividing the input image 201 into blocks, and the segmentation images are represented as partial areas a1 to a24. The segmentation positions of the segmentation images 203 are shown by dashed lines, and the partial areas surrounded by the dashed lines correspond to the respective acquisition areas when the images are acquired as segmentation images. The segmentation positions and segmentation sizes of the segmentation images are pre-implemented as ASICs in the image processing unit 104. Reference numeral 204 denotes an image showing the specific area 205 and the segmentation positions of the segmentation images 203 in a superimposed manner.

[0047] In the image 204, partial regions a14 and a20 are included within a specific region 205, and the entire partial regions a14 and a20 are targets for sharpening processing. Partial regions a7 to a9, a13, a15, a19, and a21 partially overlap with the specific region 205 (including a first region and a second region, where the first region includes the specific region and the second region does not include the specific region), and only a part of each partial region is a target for sharpening processing. The machine learning model processes the input in chunks. Thus, the input targets for the input segmentation image are partial regions a7 to a9, a13 to a15, and a19 to a21, and one of the partial regions a7 to a9, a13 to a15, and a19 to a21 is obtained in Figure 2 step S105. In other words, it is necessary to obtain a partial region included in the specific region 205 or a partial region that partially overlaps with the specific region 205 as the input segmentation image.

[0048] The segmentation regions not used as the input segmentation image (the segmentation regions that do not include the specific region 205) have defined segmentation positions and segmentation sizes, but do not need to be acquired as images. It is sufficient to process the specific region 205 using the machine learning model, and it is not necessary to acquire all the partial regions included in the specific region 205 or the partial regions that partially overlap with the specific region 205 as the input segmentation image. This is because, in the case of partial region overlap, even if a certain partial region that overlaps with the specific region is not input into the machine learning model, the entire specific region can be input into the machine learning model by inputting other partial regions.

[0049] Subsequently, in Figure 2 step S106, the processing unit (first generation unit) 104c inputs the input segmentation image obtained in step S105 into the machine learning model obtained in step S103, and generates a processed segmentation image (a third image corresponding to the input segmentation image) that has been sharpened. In the case of obtaining a partial region that partially overlaps with the specific region as the input segmentation image, the processed segmentation image includes regions that are not included in the specific region.

[0050] Subsequently, in step S107, the information acquisition unit 104b acquires the correction intensity applied to the processed segmented image from the memory 108. In this example, the correction intensity (first intensity) for the pixels included in the specific region (the region corresponding to the first region) is 1, and the correction intensity (second intensity) for the pixels not included in the specific region (the region corresponding to the second region) is 0. To provide spatial continuity of the correction intensity, the first intensity for the pixels included in the specific region may be less than 1, and the second intensity for the pixels not included in the specific region may be greater than 0. To ensure continuity at the boundary of the specific region, the correction intensity may gradually decrease towards the non-specific region near the boundary. In this case, the correction intensity is acquired as a two-dimensional map corresponding to the pixel positions of the processed image. The correction intensity may be acquired as a scalar value based on whether it is in the specific region, instead of acquiring the correction intensity as a two-dimensional map. In this case, when step S107 is executed, it is not necessary to acquire the correction intensity again.

[0051] Subsequently, in step S108, the correction intensity of the sharpening process is reflected by performing a weighted average of the processed target image and each input segmented image (combining the processed target image and each input segmented image) according to the correction intensity. In the case where the correction intensity is 1, the weighted average may not be performed, or when the correction intensity is 0, the input segmented image may be used for replacement. In the case where the correction intensity of all pixels of the processed segmented image is 1, no processing is performed, and the process proceeds to step S109.

[0052] Subsequently, in step S109, the processing unit 104c determines whether all partial regions to be set as the input segmented image have been processed. In the case where the processing of all input segmented images (partial regions) is not completed, the process returns to step S105 to acquire the unprocessed partial region as the input segmented image, and steps S106 to S108 are executed in the same manner. In the case where the processing of all input segmented images (partial regions) is completed, the process proceeds to step S110.

[0053] In step 110, the processing unit (second generation unit) 104c arranges the processed segmented images according to the positional relationship before segmentation and generates a connected (combined) output image (fourth image). In other words, the processing unit 104c generates an output image based on the processed segmented image and the processing of the specific region. When the segmented regions overlap, the segmented regions may be extracted to avoid overlap, or they may be connected by performing a weighted average on their overlapping parts. Thus, the entire image processing in this example is completed.

[0054] As described above, the input segmentation image in this example includes a first region including a specific region and a second region not including the specific region. In the region corresponding to the first region in the processed segmentation image, the pixel values are changed by setting the processing intensity of the machine learning model to a first intensity. In the region corresponding to the second region in the processed segmentation image, the pixel values are changed by setting the processing intensity to a second intensity smaller than the first intensity. Then, an output image is obtained by using the region to which the first intensity is applied and the region to which the second intensity is applied in the processed segmentation image.

[0055] The processing order of this example can be appropriately changed. The acquisition of the image capture conditions in step S102 and the acquisition of the machine learning model in step S103 only need to be executed until the machine learning model is used in step S106. In the case of using different machine learning models for each input segmentation image, multiple machine learning models can be acquired in step S103, or the necessary machine learning model can be acquired as needed each time.

[0056] The acquisition of the correction intensity in step S107 only needs to be executed before it is applied to step S108. In this example, the correction intensity is acquired for each input segmentation image, but it can also be acquired for the entire captured image, and the part corresponding to each input segmentation image can be extracted and used. The application of the correction intensity in step S108 only needs to be reflected on the output image, and it can be executed after the processed segmentation images are joined in step S110. In this case, the correction intensity can be applied by performing a weighted average on the output image and the input image, rather than on the processed target image and each input segmentation image. In other words, in this example, the processed segmentation image or the output image is obtained by changing the pixel values based on the specific region.

[0057] In this example, the information acquisition unit 104b acquires the machine learning model and the correction intensity from the memory 108 of the imaging device 100, but the present invention is not limited thereto. For example, in the case of an imaging device in which the optical system 101 is detachably mounted on the camera body, for example, in the case of an imaging device in which the optical system 101 is detachably mounted on the camera body, the information acquisition unit 104b can acquire the optical characteristics and the correction intensity stored in the memory of the lens device including the optical system 101 through communication with the imaging device. In this case, the lens device (optical system 101) includes a memory (not shown) for storing the optical characteristics and the correction intensity, and a communication unit (not shown) for transmitting the optical characteristics and the correction intensity to the camera body.

[0058] The machine learning model and the correction intensity can be pre - maintained in the server instead of being stored in the memory of the imaging device or the lens device. In this case, the machine learning model and the correction intensity can be downloaded to the imaging device or the lens device through communication as needed.

[0059] In this example, the correction intensity is obtained from the memory 108, but an index value such as a defocus amount or an object distance can also be obtained from the memory 108 and then converted into the correction intensity. Therefore, under different imaging devices, optical systems, or image - taking conditions, the same index value can obtain different correction intensities. The index value can be information about a specific area.

[0060] In this example, the correction intensity is applied by weighted - averaging the processed target image and each input segmented image, but the application of the correction intensity is not limited to this. Various modifications can be made, such as applying a blur process again to the image to which a sharpening process has been applied.

[0061] The input image can be output from the imaging device to an image - processing device disposed separately from the imaging device and can be subjected to image processing by the image - processing device. In this case, the imaging condition information and the information about each specific area can be directly or indirectly handed over from the imaging device to the image - processing device through communication. In the case of handing over through communication, it can be appropriately selected whether each piece of information is associated with the input image.

[0062] In this example, the image estimation process applied to the input image is a sharpening process, but the present invention is not limited to this, and any other image processing can be applied as long as the process is only applied to a predetermined specific area. For example, the output image can be an image obtained by applying a defocus blur conversion or a resolution enhancement to the input image.

[0063] The defocus blur conversion is a process of changing the brightness distribution of a blurred image by controlling the optical characteristics of the defocus area. The defocus blur conversion includes, for example, changing the size and shape of the blur, converting a uniformly blurred image through an ideal optical system and a blurred image with the light amount decreasing towards the peripheral part, and controlling the change of the blurred image with respect to the defocus amount.

[0064] In the case of applying defocus blur conversion, the out-of-focus area is a specific area. Therefore, the out-of-focus area can be set as a specific area based on the focus map. Since only the blur with a predetermined defocus amount is changed, the specific area can be determined based on the defocus map. In the case of applying resolution enhancement, only the in-focus area can be set as a specific area based on the focus map. In the out-of-focus area and other areas that do not contain high-frequency wave components, the effect of high-precision resolution enhancement by a machine learning model is small, and resolution enhancement by a known method such as bicubic interpolation is sufficient. As described above, the image estimation process can be a process of changing the frequency characteristics of the input image.

[0065] The effect of the process of changing the frequency characteristics of the input image changes according to the frequency characteristics of the input image and the optical characteristics of the imaging device that has captured the input image. Therefore, the processing load can be effectively reduced by setting the target area according to the effect. The process of changing the frequency characteristics of the input image can include convolution processing, which is a process equivalent to multiplying the frequency components by a gain.

[0066] In this example, the specific area is determined based on the focus map, but it can also be determined based on any other information regarding distance. For example, the determination can be based on a defocus map corresponding to the imaging-side distance of the optical system, or a parallax offset map corresponding to the defocus map. The information regarding distance can be a depth map corresponding to the subject-side distance. For example, the specific area can be determined by selecting a range where this distance information is constant, regardless of focus. In this example, the specific area can be determined by using semantic region division information (such as information indicating a person). Sharpening and resolution enhancement can be applied only to the main subject, such as a person, as the specific area, and an area excluding the specific subject (such as a typical low-contrast area like the sky) can be set as the specific area.

[0067] In this example, the specific area can be determined based on cropping information (information indicating the cropping area). Cropping is a process of extracting only a predetermined range of the image and is a process of obtaining a predetermined range of the output from the imaging sensor as a recorded image during imaging, or a process of obtaining only a predetermined range of the input image during image processing. When performing the image estimation process on the input image to be cropped when it is known that cropping will be performed, the processing load can be reduced by setting the cropping area as the specific area.

[0068] In this example, the specific area can be determined based on the image circle information regarding the imaging optical system. For example, an image captured using a full-circle fisheye lens includes an area that does not depend on the imaging of the optical system. In other words, the captured image includes an area outside the image circle. Therefore, for example, the specific area can be set only to the area within the image circle.

[0069] In this example, a specific area can be obtained based on optical performance. For example, an area with particularly low optical performance in sharpening can be set as the specific area, and an area with high optical performance and effective for high-precision estimation using a machine learning model in resolution enhancement can be set as the specific area. In defocus blur conversion, a blurred image with blur that the user does not want can be set as the specific area. Examples of cases where blur is not desired include cases where there is a large brightness distribution in the peripheral part, but the brightness distribution of the blurred image is determined based on optical characteristics. In other words, an area where such a blurred image is caused only by the imaging conditions and viewing angle of the optical characteristics can be set as the specific area. The setting of the specific area can be based on multiple pieces of information among the above-mentioned information. For example, only the in-focus area in the cropped area can be set as the specific area.

[0070] In this example, a specific area can be determined based on the content of the input image. In the case of determining a specific area based on the content of the input image, the processing load can be effectively reduced based on information that is not determined in advance, such as before imaging or before image processing. In addition, by using an image processing method that sets a specific area based on the content of the input image, an image estimation process with a low processing load can be performed without the need to implement software or an image processing circuit capable of image segmentation according to the input image. Therefore, the implementation cost can be reduced. The content refers to the subject being imaged.

[0071] In this example, a specific area can be determined based on settings that are variable during image processing or imaging (variable image processing settings or imaging settings). This determination is based on the cropped area or any other area specified by the user as the target of image processing. Alternatively, this determination can be based on the focus information at the time of image capture. These settings are not inherent to the device but are based on later conditions such as user specifications and image capture methods. By obtaining a specific area based on variable settings, the processing load can be effectively reduced according to the settings.

[0072] A segmented area can be obtained by segmentation regardless of the content of the input image, such as dividing the input image equally. In other words, the segmented area can be segmented without specifying an area of interest in the subject. As described above, this can reduce the implementation cost of software or an image processing circuit.

[0073] The segmented area can be coarser than the resolution of one pixel in the specific area. In other words, the area in the first image corresponding to one pixel in the information (e.g., focus map) used to determine the specific area is smaller than the partial image. By increasing the size of each partial image input to the machine learning model, the processing speed can be increased. In addition, by setting the specific area to be finer than the segmented area, post-processing corresponding to the fine structure of the subject can be performed in the case where each segmented area is large, thereby improving the quality of the output image after image processing.

[0074] The processing intensity of the machine learning model for each processed segmented image or output image can be set such that the pixel values of the regions not included in the specific region are changed to be smaller than the pixel values of the regions included in the specific region. In other words, each processed segmented image or output image can be obtained by changing the pixel values such that the processing intensity of the machine learning model for the regions not included in the specific region is less than the processing intensity of the machine learning model for the regions included in the specific region. As described above, by reducing the correction intensity of the regions other than the specific region, the image estimation process can be preferentially performed on the specific region. The correction intensity of the specific region only needs to be statistically greater than the correction intensity of the non-specific region, and there can be regions within the specific region where the correction intensity is less than that of the non-specific region.

[0075] In this example, the sharpening process is applied to the captured image, but the target of the image estimation process only needs to be a graph with two-dimensional pixel values that can be represented as an image. For example, the target can be a depth map or CG. The segmented regions can be uniformly segmented regions. Therefore, input data of a consistent size can be input into the machine learning model.

[0076] In this example, the segmented regions do not overlap, but they can overlap with each other. In this example, the image is a two-dimensional image. Since it is to be input into a machine learning model involving operations such as convolution, the segmented regions can be quadrilaterals. The segmentation can be performed in blocks in two two-dimensional directions, the horizontal direction and the vertical direction. This example also applies to moving images. In a moving image, a specific region can be commonly used for frames at different time points, or different specific regions can be used for each frame.

[0077] Example 2

[0078] The imaging device in Example 2 of the present disclosure will be described below. In this example, resolution enhancement processing is performed by using a specific region set based on a cropped region. The same block diagram as in Example 1 applies to the imaging device 100 in this example.

[0079] The following will refer to Figure 4 Describe the sharpening process in this example. Figure 4 is a flowchart of the sharpening process (image estimation process) in this example. Figure 4 Steps S201 and S202 in Figure 2 are the same as steps S101 and S102 described above in Example 1. In step S203, different from step S103, the information acquisition unit 104b acquires the machine learning model for performing resolution enhancement processing.

[0080] Subsequently, in step S204, the information acquisition unit 104b acquires a specific area as the target for resolution enhancement processing. In this example, in addition to the resolution enhancement processing, image cropping processing is also performed. Since only the area included in the remaining area (cropping area) after the cropping processing needs to be used as the output image, the cropping area is set as the specific area.

[0081] In this example, the cropping processing and the resolution enhancement processing are applied to the image stored in the memory 108 after imaging, but the present invention is not limited thereto. For example, a predetermined area can be designated as the cropping area at the time of image capture, and the resolution enhancement processing can be applied while imaging. The cropping area does not necessarily need to be acquired as a two-dimensional image, and only the coordinates of the four corners of the quadrilateral area can be acquired. The cropping area can be specified by the user or automatically set based on the specified magnification. Alternatively, the cropping area can be automatically set based on the subject.

[0082] Subsequently, in step S205, similar to step S105, the information acquisition unit 104b acquires the input segmented image based on the specific area. Figure 5 The relationship between the specific area and each input segmented image in this example is shown.

[0083] Figure 5 In the figure, reference numeral 301 denotes the input image, the shaded part of the image 302 is the specific area acquired in step S204, which is denoted by 305. Reference numeral 303 denotes the segmented image, and the segmentation position and partial area of the segmented image 303 are indicated by dashed lines. Reference numeral 304 denotes the image obtained by superimposing the specific area 305 and the segmented image 303. The area that needs to be acquired as the input segmented image is the partial area included in the specific area 305 or the partial area that partially overlaps with the specific area 305. In other words, the partial areas a1 to a3, a7 to a9, a13 to a15, and a19 to a21. Among these partial areas, one partial area to which the image estimation processing has not been applied is acquired as the input segmented image.

[0084] Subsequently, in Figure 4 step S206 therein, the processing unit 104c inputs the input segmented image acquired in step S205 into the machine learning model acquired in step S203, and acquires the processed segmented image to which the resolution enhancement processing has been applied. In the case where a partial area that partially overlaps with the specific area is acquired as the input segmented image, the processed segmented image includes an area that is not included in the specific area.

[0085] Subsequently, in step S207, the processing unit 104c determines whether the processing of all partial regions to be set as the input divided image is completed. In the case where the processing of all partial regions to be set is not completed, the processing returns to step S205 to obtain an unprocessed partial region as the input divided image, and step S206 is executed in the same manner. In the case where the processing of all partial regions to be set is completed, the processing proceeds to step S208.

[0086] In step S208, the processing unit 104c extracts (changes the number of pixels) the processed divided image obtained by processing the partial regions that partially overlap with the specific region. Specifically, regions that are not included in the shaded portion (specific region 305) in the illustrated image 304 are deleted from partial regions a1 to a3, a7, a9, a13, a15, and a19 to a21. Figure 5 Regions that are not included in the specific region are not included in the partial regions included in the specific region 305, and thus no extraction processing is required.

[0087] Subsequently, in step S209, the processing unit 104c arranges the processed divided image in the positional relationship before division and obtains the connected output image. Thus, an output image that only enhances the resolution of the specific region in the input image is obtained, thereby completing all the image processing in this example. The order of step S208 and step S209 can be reversed, and only the specific region can be extracted from the connected output image. Step S208 and step S209 can be processed simultaneously. In other words, in the processed divided image, only the pixels included in the specific region can be referred to generate the output image. In this case, the output image is only based on the region included in the specific region in the processed divided image. In this way, in this example, the processed divided image or the output image is obtained by changing the number of pixels based on the specific region.

[0088] Other embodiments

[0089] Embodiments of the present invention can also be implemented by a computer of a system or apparatus that reads and executes computer-executable instructions (e.g., one or more programs) recorded on a storage medium (which may also be more fully referred to as a "non-transitory computer-readable storage medium") to perform one or more functions of the above-described embodiments and / or includes one or more circuits (e.g., an application specific integrated circuit (ASIC)) for performing one or more functions of the above-described embodiments, and an embodiment of the present invention can be implemented by a method of, for example, reading and executing the computer-executable instructions from the storage medium by the computer of the system or apparatus to perform one or more functions of the above-described embodiments and / or controlling the one or more circuits to perform one or more functions of the above-described embodiments. The computer may include one or more processors (e.g., a central processing unit (CPU), a microprocessing unit (MPU)), and may include a network of separate computers or separate processors to read and execute the computer-executable instructions. The computer-executable instructions may be provided to the computer, for example, from a network or the storage medium. The storage medium may include, for example, a hard disk, a random access memory (RAM), a read-only memory (ROM), a memory of a distributed computing system, an optical disk (such as a compact disc (CD), a digital versatile disc (DVD), or a Blu-ray disc (BD) TM ), a flash device, and a memory card, among others.

[0090] Embodiments of the present invention can also be implemented by the following method, that is, by providing software (a program) that performs the functions of the above-described embodiments to a system or apparatus via a network or various storage media, and a method in which a computer or a central processing unit (CPU), a microprocessing unit (MPU) of the system or apparatus reads and executes the program.

[0091] Although the present invention has been described with reference to exemplary embodiments, it should be understood that the present invention is not limited to the disclosed exemplary embodiments. The scope of the appended claims should be given the broadest interpretation so as to cover all such variations and equivalent structures and functions.

[0092] Each example can provide an image processing method capable of reducing the processing load in image estimation processing using a machine learning model.

Claims

1. An image processing method, comprising: determining a specific area in the first image; Acquire at least one second image corresponding to some of the plurality of partial images of the first image; generating a third image corresponding to the second image by inputting the second image into a machine learning model; as well as generating a fourth image based on the first image and the third image, Characterized in that the second image includes at least a portion of the specific area.

2. The image processing method according to claim 1, characterized in that: The specific area is determined based on at least one of information about a distance corresponding to the first image, semantic area division information, cropping information, image circle information, and optical performance.

3. The image processing method according to claim 1, characterized in that: The specific area is determined based on at least one of content of the first image, variable image processing settings, and variable camera settings.

4. The image processing method according to claim 2, characterized in that: The information about the distance includes a focus map.

5. The image processing method according to claim 1, characterized in that: The plurality of partial images are acquired by equally dividing the first image.

6. The image processing method according to claim 1, characterized in that: The fourth image is an image based only on an area included in the specific area in the third image.

7. The image processing method according to claim 1, characterized in that: The third image or the fourth image is generated by processing in a manner such that the processing intensity of the machine learning model on the area not included in the specific area is lower than the processing intensity of the machine learning model on the area included in the specific area.

8. The image processing method according to claim 1, characterized in that: The second image includes a first area and a second area, the first area includes the specific area, and the second area does not include the specific area, setting a processing intensity of the machine learning model for an area in the third image corresponding to the first area to a first intensity, setting a processing intensity of an area in the third image corresponding to the second area to a second intensity that is less than the first intensity, and The fourth image is generated by using the area set to the first intensity and the area set to the second intensity in the third image.

9. The image processing method according to claim 1, characterized in that: The fourth image is an image obtained by applying at least one of sharpening, resolution enhancement, and defocus blur conversion to the first image.

10. An image processing method, comprising: determining a specific area in the first image; Acquire a second image corresponding to the specific area; generating a third image corresponding to the second image by inputting the second image into a machine learning model; generating a fourth image based on the first image and the third image, It is characterized in that the specific area is determined based on at least one of information about a distance corresponding to the first image, semantic area division information, cropping information, image circle information and optical performance.

11. An image processing device, comprising: a determining unit configured to determine a specific area in the first image; an acquisition unit configured to acquire at least one second image corresponding to some of the plurality of partial images of the first image; a first generating unit configured to generate a third image corresponding to the second image by inputting the second image into a machine learning model; as well as a second generating unit configured to generate a fourth image based on the first image and the third image, Characterized in that the second image includes at least a portion of the specific area.

12. An image processing device, comprising: a determining unit configured to determine a specific area in the first image; an acquisition unit configured to acquire a second image corresponding to the specific area; a first generating unit configured to generate a third image corresponding to the second image by inputting the second image into a machine learning model; a second generating unit configured to generate a fourth image based on the first image and the third image, It is characterized in that the specific area is determined based on at least one of information about a distance corresponding to the first image, semantic area division information, cropping information, image circle information and optical performance.

13. A camera device, comprising: The image processing device according to claim 11 or 12; as well as Image sensor. 14 . A storage medium storing a computer program for causing a computer to execute the image processing method according to claim 1 .

Citation Information

Patent Citations

  • Image processing method, image processing apparatus, image capturing apparatus, program, and storage medium

    JP2019212139A