Imaging device, control method for imaging device, program, and recording medium

The imaging device enhances subject detection accuracy by adjusting image frequency components and using different aperture settings to minimize the influence of complex background textures, thereby overcoming existing challenges in subject detection.

JP7695102B2Active Publication Date: 2025-06-18CANON KK
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2021077028
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-01-07
Filing Date
2021-04-30
Publication Date
2025-06-18
Estimated Expiration
2041-04-30

AI Technical Summary

Technical Problem

Existing subject detection techniques in imaging devices face challenges in achieving high accuracy when the background has a complex texture pattern or when the subject and background have similar texture patterns.

Method used

The proposed imaging device includes acquisition and detection means, along with control means that adjust the frequency components of images to reduce the influence of background textures, enabling more accurate subject detection. This is achieved by acquiring images with different aperture settings and performing filtering processes to blur the background.

Benefits of technology

The solution effectively reduces the impact of background textures on subject detection, leading to higher accuracy and improved performance in complex scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007695102000002
    Figure 0007695102000002
  • Figure 0007695102000003
    Figure 0007695102000003
  • Figure 0007695102000004
    Figure 0007695102000004
Patent Text Reader

Abstract

To provide an image processing apparatus that can reduce an influence of a subject other than a detection object and a texture of a background on subject detection to perform the subject detection with higher accuracy.SOLUTION: An imaging apparatus acquires an input image with a lens unit 101 and an image pick-up device 141 (S200), and performs subject detection (S201). The imaging apparatus calculates a degree of reliability of the subject detection and compares the degree of reliability with a threshold (S202). When the degree of reliability of the subject detection is less than the threshold, the imaging apparatus executes defocus calculation processing (S203) and background area determination processing (S204). The imaging apparatus executes low-pass filtering processing on the determined background area (S205) to reduce a high frequency component of the background area before performing the subject detection again (S206).SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to Imaging device a subject detection technique in

Background Art

[0002] When detecting a subject in a captured image by an imaging device, a process of extracting an image area (subject area) of the subject is performed. When extracting the subject area, if a subject other than the detection target or the background has a texture pattern similar to the detected subject or a complex texture pattern, there is a possibility that the subject cannot be detected. In Patent Document 1, a subject detection technique suitable for the amount of blur and sharpness of the main subject is disclosed. By selecting parameters of the detection means according to the amount of blur and sharpness of the subject and performing subject detection, it is possible to accurately detect even when the subject is not in perfect focus, for example. Further, in Patent Document 2, a technique is disclosed in which a distance map is used to exclude an area where there is a low possibility that the subject to be detected exists and perform detection. This can reduce the possibility that another subject or background exists in the area to be detected.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the technology of Patent Document 1, even if the subject to be detected is in focus, if the background image has a complex pattern, it may not be possible to obtain the desired detection accuracy in subject detection. Further, in the technology of Patent Document 2, artificial edges are generated by cutting out the subject area, and the accuracy may decrease in subject detection using a convolutional neural network. An object of the present invention is to reduce the influence of a subject or background texture other than the detection target on subject detection and enable higher-accuracy subject detection Imaging device and to provide the same.

Means for Solving the Problems

[0005] An imaging device according to an embodiment of the present invention includes acquisition means for acquiring an image captured by imaging means, detection means for detecting a subject from the acquired image, and determination of a subject detection result by the detection means perform filtering processing by image processing means control means for performing control to adjust frequency components of the whole or a partial region of the image, and the detection means detects a subject with respect to the image in which the frequency components are adjusted, The acquisition means acquires a first image used for image processing and a second image used for recording. When the first image is acquired by the acquisition means, the control means sets the aperture value related to the imaging means to a first value. When the second image is acquired by the acquisition means, the control means sets the aperture value related to the imaging means to a second value. which is characterized in that.

Effects of the Invention

[0006] According to the present invention, it is possible to reduce the influence of a subject or background texture other than the detection target on subject detection and enable higher-accuracy subject detection Imaging device and to provide the same.

Brief Description of the Drawings

[0007]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Modes for Carrying Out the Invention

[0008] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In each embodiment, an example of an imaging device to which an image processing apparatus according to the present invention is applied is shown. [First Embodiment] With reference to FIG. 1, the configuration of the imaging device in the present embodiment will be described. FIG. 1 is a block diagram showing a configuration example of the imaging device 100. The imaging device 100 is a digital still camera, a video camera, or the like that can capture a subject and record moving image or still image data on various media such as a tape, a solid-state memory, an optical disk, or a magnetic disk. The present invention is applicable to various electronic devices having imaging means.

[0009] Each unit in the imaging device 100 is connected via a bus 160. Each unit is controlled by a CPU (Central Processing Unit) 151 that constitutes a control unit. The CPU 151 performs the following processes and controls by executing a program.

[0010] The lens unit 101 includes optical members such as a fixed lens and a movable lens that constitute an imaging optical system. FIG. 1 shows a configuration including a fixed first-group lens 102, a zoom lens (variable magnification lens) 111, a diaphragm 103, a fixed third-group lens 121, and a focus lens (focus adjustment lens) 131.

[0011] In accordance with the instructions from the CPU 151, the aperture control unit 105 controls the adjustment of the aperture diameter of the aperture 103 during shooting by driving the aperture 103 via the aperture motor (AM) 104. The zoom control unit 113 changes the focal length of the imaging optical system by driving the zoom lens 111 via the zoom motor (ZM) 112.

[0012] The focus control unit 133 determines the driving amount of the focus motor (FM) 132 based on the amount of defocus (defocus amount) on the optical axis during the focus adjustment of the lens unit 101. Based on the determined driving amount, the focus control unit 133 drives the focus lens 131 via the focus motor (FM) 132 to control the focus adjustment state. The movement control of the focus lens 131 is performed by the focus control unit 133 and the focus motor (FM) 132 to realize autofocus (AF) control. Although the focus lens 131 is simply shown as a single lens in FIG. 1, it is usually composed of a plurality of lenses.

[0013] The light from the subject forms an image on the imaging element 141 via the lens unit 101. The imaging element 141 performs photoelectric conversion on the subject image (optical image) formed by the imaging optical system and outputs an electrical signal. The imaging element 141 has a configuration in which photoelectric conversion units with a predetermined number of pixels are arranged in the horizontal and vertical directions, and performs photoelectric conversion in the light receiving unit and outputs an electrical signal corresponding to the optical image to the imaging signal processing unit 142. The imaging element 141 is controlled by the imaging control unit 143.

[0014] The imaging signal processing unit 142 processes the signal acquired by the imaging element 141 into an image signal and acquires the image data on the imaging surface. The image data output from the imaging signal processing unit 142 is sent to the imaging control unit 143 and temporarily stored in the RAM (Random Access Memory) 154.

[0015] The image compression / decompression unit 153 reads out the image data stored in the RAM 154, compresses it, and then performs a process of recording it on the image recording medium 157. In parallel with this process, the image data stored in the RAM 154 is sent to the image processing unit 152.

[0016] The image processing unit 152 performs predetermined image processing, such as reduction processing or enlargement processing to an optimal size for the image data, calculation processing of the similarity between image data, gamma correction and white balance processing based on the subject area, etc. The image data processed to the optimal size is sent to the monitor display 150 as appropriate, and an image is displayed, and preview image display or through image display is performed. Also, the object detection result obtained by the object detection unit 162 can be superimposed on the image data and displayed. The object detection unit 162 performs a process of determining an area where a predetermined object exists in the captured image using the image signal.

[0017] By using the RAM 154 as a ring buffer, it is possible to buffer data of a plurality of images captured within a predetermined period and various detection data. The various detection data includes the detection result of the object detection unit 162 corresponding to each image data, data of the positional and postural change of the imaging device 100, etc.

[0018] The positional and postural change acquisition unit 161 includes a positional and postural sensor such as a gyro sensor, an acceleration sensor, an electronic compass, etc., and measures the positional and postural change of the imaging device 100 with respect to the shooting scene. The acquired data of the positional and postural change is stored in the RAM 154.

[0019] The operation switch unit 156 is an input interface unit including a touch panel, operation buttons, etc. The user can issue various operation instructions by selecting and operating various function icons displayed on the monitor display 150. The CPU 151 controls the imaging operation based on the operation instruction signal input from the operation switch unit 156 or the magnitude of the pixel signal of the image data temporarily stored in the RAM 154. For example, the CPU 151 determines the accumulation time of the imaging element 141 and the gain setting value when outputting from the imaging element 141 to the imaging signal processing unit 142. The imaging control unit 143 receives instructions for the accumulation time and the gain setting value from the CPU 151 and controls the imaging element 141.

[0020] The CPU 151 sends commands to the focus control unit 133 to perform AF control on a specific subject area, and also sends commands to the aperture control unit 105 to perform exposure control using the luminance value of the specific subject area.

[0021] The monitor display 150 has a display device and performs image display, rectangular display of object detection results, etc. The power management unit 158 manages the battery 159 and supplies stable power to the entire imaging device 100.

[0022] The flash memory 155 stores control programs necessary for the operation of the imaging device 100, parameters used for the operation of each part, etc. When the imaging device 100 is started up by shifting from the power-off state to the power-on state by the user's operation, the control programs and parameters stored in the flash memory 155 are read into a part of the RAM 154. The CPU 151 controls the operation of the imaging device 100 according to the control programs and constants loaded into the RAM 154.

[0023] The defocus calculation unit 163 calculates the defocus amount for an arbitrary subject in the captured image. Since the method for calculating the defocus amount is well-known, its description is omitted. The generated defocus information is stored in the RAM 154 and referred to by the image processing unit 152. In this embodiment, an example of obtaining the distribution information of the defocus amount in the captured image is shown, but there are other methods. For example, there is a method of generating a plurality of viewpoint images (disparity images) by splitting the light from the subject by the pupil, calculating the disparity amount, and obtaining the depth distribution information of the subject. The pupil-splitting type imaging device includes a plurality of microlenses and a plurality of photoelectric conversion units corresponding to each microlens, and can output signals of different viewpoint images from each photoelectric conversion unit. The depth distribution information of the subject includes data representing the distance from the imaging unit to the subject (subject distance) as an absolute value in distance values, and data indicating the relative distance relationship (depth of the image) in the image data (such as distribution data of the disparity amount). The depth direction corresponds to the depth direction with respect to the imaging unit. The plurality of viewpoint image data can also be obtained by a multi-eye camera having a plurality of imaging means.

[0024] In this embodiment, an example of using a convolutional neural network as the subject detection means based on machine learning is shown. In this specification, the convolutional neural network is abbreviated as CNN. The CNN is configured by stacking convolutional layers and pooling layers. The subject detection means outputs data of a rectangular region on the image and data of the reliability of the detection result. As an example, the reliability is output as an integer value from 0 to 255, and it is assumed that the likelihood of false detection is higher as the value of the reliability is smaller. The CPU 151 realizes the following processing using the data and program of the learned model for subject detection.

[0025] In the subject detection process using a CNN, convolution operations using filters obtained in advance by machine learning are performed multiple times. Since the convolution operation, that is, the sum-of-products operation is performed between a target pixel and the pixel values in its surrounding area, due to its characteristics, the operation result corresponding to the area of the subject to be detected is also affected by the pixel pattern of the surrounding background area. The range of the background area that affects the detection of the subject area depends on the size of the filter and the number of layers of the network.

[0026] In order to reduce the influence of the pixel pattern of the background area on subject detection, the CPU 151 executes a process of reducing the high-frequency components of the background area. This process can be realized as follows. (1) Based on defocus information, depth information, distance information, etc., perform predetermined image processing on the pixels in the area determined to be a distant area (background area) from the imaging device 100 to blur the image. The predetermined processing includes low-pass filtering processing and band-pass filtering processing performed by the image processing unit 152. (2) When the desired subject is in focus to a certain extent, the aperture control unit 105 performs control to drive the aperture 103 of the lens unit 101 in the direction of increasing the aperture diameter to increase the amount of blur in the background area. (3) The focus control unit 133 drives the focus lens 131 and changes the focus position in a predetermined direction to increase the amount of defocus in the background area.

[0027] Alternatively, by combining multiple processes, it is possible to perform a process of reducing the high-frequency components of the background area. Note that the processes of reducing the high-frequency components of the background area shown in (1) to (3) are examples of processes for adjusting the frequency components of the entire image or a partial area. The CPU 151 determines the method of the frequency component adjustment process by judging the detection result of the subject, and performs control to improve the accuracy of subject detection.

[0028] Referring to FIG. 2, the processing flow in this embodiment will be described. FIG. 2 is a flowchart showing an example of processing, and the following processing is realized by the CPU 151 executing a program to control each part of FIG. 2.

[0029] In S200, the imaging control unit 143 processes the signal acquired by the imaging element 141 and supplies input image data to each part. In the next S201, subject detection is performed on the input image. A subject detection process using a CNN is executed, and a process of outputting a rectangular region on the captured image and the reliability of the detection result is performed.

[0030] In S202, the CPU 151 determines whether a detection result with high reliability has been obtained. If it is determined that some subject has been detected and the reliability is equal to or higher than a predetermined threshold value, the process ends. In this case, after performing arbitrary processing such as AF control and frame display on the detected subject, the processing for the image of the subject ends. On the other hand, if it is determined that a detection result with high reliability has not been obtained, the process proceeds to the processing of S203.

[0031] In S203, the defocus calculation unit 163 calculates the defocus amount of each image region and outputs defocus information. In the next S204, the CPU 151 performs a background region determination process using the defocus information calculated in S203. Here, it is assumed that a region where the defocus amount is equal to or higher than a threshold value is regarded as the background region.

[0032] In S205, the CPU 151 and the image processing unit 152 perform low-pass filtering processing only on the region regarded as the background region in S204. In S206, the CPU 151 executes the subject detection process again. That is, the subject detection process is performed on the image after the low-pass filtering process. Thereafter, after performing arbitrary processing such as AF control and frame display according to the result of the subject detection, the processing for the image of the detected subject ends.

[0033] [Second Embodiment] Next, a second embodiment of the present invention will be described. In this embodiment, an example is shown in which the blurring method of an image is changed by adjusting the aperture 103 to improve the detection performance. FIG. 3 is a flowchart showing an example of processing during standby, which is a stage where the photographer is adjusting the composition. Regarding matters similar to those in the first embodiment, the same reference numerals and symbols as those already used are reused, and detailed descriptions thereof are omitted, and mainly the differences will be described. Such an omission method of description is the same in the embodiments described later.

[0034] Referring to FIG. 3, the sequence during standby will be described. At S300, the imaging control unit 143 acquires the nth input image data and supplies it to each unit. n is a variable of a natural number, and its initial value is set to 1. At S301, subject detection based on CNN is performed on the nth input image. At S302, the CPU 151 determines whether a highly reliable detection result has been obtained. If the reliability is equal to or higher than the threshold value, after performing processing such as AF control and frame display on the detected subject, the processing for the subject image is terminated and the process proceeds to S310. If the reliability is less than the threshold value at S302, the process proceeds to S303.

[0035] At S303, a defocus calculation process is performed, and the defocus calculation unit 163 supplies defocus information to each unit. At S304, the CPU 151 calculates a difference value between the maximum value and the minimum value of the defocus amount of the entire image. This difference value is used to evaluate whether there is a distance difference in the depth direction in the shooting scene. If the calculated difference value is less than the threshold value, the process proceeds to S310. If the difference value is equal to or greater than the threshold value, the process proceeds to S305. In any case, the processing for the image is terminated.

[0036] At S305, the CPU 151 performs a process of setting the aperture value to a relatively small value. For example, a process of setting an aperture value one step smaller than the current aperture value is performed. Alternatively, the minimum aperture value that can be set in the imaging device 100 may be set.

[0037] In S306, the imaging control unit 143 acquires the next frame, that is, the (n + 1)-th input image. In the next S307, subject detection based on CNN is performed on the (n + 1)-th input image. In S308, the CPU 151 determines whether a detection result with high reliability has been obtained. If the reliability is equal to or higher than the threshold value, after performing processing such as AF control and frame display on the detected subject and then ending the processing on the subject image, the process proceeds to S309. If the reliability is less than the threshold value, the process proceeds to the processing of S310.

[0038] In S309 and S310, the CPU 151 executes flag setting processing. The value of this flag is true when it indicates that the input image in which the background area is blurred for the shooting scene is advantageous for subject detection, and false when it indicates that the input image is not advantageous. In S309, the flag is set to be valid, that is, set to true. In S310, the flag is set to be invalid, that is, set to false. After S309 and S310, a series of processing is completed.

[0039] With reference to FIG. 4, the processing during continuous still image shooting will be described. FIG. 4 is a flowchart for explaining an example of a sequence during continuous still image shooting. Here, during continuous still image shooting, it is assumed that the first frame for acquiring the image (still image) data actually output to the recording medium and the second frame for acquiring the evaluation image data used for image processing inside the imaging device are alternately repeated. Also, subject detection, AF control and frame display, etc. for the detection result are performed only on the evaluation image.

[0040] In S400, the CPU 151 determines whether the target frame is the second frame and whether the evaluation image data has been acquired. If it is determined that the evaluation image data has been acquired, the process proceeds to the processing of S401. If it is determined that the recording still image data has been acquired in the first frame, the process proceeds to the processing of S405.

[0041] In S401, the CPU 151 executes the determination process of the flags set in S309 and S310 (FIG. 3). If it is determined that the flag is valid (true value), the process proceeds to S402. If it is determined that the flag is invalid (false value), the process proceeds to S405.

[0042] In S402, the CPU 151 sets a small aperture value in the same manner as S305 in FIG. 3. After that, after the input image data is acquired in S403, subject detection based on CNN is performed in S404. If some subject is detected and it is determined that the reliability is equal to or higher than the threshold value, after performing processes such as AF control and frame display for the subject, the process for the input image is terminated. In S405, the CPU 151 sets the aperture value as designated by the photographer. After the input image data is acquired in the next S406, the process for the input image is terminated.

[0043] In the present embodiment, in S304 of FIG. 3, it is determined whether there is a distance difference in the depth direction from the defocus information, and the aperture value is adjusted based on the determination result. Also, when it is determined in S401 of FIG. 4 that the flag is valid, for example, different aperture values are set for the evaluation image and the recording still image, and the input image is acquired. At this time, the CPU 151 determines the time required for opening and closing the aperture and sets the upper limit value of the continuous shooting speed. By changing the way the image is blurred by adjusting the aperture 103, it is possible to improve the detection performance.

[0044] [Third Embodiment] Referring to FIGS. 3 and 4, a third embodiment of the present invention will be described. In the present embodiment, a configuration is shown in which it is controlled whether to decrease the aperture value by using information on the degree of variation in the subject detection results for each frame. For example, assume a case where subject detection is performed for two consecutive frames. If a subject detection result having a reliability equal to or higher than a predetermined threshold value is obtained for any of the frames, the CPU 151 performs the same control as in the above embodiment.

[0045] The CPU 151 calculates the difference value of the reliability of the subject detection results for two consecutive frames and compares it with a predetermined difference threshold. If the difference value of the reliability is equal to or greater than the difference threshold, the CPU 151 determines that the detection results are fluctuating. At this time, for the frame in which the subject is not detected, the CPU 151 calculates the difference value with the reliability value set to zero.

[0046] When it is determined that the subject detection results are fluctuating, the CPU 151 determines that it is possible to reduce the influence of the background pattern in the shooting scene on the performance of the subject detection means by reducing the aperture value and blurring the background image. In this case, in S309 of FIG. 3, the CPU 151 effectively sets the flag and then executes the continuous shooting sequence of S400 to S406 in FIG. 4.

[0047] Next, an application example to an imaging device that can control focus and defocus for each region, such as a light field camera, will be described. The light field camera can focus on a desired region or position by splitting incident light with a microlens array arranged on the imaging element and acquiring light intensity information and incident direction information.

[0048] As an example, assume a shooting scene in which there is a certain depth difference within the subject region to be detected. In such a shooting scene, the subject may not be detected because the entire subject region is not in focus. Alternatively, the focus may be off from the main subject, and the main subject may not be detected because the image of the entire main subject is weakly blurred. In such a case, the CPU 151 determines that the subject is more likely to be detected by performing focus (focus adjustment) control only on the region where the defocus amount is within a predetermined range. The reason for limiting the region to the region where the defocus amount is within the predetermined range is that if the focus control is performed up to the background region when the background has a complex pattern, the subject may not be detected as a result. The CPU 151 regards the region where the defocus amount is equal to or greater than the threshold as the background region, and does not perform any processing on the region, or performs control to make the defocus amount even larger.

[0049] The processing of this embodiment is performed without distinguishing between the evaluation image and the recording still image, or is performed on the evaluation image as in FIG. 4 and not on the recording still image.

[0050] Next, an example of an imaging device having a shooting mode for automatically determining the aperture value will be described. In this shooting mode, the CPU 151 determines an aperture value (temporary value) by a known method. At this time, as described in S300 to S310 in FIG. 3, the CPU 151 determines whether the subject detection performance becomes higher when the value is made smaller than the aperture value (temporary value). When the CPU 151 determines that the subject detection performance is higher than the current level, it further executes a process of reducing the aperture value, and accordingly adjusts the set values such as the shutter speed. Also, the CPU 151 determines the distance difference in the depth direction in the captured image in the same manner as S304 in FIG. 3. When there is a distance difference equal to or greater than the threshold value (or when the difference value between the maximum value and the minimum value of the defocus amount of the entire image is equal to or greater than the difference threshold value), the CPU 151 repeatedly performs a process of further reducing the aperture value (temporary value) to determine the final aperture value.

[0051] According to the above embodiment, in a scene including a complex background pattern, it is possible to provide an imaging device that reduces the influence of the complex background pattern on subject detection and enables more accurate subject detection. Note that the subject detection process based on machine learning shown in the above embodiment is an example. Various subject detection processes capable of calculating the reliability of subject detection (such as the reliability of the correlation operation in phase difference detection) from not only the learning model for subject detection but also the defocus amount, the amount of image displacement of a plurality of viewpoint images, etc. can be adopted.

[0052] [Fourth Embodiment] Referring to FIGS. 5 to 8, a fourth embodiment of the present invention will be described. In this embodiment, a defocus map (or distance map) in which a defocus value (or distance value) is stored for each pixel is used to limit the region where the subject exists. By significantly reducing false detection due to texture patterns similar to the subject or complex texture patterns and also using the contour shape of the subject, the detection accuracy can be improved. At this time, after cutting out the region including the detection target, a process of smoothing the vicinity of the region boundary or a process of gradually attenuating the pixel value in the region other than the detection target is performed. Thereby, the generation of artificial edges that leads to a decrease in the accuracy of the CNN can be suppressed.

[0053] FIG. 5 is a block diagram showing a part of the configuration of the image processing unit 152. The image processing unit 152 includes a region division unit 500 and a pixel value adjustment unit 501. FIG. 6 is a flowchart showing an example of the process. The following processes are realized by the CPU 151 executing a program to control each part of FIG. 5. The processes of S200 and S203 have been described with reference to FIG. 2, and the process of S301 has been described with reference to FIG. 3. After S203, the process proceeds to S600.

[0054] In S600, the region division unit 500 divides the region of the input image using the defocus map. Here, it is assumed that the division is performed based on the distribution histogram of the defocus values, but existing clustering and region division methods such as the k-means method and superpixel may also be used.

[0055] In S601, the pixel value adjustment unit 501 performs edge suppression processing. Based on the information of the divided region in S600, the pixel value adjustment unit 501 performs low-pass filtering processing or multiplication of weights on the input image to suppress the generation of edges.

[0056] Details of the processes performed by the region division unit 500 and the pixel value adjustment unit 501 will be described using the conceptual diagram of FIG. 7 and the flowchart of FIG. 8. Hereinafter, in S601 of FIG. 6, the case where low-pass filtering processing is performed will be described. FIG. 7 shows a captured image 700 and a defocus map 701 in which defocus values are stored for each pixel. An example in which a subject 710 is captured in the captured image 700 is shown. Note that the number of pixels of the captured image 700 and the number of pixels of the defocus map 701 do not have to be the same. Hereinafter, for convenience of explanation, it is assumed that the defocus map 701 has been enlarged (or reduced) by an appropriate interpolation method and has the same number of pixels as the captured image 700. Further, instead of the defocus map 701, a distance map in which distance values from the imaging device to the subject are stored for each pixel may be used. In the defocus map 701 of FIG. 7, a region where the pixel value is 0 represents a focused region, and the larger the pixel value, the more distant the region is from the in-focus position.

[0057] In S600 of FIG. 6, the region division unit 500 divides the defocus map 701 of FIG. 7 into a region 711 and a region 712. Here, it is assumed that the focused region 711 is a region to be detected (hereinafter referred to as a detection region). The image 702 in FIG. 7 is an image obtained by cutting out only the region corresponding to the detection region 711 in the captured image 700, but edges that do not exist in the original captured image 700 are generated near the subject 710. Therefore, in the present embodiment, as shown in the image 703, a process of cutting out a region 713 that includes the detection region 711 from the captured image 700 is executed. Then, by performing low-pass filtering processing on the vicinity of the boundary of the region 713, the generation of edges can be suppressed. Hereinafter, the flow of the process will be described with reference to FIG. 8.

[0058] In S800, the region division unit 500 determines a detection region. Although FIG. 7 shows an example in which there is one detection region 711, there may be a plurality of detection regions 711. In that case, the process of FIG. 8 is executed for each detection region. For example, it is possible to perform clustering of defocus values and sequentially set clusters corresponding to the respective defocus values as detection regions.

[0059] In S801, in the image 700, a process of cutting out the region 713 is performed so as to include the detection region 711. The image obtained as a result of the process is defined as the image 703 (FIG. 7). Hereinafter, the region obtained by removing the detection region 711 from the cut-out region 713 is referred to as the margin region. The size of the margin region is determined by the receptive field of the CNN and the presence or absence of occlusion with respect to another subject. The receptive field of the CNN represents the range in which the detector convolves pixel values. It is desirable that the range for convolving the information of the subject 710 does not include the edges generated by the cutting out. Therefore, there is a method of setting the width of the margin region in proportion to the size of the receptive field. Further, when another subject overlaps in front of the detection region 711, if the margin region is made small, there is a possibility that the subject is missing and the accuracy is reduced. Therefore, when occlusion occurs in the detection region 711, the setting is made such that the width of the margin region is made larger.

[0060] In S802, the pixel value adjustment unit 501 applies a low-pass filter to the target pixel within the margin area. Thereby, while maintaining the pixel values in the detection area 711, the margin area can be blurred, and the generation of artificial edges can be suppressed. Note that the number of taps of the low-pass filter may be changed according to the distance on the image from the boundary of the detection area 711. Also, a process of switching whether to apply the low-pass filter may be performed according to whether the distance on the image from the boundary of the detection area 711 exceeds a predetermined threshold value. The boundary of the detection area 711 can be calculated by extracting the edges after region division. The reason for changing the number of taps and the presence or absence of the filtering process based on the distance from the boundary of the detection area 711 is that, near the boundary of the detection area 711, due to the defocus value error, the pixels constituting the subject are likely to be misclassified into the margin area. Therefore, it is desirable to reduce the number of taps of the filter or not perform the filtering process near the boundary of the detection area 711. As described above, by changing the number of taps and the presence or absence of the filtering process based on the distance from the boundary of the detection area 711, it becomes possible to smooth the pixel values in other areas while maintaining the pixel values of the subject to be detected.

[0061] Furthermore, the number of taps of the low-pass filter may be determined based on the difference between the average value of the defocus values in the detection area 711 and the average value of the defocus values around the target pixel. For the same reason as described above, there is a possibility that the pixels constituting the subject are misclassified into the margin area due to the defocus value error. However, near the boundary of the cut-out area 713, in order to suppress the generation of edges, it is assumed that a low-pass filter having a number of taps equal to or more than a certain value is applied.

[0062] In S803, a determination process is executed to determine whether the pixel value adjustment unit 501 has processed all the pixel values within the margin area. When it is determined that the processing by the pixel value adjustment unit 501 has been completed, a series of processes is terminated. When it is determined that the processing by the pixel value adjustment unit 501 has not been completed, the process proceeds to S804. After the process of updating the pixel of interest (the process of changing the position of the pixel of interest) is performed in S804, the process returns to S802 and continues.

[0063] As described above, an example of applying the low-pass filter in S601 of FIG. 6 has been explained. Without being limited to this example, a process of multiplying a weight based on the distance from the boundary of the detection region 711 by the pixel value and gradually decreasing the pixel value may be performed. In this case, for example, the region 712 excluding the detection region 711 from the image of the subject 710 is set as the margin region. A process of multiplying all the pixels in the region corresponding to the margin region (region 712) in the image 700 by a weight (denoted as w) defined by the following formula is executed.

Equation

[0064] The above method is an example, and any other method may be used as long as it can suppress the generation of artificial edges due to cutting out the region while maintaining the pixel values of the region to be detected.

[0065] According to the present embodiment, when limiting the region where the subject exists, the influence of the edges generated by cutting out can be suppressed, so that the influence of the texture patterns of the subjects and backgrounds other than the detection target can be reduced, and the detection performance can be improved.

[0066] [Fifth Embodiment] Referring to FIG. 9, a fifth embodiment of the present invention will be described. In the fourth embodiment, the operation during inference was described. In this embodiment, a case will be described in which the same processing as in the fourth embodiment is performed on the image during machine learning. Since the characteristics of the image during learning and inference can be made to match, the accuracy of the detector can be further improved. Also, since the contour information of the subject is learned at the same time, it is possible to distinguish the recognition target depicted in a painting or a photograph from an actual recognition target.

[0067] Specifically, by executing the processing shown in FIGS. 7 and 8, a process of acquiring a learning image is performed. Even when the subject is cut out using the defocus map during inference without using the method of this embodiment, the characteristics of the learning image can be made uniform, but if occlusion occurs in the subject, the accuracy is likely to decrease. The reason for this will be described with reference to FIG. 9.

[0068] FIG. 9 is a schematic diagram showing a photographed image 900, a defocus map 901, a cut-out image 902, and an image 903 for machine learning. Images of subjects 910 and 911 shown in the photographed image 900 are shown. Since the subject 911 is on the front side (camera side), there is a situation where occlusion occurs in a part of the subject 910. In the defocus map 901 in which the defocus value is stored for each pixel, regions 912 and 913 correspond to the images of the subjects 910 and 911, respectively. Regions 912 and 913 represent the result of dividing the defocus map 901 into regions.

[0069] The image 902 is the result of cropping the region 912, and there are deficiencies in part of the outline of the subject. When simply cropping the part corresponding to the subject 910 as in the image 902, machine learning will be performed by mixing an image with deficiencies in the outline of the subject to be detected and an image without deficiencies. Therefore, it may lead to a decrease in detection accuracy. On the other hand, in the present embodiment, within the region 914 including the image of the subject 910 as in the image 903, a process of applying a low-pass filter to the regions other than the subject is executed. Therefore, no deficiencies occur in the outline of the subject, and an image can be obtained in a state close to a normal image to perform machine learning. According to the present embodiment, by matching the characteristics of the images during learning and inference, the detection performance can be further improved compared to the fourth embodiment.

[0070] [Sixth Embodiment] Next, the sixth embodiment of the present invention will be described. In the present embodiment, a case where the processes of the first to third embodiments and the process of the fourth embodiment are simultaneously performed will be described. The first threshold for the area of the background region in the image is denoted as Th1, the second threshold for the total area of the regions other than the background is denoted as Th2, and the third threshold for the number of divided regions other than the background is denoted as Th3.

[0071] The process in the present embodiment will be described with reference to FIGS. 2 and 6. When the area of the background region in the image is greater than or equal to the threshold Th1, the filtering process of S205 in FIG. 2 is executed, and when it is less than the threshold Th1, the process of S205 is skipped. Also, when the total area of the regions other than the background is greater than or equal to the threshold Th2, or when the number of divided regions other than the background is greater than or equal to the threshold Th3, the process of S601 in FIG. 6 is executed. When the total area of the regions other than the background is less than the threshold Th2 and the number of divided regions other than the background is less than the threshold Th3, the process of S601 is skipped. In this way, it is possible to achieve both high-speed processing and high-performance detection.

[0072] [Other Embodiments] The present invention can also be realized by supplying a program that implements one or more functions of the above-described embodiments to a system or apparatus via a network or a storage medium, and causing one or more processors in a computer of the system or apparatus to read and execute the program. It can also be realized by a circuit (for example, an ASIC) that implements one or more functions.

Explanation of Signs

[0073] 100 Imaging device 105 Diaphragm control unit 133 Focus control unit 141 Image sensor 142 Image signal processing unit 143 Imaging control unit 151 CPU 152 Image processing unit 162 Object detection unit 163 Defocus calculation unit

Claims

1. An acquisition means for acquiring an image captured by an imaging means; A detection means for detecting a subject from the acquired image; A control means for determining a detection result of a subject by the detection means and performing control to perform filtering processing by an image processing means to adjust frequency components of the whole or a partial region of the image, and comprising: The detection means performs detection of a subject on the image in which the frequency components are adjusted; The acquisition means acquires a first image used for image processing and a second image used for recording; When the first image is acquired by the acquisition means, the control means sets an aperture value related to the imaging means to a first value, and when the second image is acquired by the acquisition means, the control means sets an aperture value related to the imaging means to a second value. An imaging device characterized by the above.

2. The control means determines a method for adjusting the frequency components and performs control to enhance the accuracy of detection of a subject by the detection means. The imaging device according to claim 1, characterized by the above.

3. The control means determines a background region for the subject and performs control to reduce high-frequency components of the background region. The imaging device according to claim 1 or claim 2, characterized by the above.

4. The control means performs control to adjust the frequency components by adjusting a value of an aperture included in the imaging means. The imaging device according to any one of claims 1 to 3, characterized by the above.

5. The control means performs control to adjust the frequency components by controlling a focus adjustment lens. The imaging device according to any one of claims 1 to 3, characterized by the above.

6. When the reliability of the detection result of the subject is less than the threshold value, the control means changes the aperture value related to the imaging means, and the detection means detects the subject from the acquired image. The imaging device according to any one of claims 1 to 5, characterized in that.

7. When the distance difference in the depth direction in the captured image is equal to or greater than the threshold value, the control means changes the aperture value related to the imaging means, and the detection means detects the subject from the acquired image. The imaging device according to any one of claims 1 to 5, characterized in that.

8. When the aperture value related to the imaging means is set to the first value, the detection means detects the subject from the first image acquired by the acquisition means. The imaging device according to any one of claims 1 to 7, characterized in that.

9. The control means calculates a difference value of the reliability of the detection results of the subject for a plurality of consecutive frames, and when it is determined that the difference value is equal to or greater than the threshold value, performs control to decrease the aperture value related to the imaging means. The imaging device according to any one of claims 1 to 8, characterized in that.

10. It includes a calculation means for calculating information on the defocus amount, depth, or distance regarding the subject in the captured image. The control means performs control to adjust the frequency component using the information calculated by the calculation means. The imaging device according to any one of claims 1 to 9, characterized in that.

11. It includes a calculation means for calculating the defocus amount regarding the subject in the captured image. The control means performs control to focus the imaging optical system on an area where the calculated defocus amount is within a predetermined range. The imaging device according to any one of claims 1 to 9, characterized in that.

12. When a shooting mode for automatically determining the aperture value of the imaging optical system is selected and the distance difference in the depth direction in the captured image is equal to or greater than a threshold value, the control means performs control to reduce the aperture value of the imaging optical system. The imaging device according to any one of claims 1 to 11, characterized in that.

13. The imaging device further includes extraction means for extracting a region of the subject based on the defocus amount, depth, or distance information. The control means performs control to adjust the frequency components for a part of the region including the region of the subject. The imaging device according to claim 10, characterized in that.

14. The detection means is constituted by a convolutional neural network. The imaging device according to any one of claims 1 to 13, characterized in that.

15. The extraction means determines the size of the region including the region of the subject based on the receptive field of the detection means or the presence or absence of occlusion with another subject. The imaging device according to claim 13, characterized in that.

16. The control means performs control to adjust the frequency components based on any one or more of the defocus amount, depth, or distance information. The imaging device according to claim 13 or claim 15, characterized in that.

17. The detection means performs machine learning on the image in which the frequency components are adjusted. The imaging device according to claim 14, characterized in that.

18. The control means changes the adjustment process for the image by comparing any one or more of the area of the background region in the image, the total area of the regions different from the background region, or the number of divided regions other than the background with a threshold value. The imaging device according to any one of claims 13 to 17, characterized in that.

19. The control means performs control to suppress edges on a part of the region including the region of the subject. The imaging device according to any one of claims 13 to 18, characterized in that.

20. A control method executed by an imaging device capable of detecting a subject, comprising: An acquisition step in which an acquisition means acquires an image captured by an imaging means; A step in which a detection means detects a subject from the acquired image; A control step in which the control means determines a detection result of the subject by the detection means and performs filtering processing by an image processing means to adjust frequency components of the whole or a partial region of the image; A step in which the detection means detects a subject from the image whose frequency components are adjusted, and having: In the acquisition step, a first image used for image processing and a second image used for recording are acquired by the acquisition means, In the control step, when the first image is acquired by the acquisition means, the control means sets an aperture value related to the imaging means to a first value, and when the second image is acquired by the acquisition means, the control means sets an aperture value related to the imaging means to a second value. A control method for an imaging device, characterized in that.

21. A program for causing a computer of an imaging device to execute each step according to claim 20.

22. A computer-readable recording medium on which the program according to claim 21 is recorded.

Citation Information

Patent Citations

  • Microprocessor

    JP1988058552A

  • Image processing device and image processing method

    JP2009111921A

  • Image processing apparatus, method and program

    JP2011253354A

  • Image pickup device

    JP2014155042A

  • Image processing apparatus, image processing method, and imaging apparatus

    JP2019186911A