Image processing apparatus and control method thereof

The image processing apparatus addresses autofocus challenges by selecting focus target areas based on human body detection and depth information, maintaining stable focus despite depth changes and obstructions.

JP7844108B2Active Publication Date: 2026-04-13CANON KK
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
CANON KK
Filing Date
2021-06-09
Publication Date
2026-04-13

AI Technical Summary

Technical Problem

Existing autofocus methods struggle when specific body parts are obstructed or when there is a large difference in depth within the focus area, making it difficult to maintain accurate focus, especially on body parts like the torso or arm.

Method used

An image processing apparatus that determines a focus target area by detecting human body regions, acquiring depth information, and selecting suitable sub-regions based on comparison with a reference area using depth and image features to maintain focus.

Benefits of technology

Enables continuous and accurate autofocus even when there is a significant change in depth within the focus area, ensuring stable focus on human body parts and other objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007844108000001
    Figure 0007844108000001
  • Figure 0007844108000002
    Figure 0007844108000002
  • Figure 0007844108000003
    Figure 0007844108000003
Patent Text Reader

Abstract

To enable a selection of a focus target area on which AF can be suitably executed.SOLUTION: An image processing apparatus is operable to determine a focus target area of an image capturing apparatus. The image processing apparatus includes: obtainment means configured to obtain a first area to be a focus target in a first image captured by the image capturing apparatus at a first point in time; detection means configured to detect a second area to be a focus target candidate from a second image captured by the image capturing apparatus at a second point in time succeeding the first point in time; and determination means configured to determine, on the basis of the first area and the second image, a focus target area in the second image from among one or more partial areas in the second area.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technique for selecting an image area.

Background Art

[0002] In shooting with a camera, there is an autofocus (AF) function that automatically focuses. As a method for selecting an area to be focused on (hereinafter referred to as a focusing target area) when shooting, there are a method in which a user manually selects using a touch panel or the like, and a method in which it is automatically selected based on detection results such as face detection and object detection. Regardless of the method by which the focusing target area is selected, the position and shape on the image may change due to the movement of an object within the selected focusing target area or the camera itself. At this time, by tracking or continuously detecting the selected focusing target area, it is possible to continue AF on the area desired by the user.

[0003] Patent Document 1 discloses a method of detecting a pupil area and performing AF using it as a focusing target area. According to this method, since a focusing target area with a constant distance from the camera is used, it is possible to focus accurately.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, the method described in Patent Document 1 requires the detection of specific body parts, such as the pupil. Therefore, it cannot be applied when these specific body parts cannot be observed due to obstruction by other objects. Furthermore, in the method described in Patent Document 1, it is difficult to focus accurately when there is a large difference in depth (distance from the camera) within the area to be focused. Therefore, it is difficult to apply when the area to be focused should be a body part of a certain size, such as the torso or arm.

[0006] This invention has been made in view of these problems, and aims to provide a technology that enables the selection of a focus target area in which autofocus can be suitably performed. [Means for solving the problem]

[0007] To solve the above-mentioned problems, the image processing apparatus according to the present invention has the following configuration. That is, the image processing apparatus that determines the focus target area of ​​the imaging apparatus determines the focus target in the first image captured by the imaging apparatus at a first time point. and point to the head of a person Acquisition means for acquiring a first region, depth information acquisition means for acquiring depth information for each region of the image captured by the imaging device, and a candidate for focus target from a second image captured by the imaging device at a second time point following the first time point. and the area including the torso of the person The system includes detection means for detecting a second region, and determination means for determining a focus target region in the second image from one or more subregions of the second region based on depth information corresponding to the first region and depth information corresponding to the second region. [Effects of the Invention]

[0008] According to the present invention, it is possible to provide a technique that enables the selection of a target area for which autofocus can be suitably performed. [Brief explanation of the drawing]

[0009] [Figure 1] This diagram shows the relative positions of the camera and the subject. [Figure 2]This figure shows an example of the configuration of the AF system in the first embodiment. [Figure 3] This is a flowchart illustrating the processes performed by the AF system in the first embodiment. [Figure 4] This diagram illustrates the comparison between the reference region and each subregion (S108). [Figure 5] This figure shows an example of detecting the head and torso of a human body. [Figure 6] This diagram shows the hardware configuration of the image processing device. [Figure 7] This is a flowchart illustrating the processes performed by the AF system in the second embodiment. [Figure 8] This diagram illustrates a comparison between the reference region and a sub-region of the head. [Figure 9] This diagram illustrates the comparison between the reference area and the torso area. [Modes for carrying out the invention]

[0010] The embodiments will be described in detail below with reference to the attached drawings. Note that the following embodiments do not limit the invention as defined in the claims. While the embodiments describe multiple features, not all of these features are essential to the invention, and the features may be combined in any way. Furthermore, in the attached drawings, identical or similar configurations are given the same reference numerals, and redundant descriptions are omitted.

[0011] (First Embodiment) As a first embodiment of the image processing apparatus according to the present invention, an autofocus (AF) system including a shooting device and a region selection device will be described below as an example. In particular, in the first embodiment, the AF system detects a human body based on an image acquired from the shooting device and extracts a focus target region from the detected human body region.

[0012] <Device configuration> FIG. 1 is a diagram showing the positional relationship between a photographing device (camera) and a subject. FIG. 1(a) exemplarily shows a subject G-1 which is a human body in a standing position, and FIG. 1(b) exemplarily shows a subject G-2 which is a human body in a supine position. The camera is located in the lower left direction of the figure, and a plurality of dotted lines indicating the distance (depth) from the camera are shown.

[0013] As shown in FIG. 1, the distance from the camera to each part of the body changes depending on the posture of the subject. For example, when the subject is in a standing position as shown in the subject G-1, the distance from the camera does not change significantly for any part. On the other hand, when lying substantially parallel to the line of sight of the camera as shown in the subject G-2, the depth changes significantly depending on the part of the body. And when the change width of the depth within the subject is wider than the depth of field of the imaging device, generally it is not possible to focus the entire subject. As a result, there may be cases where suitable AF cannot be continuously executed. Therefore, in the first embodiment, an example of enabling continuous and suitable AF even when the change width of the depth within the subject is wider than the depth of field of the imaging device will be described.

[0014] <Device Configuration> FIG. 2 is a diagram showing an example of the configuration of an AF system in the first embodiment. As shown in FIG. 2, the AF system includes a photographing device 10 and a region selection device 20.

[0015] The photographing device 10 is a camera device that images the scenery of the surrounding environment. The photographing device 10 includes an image acquisition unit 11 and a distance measurement unit 12. Examples of the photographing device 10 include a digital single-lens reflex camera, a smartphone, a wearable camera, a network camera, a web camera, etc. However, it is not limited to these examples, and any device that can image the surrounding scenery may be used.

[0016] The image acquisition unit 11 images the scenery around the imaging device 10 using an imaging element or the like, and outputs it to the region selection device 20. The image acquired by the image acquisition unit 11 may be RAW data before demosaicking processing, or an image in which all pixels have RGB values by demosaicking or the like. Further, it may be an image for live view.

[0017] The distance measurement unit 12 has a distance measurement function for measuring depth information, which is the distance between the imaging device 10 and the subject, and outputs the measured depth information to the region selection device 20. The depth information should be capable of being associated with each pixel or each region of the image acquired by the image acquisition unit 11. Further, the depth information is any information correlated with the length of the spatial distance. For example, it may be the spatial length itself, or the defocus amount based on a phase difference sensor or the like that detects the phase difference of incident light. Further, it may be the amount of change in the contrast of the image when the focal plane of the lens is moved.

[0018] The region selection device 20 (image processing device) detects the region of the human body based on the image and depth information input from the imaging device 10. Then, it selects a focusing target region to be focused on from the detected regions. The region selection device 20 includes a detection unit 21, a partial region extraction unit 22, a reference region acquisition unit 23, a comparison unit 24, and a selection unit 25. In FIG. 2, the region selection device 20 is shown as being separate from the imaging device 10, but it may be configured as an integrated device. Further, when configured as a separate body, it may be connected by a wired or wireless communication function. Further, each functional unit of the region selection device 20 can also be realized by a central processing unit (CPU) executing a software program.

[0019] Figure 6 shows the hardware configuration of the information processing device. The CPU 1001 uses RAM 1003 as work memory to read and execute the OS and other programs stored in ROM 1002 and storage device 1004. It then controls each component connected to the system bus 1009 to perform calculations and logical decisions for various processes. The processes executed by the CPU 1001 include information processing as defined in this embodiment. The storage device 1004 is a hard disk drive or external storage device, and stores programs and various data related to the information processing of this embodiment. The input unit 1005 is an input device such as an imaging device like a camera, buttons for inputting user instructions, a keyboard, or a touch panel. Note that the storage device 1004 is connected to the system bus 1009 via an interface such as SATA, and the input unit 1005 is connected via a serial bus such as USB, but the details of these connections are omitted. The communication interface 1006 communicates with external devices wirelessly. The display unit 1007 is a display.

[0020] The detection unit 21 detects the human body region in the image and outputs it to the partial region extraction unit 22. The human body region may correspond to the entire body, or to specific parts such as the face or torso. The method for detecting the human body region is not limited to any particular method. For example, an object detection method such as the one described in "Joseph Redmon, Ali Farhadi, "YOLOv3: An Incremental Improvement", arXiv e-prints (2018)" can be used. Alternatively, detection may be based on the contour shape of the head, limbs, etc. Furthermore, detection may be based on motion information extracted from time-series images. In addition, a configuration that detects from heat source information based on far-infrared radiation, etc., is also possible.

[0021] The human body region to be detected may correspond to a predetermined part, or it may correspond to a part selected by the user. For example, a function to set categories of body parts such as face and torso may be provided, and detection processing may be performed according to the category set by the user. Alternatively, for example, the corresponding detection processing may be automatically set based on information of the region selected by the user on a touch panel.

[0022] When the detection unit 21 detects multiple human body regions, it may select and output one or more corresponding human body regions based on a reference region, as described later. For example, the selection may be based on the distance from the reference region or the image similarity.

[0023] The partial region extraction unit 22 extracts one or more partial regions that meet predetermined conditions based on the human body region or depth information, and outputs them to the comparison unit 24. The extraction of partial regions is performed based on the similarity of distance and depth information in image space.

[0024] For example, if depth information for each pixel within a human body region is input, a set of pixels with similar depth information and close distances in image space may be extracted as a subregion. In this case, closed regions contained within a subregion may be merged into the subregion or extracted as separate subregions. Alternatively, thresholds may be set for the distance between pixels or the area of ​​the subregion for extraction.

[0025] Alternatively, for example, when the distance measuring unit 12 measures depth information for a specific distance measuring area, it may extract as a sub-area from among the distance measuring areas within the human body area in which the dispersion of depth information is below a threshold.

[0026] The reference area acquisition unit 23 acquires a reference area, which is a specific area targeted for AF, and outputs it to the comparison unit 24. The method of acquiring the reference area is not limited to a specific method. For example, the photographer may tap the camera's live view screen, and the tapped area may be used as the reference area. Alternatively, for example, a face detection method may be used to select the area of ​​a face that is closer to the center of the screen and appears larger as the reference area. When using a detection method, the detection method may be the same as that of the detection unit 21, or a different method may be used. Here, the reference area (first area) is assumed to be held as the area to be focused in the first image taken at the first point in time, which is when AF starts. At the start of shooting, a predetermined area of ​​the initial image (for example, the center of the screen) is set in advance as the reference area.

[0027] The comparison unit 24 compares the partial region input from the partial region extraction unit 22 with the reference region input from the reference region acquisition unit 23, and outputs the comparison result to the selection unit 25.

[0028] The elements that the comparison unit 24 compares (hereinafter referred to as comparison elements) include depth information and may be multiple different types of information. For example, they may include the relative position in image space between a reference region and a subregion, or they may include the pixel value on the image corresponding to the region. Furthermore, when using the relative position in image space, the magnitude of the relative position may be normalized based on the depth information. For example, the relative position value may be made smaller for regions that are far from the camera, and larger for regions that are far from the camera.

[0029] However, the comparison method of the comparison unit 24 is not limited to a specific method. For example, the comparison elements of each region may be averaged, and the difference in the average values ​​may be output as the comparison result. Also, if the shapes of the sub-region and the reference region are the same, the difference in the comparison elements for each corresponding pixel may be output as the comparison result. In addition, the distribution of depth information within the region may be compared, and an index such as Kullback-Leibler divergence may be output as the comparison result.

[0030] Furthermore, the comparison unit 24 may compare the estimated depth information with the depth information of a partial region by estimating the amount of change in depth information from the time the reference region was acquired, based on the time change in the depth information of the object of interest. In this case, more stable autofocus may be achieved even for subjects whose distance from the camera changes dynamically over time.

[0031] The selection unit 25 selects the focus area for the current input image based on the comparison result between the reference area and each sub-area input from the comparison unit 24, and outputs it to the imaging device 10. One example of how the selection unit 25 selects the focus area is to select a sub-area that is similar to the reference area (i.e., the difference in depth from the reference area is relatively small) based on the comparison result with the reference area. This method makes it possible to continue AF in an area similar to the reference area.

[0032] Furthermore, instead of using only the comparison results, the selection process may evaluate the priority of each sub-region to be selected as the area to focus on, and select those with higher priority. For example, if there are multiple sub-regions whose depth difference from the reference region is less than a predetermined value, the larger the area, the higher the priority. Alternatively, for example, the closer to the center of the image, the higher the priority. In addition, the selection unit 25 may determine the area to focus on in the second image based on multiple criteria. For example, the multiple criteria may include the similarity between the first region and the second region, and criteria for the area of ​​one or more sub-regions.

[0033] Furthermore, the selection criteria for the area to be focused may be changed depending on the elapsed time since the reference area was acquired. For example, if the elapsed time is short, priority may be given to areas with similarity to the reference area based on the comparison result, while as the elapsed time increases, other priority criteria such as the area of ​​a sub-region may be given more weight to the selection. Also, if the comparison result includes multiple types of elements, the selection may be made by considering the similarity of each element with different weights.

[0034] <Device Operation> Figure 3 is a flowchart illustrating the processes performed by the AF system in the first embodiment. S101 to S111 each represent specific processes, which are generally executed in order. However, the AF system does not necessarily have to perform all the processes described in this flowchart, and the execution order of the processes may change. Furthermore, multiple processes may be executed in parallel.

[0035] In step S101, the image acquisition unit 11 acquires the image (first image) at the time AF starts (time t-1). For example, it acquires the RGB image of the live view. In step S102, the distance measuring unit 12 measures the depth information at the time AF starts. For example, if the distance measuring unit 12 is equipped with a phase difference sensor, it measures the amount of defocus. The measured depth information will be subsequently acquired (depth information acquisition) by the area selection device 20.

[0036] In step S103, the reference area acquisition unit 23 acquires the reference area (first area) at the time AF starts (first time point). That is, it acquires (acquires reference) the area that was used as the focus target area at the time AF started. For example, it acquires the area selected by the user on the touch panel, or the area of ​​a face or human body that is automatically detected. When acquiring the reference area based on a detection process, the detection unit 21 or the like may be used. Also, if there are multiple candidate reference areas, the selection unit 25 or the like may be used to select the reference area. Furthermore, if shooting is continuous, the previously focused area may be acquired.

[0037] In step S104, the image acquisition unit 11 acquires an image (second image) at the time (time t, second time point) when the focus target area is selected. The time when the focus target area is selected is a time that follows the AF start time. Note that the first time point and the second time point do not have to be consecutive times; for example, the focus position may be changed at regular time intervals. In step S105, the distance measuring unit 12 measures depth information at the time when the focus target area is selected. The measured depth information will be subsequently acquired (depth information acquisition) by the area selection device 20.

[0038] In step S106, the detection unit 21 detects candidate human body regions (second regions) from the image acquired in S104. For example, it detects a predetermined object to be focused on (regions such as a human face or whole body, animals such as dogs or cats, cars or buildings). A region with similar image features to the reference region may be detected as the second region. For example, the second region may be detected using deep learning or semantic segmentation. Alternatively, a region with a depth within a predetermined range may be detected as the second region based on depth information corresponding to the image acquired in S104. For example, when there is only one subject, the region in the foreground (i.e., a region showing a similar depth) may be detected. In this case, S107 may be skipped. Then, in step S107, the partial region extraction unit 22 extracts one or more partial regions that satisfy predetermined conditions from the human body regions detected in S106. For example, a sub-region is extracted based on depth information; specifically, a sub-region is extracted in which the depth indicated by the depth information falls within the range of the depth of field.

[0039] The process in S107 will be explained with reference to Figure 1. In S107, the criteria for extracting sub-regions are changed based on the depth of field. For example, when subject G-1 (an upright human body) is detected, the distance from the camera within the human body region is almost constant (for example, the difference in depth calculated is 50 cm or less). Therefore, generally, the entire human body region can be extracted as a single sub-region. However, if the depth of field of the imaging device is narrow (for example, a few centimeters), sub-regions that are close to the imaging device, such as the eyes, nose, hands, and feet, may be extracted separately.

[0040] One method for extracting subregions that are close to the camera is to use clustering techniques such as K-means. Specifically, depth information is used to extract neighboring pixel clusters, and each cluster is then extracted as a subregion. In this process, the distance between pixels in the image may or may not be considered.

[0041] In step S108, the comparison unit 24 compares the reference region acquired in S103 with each subregion extracted in S107. For example, it compares the distribution of depth information between the reference region and each subregion, and the image features in the reference region with the image features in the extracted subregions.

[0042] Figure 4 illustrates the comparison between the reference region and each subregion in S108. Figure 4(a) shows region G-4, which is the reference region obtained from the first image. Figure 4(b) shows regions G-5a and G-5b, which are subregions extracted from the second image. In S108, for example, the average depth may be calculated from the depth information for both the reference region and each subregion, and the absolute value of the difference in the average depth from the depth information between the reference region and the subregion may be output as the comparison result. In the example shown in Figure 4, region G-5a is closer to region G-4 than region G-5b. Therefore, the difference in depth information is output as a smaller difference in the comparison result. When comparing using image features, the similarity between the image features in the reference region and the image features in the extracted subregions is compared with a pre-set threshold. If the similarity is above the threshold, they are similar and are likely to be the same part. On the other hand, if it is below the threshold, they are not similar and are likely to be different parts.

[0043] In step S109, the selection unit 25 selects a sub-region based on the comparison results and determines it as the focus area. Here, the sub-region having the average depth of field closest to the average depth of field in the reference region is determined as the focus area. Instead of the average depth of field in the region, the depth of field at a representative position may be used. Alternatively, the sub-region with the smallest difference in depth of field (smaller than a predetermined value) may be determined as the focus area. Determining a region with a similar depth of field in this way also leads to a reduction in focusing time. Furthermore, as an example of a method for selecting the focus area, for example, the sub-region most similar to the reference region may be selected. Alternatively, among the sub-regions where the difference from the reference region is below a threshold in the comparison results, the sub-region with the largest area may be selected. Multiple selection criteria may be combined.

[0044] In step S110, the region selection device 20 determines whether to continue the AF process. If the AF process is to be continued, the process proceeds to S111; otherwise, the process ends. In step S111, the region selection device 20 replaces the second image with the first image. Then the process returns to S103 and is repeated.

[0045] Through the above process, it becomes possible to suitably select the region in the current image (time t) that corresponds to the reference region that was selected as the target of AF immediately before (time t-1).

[0046] As described above, according to the first embodiment, information from the reference area is used to select a sub-region within the detected human body region that is suitable for AF as the focus target region. In particular, a sub-region with a smaller difference (more similar) from the reference area is selected. As a result, the AF system can suitably continue to perform AF even when there is a large difference in depth within the detected region.

[0047] Although the above explanation focused on detecting the human body, it can also be applied to detection targets other than the human body. For example, it can be used to detect animals other than humans, or specific objects such as vehicles. In addition to being usable for digital camera photography, it can also be used in systems that change the focus position after shooting through post-processing.

[0048] (Second Embodiment) In the second embodiment, we will describe a configuration that handles cases where the reference region and the focus target region belong to different part categories. Below, we will describe an example in which the head and torso are detected based on an image acquired from a camera or the like, and when the head is not detected, the focus target region is selected from a sub-region of the torso.

[0049] In this context, the term "torso" refers to the torso portion of a person's body, excluding the head. However, the torso may also include the neck, limbs, etc. Furthermore, it may refer to only a portion of the body, such as the chest or abdomen, rather than the entire torso.

[0050] The configuration of the AF system in the second embodiment is almost the same as that of the first embodiment (Figure 2). However, since the operation of each functional part differs from that of the first embodiment, the differences from the first embodiment will be described below.

[0051] Figure 7 is a flowchart illustrating the processes performed by the AF system in the second embodiment. S201 to S215 each represent specific processes.

[0052] In steps S201 and S202, the first image and its depth information are acquired. The processing performed in S201 and S202 is the same as in S101 and S102 in the first embodiment, so a description is omitted.

[0053] In step S203, the detection unit 21 detects the head region and torso region of the human body from the first image and outputs them to the partial region extraction unit 22 and the reference region acquisition unit 23. The method for detecting the torso region is not limited to a specific method. For example, the torso may be detected directly using a semantic region segmentation method, or the torso may be detected by detecting parts included in the torso, such as the shoulders and waist, using an object detection method.

[0054] Figure 5 shows examples of human body detection. Figure 5(a) shows an example where the detection unit 21 detects the head and torso. Regions G-8 and G-6a are the detected head and torso. On the other hand, Figure 5(b) shows an example where the head is obscured by a structure (region G-9) and the detection unit 21 detects only the torso G-6b.

[0055] The size and position of the parts detected as the torso, such as regions G-6a and G-6b, do not necessarily correspond to the entire trunk. Furthermore, they do not necessarily need to be represented by rectangles; they may be any shape, such as ellipses or polygons. They may also be represented as a distribution.

[0056] In step S204, the reference area acquisition unit 23 acquires the reference area to be AF from the first image and outputs it to the comparison unit 24. The position of the reference area is specified by the user or using the detection result of the detection unit 21, in the same manner as in S103 of the first embodiment. In the following description, the head is described as being preferentially selected as the reference area, but the present invention is not limited to head priority. If the torso is preferred, the head and torso should be swapped in the following description.

[0057] If the specified location is a head region as shown in region G-8 in Figure 5(a), that head region is designated as the reference region. The head region acquired by the reference region acquisition unit 23 may be a part of the head, such as the face, rather than the entire head, as shown in region G-8. If the specified location is a torso region as shown in region G-6a in Figure 5(a), region G-8, which is the head region closest to the specified location, is selected as the reference region. If the specified location is a torso region as shown in region G-6b in Figure 5(b), and no corresponding head region exists, region G-6b at the specified location is selected as the reference region.

[0058] From this point forward, the explanation for cases where the torso region is specified as the reference area will not be illustrated, but the processing is the same as when the head region is specified as the reference area. If images are being taken consecutively, the focus target area selected at the previous time will be used as the reference area.

[0059] In steps S205 and S206, a second image and its depth information are acquired. The processing in S205 and S206 is the same as in S104 and S105 of the first embodiment, so a description is omitted.

[0060] In step S207, the detection unit 21 detects the head region and the torso region from the second image. In step S208, the partial region extraction unit 22 extracts the respective partial regions from the head region and the torso region detected from the second image based on depth information. The specific method for extracting the partial regions has already been explained in detail in S107 of the first embodiment, so it will not be explained here.

[0061] In step S209, the comparison unit 24 performs a comparison process between the reference region acquired from the first image and the subregion belonging to the head region detected in the second image, using the depth information in each image. The specific method for comparing the reference region and the subregion is the same as the method described in S108 of the first embodiment.

[0062] In step S210, the selection unit 25 verifies the comparison result obtained in S209. The verification determines whether a portion of the head exists within a predetermined range of the second image. Here, the predetermined range is the tracking range in continuous shooting, but its size is not limited to a specific range. For example, since the tracking target is the human body, the size of the predetermined range is set based on the reasonable movement speed of the human body and the frame rate of continuous shooting. If the verification result shows that a portion of the head exists within the predetermined range, the process proceeds to S211; otherwise, it proceeds to S212. Cases where a portion of the head does not exist within the predetermined range include, for example, when the head region is obscured, as shown in Figure 9(b), or when head detection fails.

[0063] In step S211, the selection unit 25 selects the area to be focused. For example, following the same procedure as in S109 of the first embodiment, the area to be focused is selected from the head portion areas G-11a and G-11b in the second image shown in Figure 8(b) relative to the reference area G-10 in the first image shown in Figure 8(a). Then, the process proceeds to S214.

[0064] In step S212, the comparison unit 24 performs a comparison process between the reference region acquired from the first image and the subregion belonging to the torso region detected in the second image, using the depth information from each image. The specific method for comparing the reference region and the subregion is the same as the method described in S108 of the first embodiment. Once the comparison process is complete, the process proceeds to S213.

[0065] In step S213, the selection unit 25 determines the area to be focused from the comparison results obtained in S212. For example, following the same procedure as in S109 of the first embodiment, the area to be focused is selected from the body portion areas G-7a and G-7b in the second image shown in Figure 9(b) for the reference area G-10 in the first image shown in Figure 9(a). Then, the process proceeds to S214.

[0066] In step S214, the region selection device 20 determines whether to continue the AF process. If the AF process is to be continued, the process proceeds to S215; otherwise, the process ends. In step S215, the region selection device 20 replaces the second image with the first image. Then, the process returns to S204 and is repeated.

[0067] As described above, according to the second embodiment, if a part that was a reference area in the preceding image cannot be detected in the image to be processed due to occlusion or the like, a partial area within the detected human body area suitable for AF is selected as the area to be focused. As a result, the AF system can suitably continue to perform AF even when occlusion or the like is present.

[0068] In this embodiment, the detection unit 21 detects the torso region and the reference region acquisition unit 23 acquires a reference region based on the head region. However, this method is applicable to combinations other than the head and torso. For example, it could be a combination of a face and the whole body, or a combination of a single person and a densely packed group of people. Alternatively, it could be a combination of a car's license plate and the entire car body.

[0069] (Other examples) The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.

[0070] The invention is not limited to the embodiments described above, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, claims are attached to disclose the scope of the invention. [Explanation of symbols]

[0071] 10 Imaging device; 11 Image acquisition unit; 12 Distance measuring unit; 20 Region selection device; 21 Detection unit; 22 Partial region extraction unit; 23 Reference region acquisition unit; 24 Comparison unit; 25 Selection unit

Claims

1. An image processing device for determining the focus area of ​​an imaging device, An acquisition means that acquires a first region that is the object of focus in the first image captured by the imaging device at a first time point and indicates the head of a person, Depth information acquisition means for acquiring depth information for each region of the image captured by the aforementioned imaging device, A detection means for detecting a second region, which is a candidate for focus target and includes the torso of a person, from a second image captured by the imaging device at a second time point following the first time point, A determination means for determining a focus target region in the second image from one or more subregions of the second region based on depth information corresponding to the first region and depth information corresponding to the second region, An image processing apparatus characterized by comprising:

2. The determination means determines the focus target region based on the difference in depth calculated from the depth information corresponding to the first region and the depth information corresponding to the second image. The image processing apparatus according to feature 1.

3. The determination means determines, among the one or more sub-regions, the sub-region in which the difference between the depth of the first region and the depth of each of the one or more sub-regions is less than a predetermined threshold, as the focus target region. The image processing apparatus according to claim 2.

4. The determination means, when there are multiple sub-regions whose depth difference from the first region is less than a predetermined value, determines the sub-region with a relatively large area or a relatively small depth among the multiple sub-regions to be the region to be focused. The image processing apparatus according to claim 3.

5. The depth information is based on the distance between the imaging device and the subject or the phase difference of the incident light at the image sensor of the imaging device. The image processing apparatus according to any one of claims 1 to 4.

6. The detection means detects a region in the second image that has image features similar to those extracted from the first region as the second region. The image processing apparatus according to any one of claims 1 to 5.

7. The system further comprises extraction means for extracting one or more sub-regions from the second region that satisfy predetermined conditions, The determination means determines the area to be focused in the second image from one or more sub-regions extracted by the extraction means. The image processing apparatus according to any one of claims 1 to 6.

8. The extraction means extracts one or more sub-regions from the second region in which the depth indicated by the depth information corresponding to the second image falls within a predetermined range. The image processing apparatus according to feature 7.

9. The extraction means changes the predetermined conditions of one or more sub-regions based on the depth of field when the imaging device captures the second image. The image processing apparatus according to claim 7 or 8.

10. The determination means determines the area to be focused in the second image based on a plurality of criteria, The aforementioned multiple criteria include the similarity between the first region and the second region, and the respective criteria for the area of ​​one or more subregions. The image processing apparatus according to any one of claims 1 to 9.

11. The determination means changes the plurality of criteria for the focus target area based on the elapsed time from the first time point. The image processing apparatus according to feature 10.

12. The second region is the region representing the head and / or torso of the person. The image processing apparatus according to any one of claims 1 to 11.

13. The determination means prioritizes determining the area of ​​one or more sub-regions that is larger than a predetermined value as the area to be focused in the second image. The image processing apparatus according to any one of claims 1 to 12.

14. A control method for an image processing device that determines the focus area of ​​an imaging device, An acquisition step in which a first region is acquired in the first image captured by the imaging device at a first point in time, which is the object of focus and indicates the head of a person, A depth information acquisition step is performed to acquire depth information for each region of the image captured by the aforementioned imaging device. A detection step in which, at a second time point following the first time point, a second region is detected from a second image captured by the imaging device that is a candidate for the object to be focused on and includes the torso of a person, A determination step of determining the area to be focused in the second image from one or more sub-regions of the second region based on depth information corresponding to the first region and depth information corresponding to the second region, A control method characterized by including

15. A program for causing a computer to function as one of the means of an image processing apparatus according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Device and method of imaging and program

    JP2010114752A

  • Imaging apparatus and method for controlling the same

    JP2012128288A

  • Imaging device

    JP2014048393A

  • Focus adjustment device and imaging apparatus

    JP2015052799A

  • Imaging apparatus and control method of the same

    JP2016136683A