Image processing device, control method thereof, imaging device, and program

The image processing apparatus enhances subject detection by confirming face specification and expanding detected regions to include non-detected parts, addressing misidentification issues in conventional systems.

JP2025172957AActive Publication Date: 2025-11-26CANON KK
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2025149959
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-10
Publication Date
2025-11-26
Estimated Expiration
2041-02-24

AI Technical Summary

Technical Problem

Conventional image processing systems face challenges in accurately specifying parts related to a subject when multiple parts are detected, leading to potential misidentification of the desired subject during user interaction.

Method used

An image processing apparatus that includes a detection means to identify face and torso parts of a subject, with a determination means to confirm the face as specified when a user designates an area between the face and torso, ensuring accurate subject specification.

Benefits of technology

Enables precise specification of subject parts in images, even when user interaction may inadvertently target undetected areas, by expanding the detected region to include non-detected parts based on user input.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025172957000001_ABST
    Figure 2025172957000001_ABST
Patent Text Reader

Abstract

To enable designation between parts of a subject within an image in image processing for detection of the subject.SOLUTION: An imaging device 100 includes a subject detection unit 161 for detecting a subject in an input image. The subject detection unit 161 detects at least two parts in the image (S402), and determines a priority part among the detected parts (S403). The subject detection unit 161 performs region expansion processing (S404, S405) for expanding a determined priority region, thereby expanding a region in a direction from the priority region to other detected parts. In the detection of a plurality of parts, undetectable parts are also treated as extended regions of the detected part, whereby if a user designates a region by mistake and the designated region is included in the expanded region, it is possible to determine that the user has designated the region.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to techniques for detecting and determining objects within an image. [Background technology]

[0002] The imaging device acquires the feature amount of a specific area of ​​the image from the captured image data, focuses on the detected subject, and captures the image with optimal brightness, color, etc. Patent Document 1 discloses a technology that continues to estimate the position of the subject's face even when the subject's facial orientation changes. By detecting multiple parts of a single subject, the subject can be continuously detected with high accuracy. Furthermore, when multiple subjects are detected, the user can select the subject they want to photograph by performing a touch operation. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2009-81714 Summary of the Invention [Problem to be solved by the invention]

[0004] In conventional technology, if there is a part that is not subject to detection among multiple detected parts, and the user designates that part by performing a touch operation or the like, there is a possibility that the desired subject will not be selected. An object of the present invention is to enable specification of parts relating to an object in an image in image processing for object detection. [Means for solving the problem]

[0005] An apparatus according to one embodiment of the present invention is an image processing apparatus that acquires an image and detects a subject, and is characterized by comprising: a detection means that detects the face and torso parts of the same subject in the image; and a determination means that, when a user specifies an area between the face and the torso, determines that the face has been specified. [Effects of the Invention]

[0006] According to the present invention, it is possible to specify parts relating to a subject in an image in image processing for subject detection. [Brief explanation of the drawings]

[0007] [Figure 1] 1 is a block diagram showing the configuration of an imaging apparatus according to a first embodiment. [Figure 2] FIG. 2 is a block diagram showing the configuration of a subject detection unit according to the first embodiment. [Figure 3] 1 is a flowchart of the overall processing of the first embodiment. [Figure 4] 10 is a flowchart of subject detection according to the first embodiment. [Figure 5] FIG. 10 is a diagram relating to region expansion in the first embodiment. [Figure 6] FIG. 10 is a block diagram showing the configuration of a subject detection unit according to a second embodiment. [Figure 7] 10 is a flowchart of subject detection according to the second embodiment. [Figure 8] FIG. 10 is a diagram relating to region expansion in the second embodiment. [Figure 9] 10 is a flowchart of inter-site connection processing according to the second embodiment. [Figure 10] FIG. 10 is a diagram relating to inter-site connection processing in the second embodiment. [Figure 11] 10 is a flowchart of subject detection according to the third embodiment. [Figure 12] FIG. 10 is a diagram relating to region expansion in the third embodiment. [Figure 13] 13 is a flowchart of subject detection according to a modified example of the third embodiment. [Figure 14]FIG. 10 is a diagram relating to region expansion in a modified example of the third embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0008] The present invention will be described in detail below with reference to the accompanying drawings. Each embodiment illustrates an example of an imaging device to which the image processing device according to the present invention is applied. The present invention is applicable to video cameras and digital still cameras with subject detection functions, as well as various electronic devices with imaging means.

[0009] [Embodiment 1] 1 is a block diagram showing an example of the functional configuration of an imaging device 100 according to this embodiment. The imaging device 100 is a digital still camera, video camera, or the like that can capture an image of a subject and record moving and still image data on a recording medium. Each unit within the imaging device 100 is connected via a bus 160. Each unit is controlled by a CPU (Central Processing Unit) 151 that constitutes a control unit. The CPU executes a program to perform the following processing and control.

[0010] The photographing lens unit 101 includes optical components such as fixed lenses and movable lenses that constitute the imaging optical system. FIG. 1 shows a configuration including a fixed first group lens 102, a zoom lens (lens for varying magnification) 111, an aperture 103, a fixed third group lens 121, and a focus lens (lens for adjusting focus) 131. For ease of explanation, the lenses 102, 111, 121, and 131 are illustrated as single lenses, but each may be composed of multiple lenses. The photographing lens unit 101 may also be configured as a detachable interchangeable lens.

[0011] The aperture control unit 105 controls the light amount adjustment during shooting by adjusting the aperture diameter of the aperture 103 by driving the aperture 103 via the aperture motor (AM) 104 in accordance with instructions from the CPU 151. The zoom control unit 113 changes the focal length of the photographing lens unit 101 by driving the zoom lens 111 via the zoom motor (ZM) 112.

[0012] The focus control unit 133 controls the driving of the focus motor (FM) 132. The focus control unit 133 calculates the defocus amount and defocus direction of the photographing lens unit 101 based on the phase difference between a pair of focus detection signals (an A image signal and a B image signal) obtained from the image sensor 141. The focus control unit 133 converts the defocus amount and defocus direction into the drive amount and drive direction of the focus motor (FM) 132. The focus control unit 133 controls the operation of the focus motor (FM) 132 based on the drive amount and drive direction, and drives the focus lens 131 to control the focus state of the photographing lens unit 101. In this way, the focus control unit 133 performs automatic focus detection and adjustment (phase difference AF) using a phase difference detection method. Alternatively, the focus control unit 133 calculates a contrast evaluation value from the image signal obtained from the image sensor 141 and performs AF control using a contrast detection method.

[0013] Light from the subject is focused on the image sensor 141 via the photographing lens unit 101. The image sensor 141 performs photoelectric conversion on the subject image (optical image) focused by the imaging optical system and outputs an electrical signal. Each of the multiple pixel units arranged in the image sensor 141 has a photoelectric conversion unit. For example, the image sensor 141 has m pixels in the horizontal direction and n pixels in the vertical direction, that is, m×n pixel units arranged in a matrix. Each pixel unit is provided with a microlens and two photoelectric conversion units. Reading of signals from the image sensor 141 is controlled by the imaging control unit 143 in accordance with instructions from the CPU 151. The image sensor 141 outputs an electrical signal to the imaging signal processing unit 142.

[0014] The imaging signal processing unit 142 performs signal processing to convert the signal acquired by the imaging element 141 into an image signal and acquire image data on the imaging surface. The imaging signal processing unit 142 performs signal processing such as noise reduction, A / D conversion, and automatic gain control. The image data output from the imaging signal processing unit 142 is sent to the imaging control unit 143. The imaging control unit 143 stores the image signal received from the imaging signal processing unit 142 in a RAM (random access memory) 154.

[0015] The image processing unit 152 applies predetermined image processing to the image data stored in the RAM 154. The predetermined image processing includes, but is not limited to, so-called development processing such as white balance adjustment processing, color interpolation (demosaic) processing, and gamma correction processing, as well as signal format conversion processing and scaling processing. The image processing unit 152 also generates information related to subject brightness to be used for automatic exposure control (AE). Information related to a specific subject area within an image is supplied to the image processing unit 152 from the subject detection unit 161 and is used, for example, for white balance adjustment processing. When performing AF control using a contrast detection method, the image processing unit 152 generates an AF evaluation value. The image processing unit 152 stores the processed image data in the RAM 154.

[0016] The image compression / decompression unit 153 reads and compresses the image data stored in the RAM 154, and then records the data on the image recording medium 157. In parallel with this process, the image data stored in the RAM 154 is sent to the image processing unit 152. When recording the image data saved in the RAM 154, the CPU 151 performs a process of adding a predetermined header, etc. to the image data, and generates a data file according to the recording format. At this time, the CPU 151 performs a process of encoding the image data using the image compression / decompression unit 153 as necessary, and compressing the amount of information. The CPU 151 then performs a process of recording the generated data file on the image recording medium 157.

[0017] The flash memory 155 stores control programs required for the operation of the imaging device 100, setting values ​​used for the operation of each unit, GUI data, user setting values, etc. When the imaging device 100 is started up by transitioning from a power-off state to a power-on state through a user operation, the control programs and parameters stored in the flash memory 155 are loaded into a part of the RAM 154. The CPU 151 controls the operation of the imaging device 100 in accordance with the control programs and constants loaded into the RAM 154.

[0018] The CPU 151 executes AE processing to automatically determine exposure conditions (shutter speed or accumulation time, aperture value, sensitivity) based on information about the brightness of the subject. Information about the brightness of the subject can be acquired, for example, from the image processing unit 152. The CPU 151 can also determine exposure conditions based on the area of ​​a specific subject, such as a person's face. The CPU 151 controls exposure based on the electronic shutter speed (accumulation time) and the magnitude of gain. The CPU 151 notifies the imaging control unit 143 of the determined accumulation time and magnitude of gain. The imaging control unit 143 controls the operation of the image sensor 141 so that photography is performed in accordance with the notified exposure conditions.

[0019] Monitor display 150 has a display device such as an LCD (liquid crystal display) or an organic EL (electroluminescence) display, and performs display processing of image data, etc. For example, when displaying image data stored in RAM 154, CPU 151 performs scaling processing on the image data in image processing unit 152 so that the image data matches the display size of monitor display 150. The processed image data is written to an area of ​​RAM 154 used as video memory (VRAM area). Monitor display 150 reads the image data to be displayed from the VRAM area of ​​RAM 154 and displays it on the screen.

[0020] The imaging device 100 causes the monitor display 150 to function as an electronic viewfinder (EVF) by instantly displaying captured moving images on the monitor display 150 while in still image standby mode or during moving image recording. The moving images and their frame images displayed on the monitor display 150 functioning as an EVF are called live view images or through images. Furthermore, when capturing a still image, the imaging device 100 displays the most recently captured still image on the monitor display 150 for a certain period of time so that the user can check the capture results. The display operation is controlled according to commands from the CPU 151.

[0021] The operation unit 156 includes switches, buttons, keys, a touch panel, an eye-gaze input device, etc., which allow the user to input operation signals to the imaging device 100. The operation unit 156 outputs operation signals to the CPU 151 via a bus 160. The CPU 151 controls each unit to realize operations according to the operation signals. The image recording medium 157 is a memory card or the like that records data files in a predetermined format, etc. The power management unit 158 ​​manages a battery 159 and provides a stable power supply to the entire imaging device 100.

[0022] The subject detection unit 161 has the following functions when the imaging target is a specific subject (for example, a person). Body part detection function that detects specific body parts such as the face or torso of a specific subject. - Priority area determination function that determines the priority area from the detected areas. An extension function that prioritizes areas of the subject that are not detected by the body part detection function. The detection results of the subject detection unit 161 can be used, for example, to automatically set the focus detection area. As a result, a tracking AF function for a specific subject area can be realized. AE processing is performed based on the luminance information of the focus detection area, and image processing (for example, gamma correction processing, white balance adjustment processing, etc.) is performed based on the pixel values ​​of the focus detection area. The CPU 151 performs processing to superimpose an index indicating the current position of the subject area on the displayed image and display it on the monitor display 150. The index is, for example, a rectangular frame surrounding the subject area. The configuration and operation of the subject detection unit 161 will be described in detail later.

[0023] The position and orientation change acquisition unit 162 includes a position and orientation sensor such as a gyro sensor, an acceleration sensor, or an electronic compass, and measures a change in the position and orientation of the imaging device 100 relative to the captured scene. The acquired data on the position and orientation change is stored in the RAM 154 and is referenced by the subject detection unit 161.

[0024] The configuration of subject detection unit 161 will be described with reference to Fig. 2. Fig. 2 is a block diagram showing an example of the functional configuration of subject detection unit 161. Subject detection unit 161 includes a multiple part detection unit 201, a priority part determination unit 202, a priority part expansion unit 203, and an expansion range adjustment unit 204.

[0025] The multiple part detection unit 201 sequentially acquires time-series image signals from the image processing unit 152 and detects at least two parts of the subject of the imaging target included in each image. The detection results include information such as area information, coordinate information, and reliability within the image. Area information includes information indicating range, area, boundary, etc. Coordinate information, which is position information, includes information on coordinate points and coordinate values.

[0026] The priority region determination unit 202 determines a region to be prioritized (priority region) from among the regions of the subject detected by the multiple region detection unit 201. The priority region expansion unit 203 expands the region of the subject from the priority region determined by the priority region determination unit 202. The priority region expansion unit 203 sets the expanded region as an expanded region of the priority region.

[0027] The extension range adjustment unit 204 adjusts the range of the extension region by the priority portion extension unit 203. Information about the region that has been extended by the priority portion extension unit 203 and then adjusted by the extension range adjustment unit 204 is used by the CPU 151 and the like for various processes.

[0028] The operation of the imaging device 100 involving subject detection processing will be described with reference to the flowchart of Fig. 3. The following processing is realized by the CPU 151 executing a program.

[0029] In S301, the CPU 151 determines whether the power supply to the imaging device 100 is ON or OFF. If it is determined that the power supply is OFF, the processing is terminated, and if it is determined that the power supply is ON, the processing proceeds to S302. In S302, the CPU 151 executes imaging processing for one frame. A pair of parallax image data and data for one screen of the captured image are generated by the imaging element 141. The parallax image is made up of a plurality of viewpoint images with different viewpoints. For example, the imaging element 141 has a plurality of microlenses and a plurality of photoelectric conversion units corresponding to each microlens. An A image signal is acquired from the first photoelectric conversion unit of each pixel unit, and a B image signal is acquired from the second photoelectric conversion unit. Viewpoint image data based on the A image signal and the B image signal is generated. The generated data is stored in the RAM 154. Next, the processing proceeds to S303.

[0030] In S303, CPU 151 executes subject detection using subject detection unit 161. Details of the subject detection process will be described later. Subject detection unit 161 notifies CPU 151 of detection information such as the position and size of the subject area. The detection information is stored in RAM 154. CPU 151 sets a focus detection area based on the notified subject area. Next, the process proceeds to S304.

[0031] In S304, the CPU 151 causes the focus control unit 133 to execute focus detection processing. The focus control unit 133 acquires focus detection signals from multiple pixel units included in the focus detection area of ​​a pair of parallax images. An image A signal is generated from signals obtained from first photoelectric conversion units in multiple pixel units arranged in the same row, and an image B signal is generated from multiple signals obtained from second photoelectric conversion units. The focus control unit 133 calculates the correlation between the images A and B while shifting the relative positions of the images A and B, and determines the relative position at which the similarity between the images A and B is highest as the phase difference (shift amount) between the images A and B. Furthermore, the focus control unit 133 converts the phase difference into a defocus amount and a defocus direction. Next, the process proceeds to S305.

[0032] In S305, the focus control unit 133 drives the focus motor (FM) 132 according to the lens drive amount and drive direction corresponding to the defocus amount and defocus direction calculated in S304. A focus adjustment operation is performed by controlling the movement of the focus lens 131. When the lens drive process is completed, the process returns to S301.

[0033] Thereafter, the processes of S302 to S305 are repeatedly executed until it is no longer determined in S301 that the power switch is ON. Note that, although the subject detection process is executed for each frame in Fig. 3, the subject detection process may be executed every few frames in order to reduce the processing load and power consumption.

[0034] The processing performed by the subject detection unit 161 will be described with reference to the flowchart of Fig. 4 and Fig. 5. First, an input image is supplied from the imaging control unit 143 to the subject detection unit 161 in S401.

[0035] In S402, the multiple body part detection unit 201 detects body parts of the subject from the image input from the imaging control unit 143. At least two body parts are detected at this time. Information on the detection result includes area information and coordinate information within the image. Other information includes reliability and the like. Known methods are used for body part detection. For example, there is a method of detecting the body parts by extracting features of a specific subject using a convolutional neural network (hereinafter referred to as CNN). There is also a method of registering the subject or body part of the subject to be detected as a template in advance and performing body part detection processing by template matching. An example is shown with reference to FIG. 5. FIG. 5 is a schematic diagram for explaining the detection of body parts of a subject. An example is shown in which two body parts, a face detection area 501 and a torso detection area 502, are detected from a person who is the subject.

[0036] In S403 of Fig. 4, the priority part determination unit 202 determines a priority part from among the multiple parts detected in S402. In this embodiment, the priority part is determined according to a setting value previously set by the user via the operation unit 156 as a method for determining the priority part. For example, if face priority is set, in the example of Fig. 5, the face detection area 501 is determined as the priority part. Note that the method for determining the priority part is not limited to setting by user operation, and there is also a method for determining, for example, a part with high reliability calculated in S402 as the priority part.

[0037] In S404, the priority region expansion unit 203 performs a first expansion process on the region. The region is expanded from the priority region determined in S403 toward regions that were not determined to be priority. In the example of Fig. 5, expansion processing is performed from a face detection region 501, which is a priority region, toward a torso detection region 502, which is not a priority region, to region 503.

[0038] In S405, the priority region expansion unit 203 performs a second expansion process on the region, expanding the region in a direction different from the region expanded in S404. In the example of Fig. 5, the region is expanded from the face detection region 501, which is the priority region, in directions (three directions) different from region 503, and region 504 is obtained. The expansion range of region 504 is narrower than the expansion range of region 503 expanded in S404.

[0039] The areas expanded in S404 and S405 are treated as expanded areas of the priority part. In the example of Fig. 5, areas 503 and 504 are treated as expanded areas of face detection area 501, which is the priority part. For example, when the user specifies a subject by touching area 503 via operation unit 156 or by inputting line of sight, CPU 151 determines that a person's face (detection area 501) has been specified, and executes expansion processing corresponding to face detection area 501.

[0040] In S406, the extension range adjustment unit 204 performs a process to adjust the extension range of the region extended in S404 and S405. For example, the extension range adjustment unit 204 can estimate camera shake based on motion information of the imaging device acquired by the position and orientation change acquisition unit 162 and adjust the extension range according to the camera shake. As camera shake increases, it becomes more difficult to accurately specify a subject. Therefore, the extension range adjustment unit 204 increases the extension range to make it easier to specify a priority portion. As another example, consider a case where a user specifies a subject by eye-gaze input via the operation unit 156. Due to the nature of the eye gaze, variations may occur even when gazing at a fixed point, making it difficult to accurately specify a subject using only the eye gaze. The extension range adjustment unit 204 adjusts the extension range based on variations in eye gaze position, making it easier to specify a priority portion. Methods for estimating camera shake and variations in eye gaze position are well known, so detailed descriptions of these methods will be omitted.

[0041] In this embodiment, the case where the specific subject is a human face and torso has been described. However, the present invention is not limited to this, and can be applied to detecting the face and torso of an animal such as a cat, or inanimate parts of an automobile (such as a headlamp and hood), and can also be applied to detecting a combination of animate and inanimate objects, such as a person riding a motorcycle and the motorcycle itself. There are no particular limitations on the subject and its parts to be detected. This also applies to the embodiments described below.

[0042] According to this embodiment, a portion of a subject that cannot be detected in the subject detection process can be treated as an extended region of the detected subject and processed accordingly. When a user specifies a subject displayed on a display unit (such as a rear LCD display unit or an EVF) by touch operation or line of sight, there is a possibility that an undetectable portion may be specified by mistake due to influences such as camera shake or variations in the specification. Even in such cases, specification based on a detected portion is possible. For example, if a user specifies an incorrect region, the imaging device can determine that the user specified that region if the specified region is included in the extended region.

[0043] [Embodiment 2] Next, a second embodiment of the present invention will be described. In this embodiment, the same symbols and signs as those in the first embodiment will be used for the configurations and processes common to the first embodiment, and detailed descriptions of these will be omitted, with the differences being mainly described. This method of omitting descriptions will also be used in the embodiments and modified examples described below.

[0044] 6 is a block diagram showing an example of the functional configuration of the subject detection unit 161 of this embodiment. The subject detection unit 161 includes a multiple part detection unit 601, an inter-part connection unit 602, a priority part determination unit 603, a priority part expansion unit 604, and an expansion range adjustment unit 605. The difference from embodiment 1 is that the inter-part connection unit 602 has been added.

[0045] Image signals are sequentially supplied in time series from the image processing unit 152 to the multiple part detection unit 601. The multiple part detection unit 601 detects at least two parts of the subject being imaged included in each image. Information on the detection results includes information such as area information, coordinate information, and reliability within the image.

[0046] The inter-part connection unit 602 performs processing to connect parts that are the same subject among the parts detected by the multiple part detection unit 601. The priority part determination unit 603 determines which part is to be prioritized as the subject among the parts detected by the multiple part detection unit 601. The priority part expansion unit 604 calculates an expansion area for the priority part determined by the priority part determination unit 603. The expansion range adjustment unit 605 adjusts the range of the expansion area determined by the priority part expansion unit 603.

[0047] The processing performed by the subject detection unit 161 will be described with reference to the flowchart of Fig. 7 and Fig. 8. First, in S701, an input image is supplied from the imaging control unit 143 to the subject detection unit 161.

[0048] In S702, the multiple body part detection unit 601 detects body parts of the subject from the image input from the imaging control unit 143. The multiple body part detection unit 601 detects at least two body parts. The detection result includes area information and coordinate information within the image. Other information such as reliability may also be included. Fig. 8 is a schematic diagram showing an example of detection of body parts of a person, which is a specific subject. In the example of Fig. 8(A), two body parts are detected: a detection area 801 and barycenter coordinates 802 of the person's face, and a detection area 803 and barycenter coordinates 804 of the person's torso.

[0049] In S703 of Fig. 7, the inter-part connection unit 602 performs processing to connect parts of the same subject among the multiple parts detected in S702. In the example of Fig. 8(A), the point of center of gravity 802 of the face and the point of center of gravity 804 of the torso are connected by the connection unit 805. The method of connecting parts will be described in detail later.

[0050] In S704, the priority part determination unit 603 determines a priority part from among the multiple parts detected in S702. The method for determining the priority part is the same as in S403 in FIG. 4. In FIG. 8, a person's face is determined as the priority part. In S705, the priority part expansion unit 604 expands the area from the priority part determined in S704 in the direction of the connection made in S703. In the example of FIG. 8(B), the area is expanded from the detection area 801 corresponding to the person's face, which is the priority part, in the direction of the connection part 805. The expansion area 806 in this case is the area between the face detection area 801 and the torso detection area 803.

[0051] In S706, the priority region expansion unit 604 expands the region in a direction different from the region expanded in S705. In the example of FIG. 8(B), the region is expanded from the face detection region 801, which is the priority region, in directions (three directions) different from the connection portion 8056. In this case, the expanded region 807 is an area obtained by expanding the face detection region 801 in two directions, one opposite to the direction of the connection portion 805 and one perpendicular to the direction of the connection portion 805. The expansion range of the region 807 is narrower than the expansion range of the region expanded in S705. In S707, the expansion range adjustment unit 605 adjusts the expansion range of the region expanded in S705 and S706. The adjustment method is the same as in S406 of FIG. 4.

[0052] The inter-part connection method in S703 of Fig. 7 will be described with reference to Fig. 9 and Fig. 10. Fig. 9 is a flowchart of the inter-part connection process. Fig. 10 is a schematic diagram showing a scene in which two people 1010, 1020 exist in an image. The face centroid 1011, torso centroid 1012, and connection part 1013 of person 1010, and the torso centroid 1022 of person 1020 are shown. There are two sets of each part shown on the screen. An example is shown in Fig. 10 in which the face centroid 1011 of person 1010 is the connection source and the torso centroid 1012 is the connection destination.

[0053] In S901 of Fig. 9, the inter-part connection unit 602 performs connection processing between parts. The process of searching for one connection destination from the connection source and making the connection is executed. In S902, the inter-part connection unit 602 searches the entire screen and determines whether the search for the connection destination has been completed. If it is determined that the search for the connection destination has been completed, the process proceeds to S903, and if it is determined that the search for the connection destination has not been completed, the process returns to S901 and continues the connection processing.

[0054] 10, the torso center of gravity 1012 of person 1010 and the torso center of gravity 1022 of person 1020, that is, two centers of gravity, are detected. The processes of S901 and S902 in FIG. 9 are repeated twice. A connection section 1013 shows the connection result from the face center of gravity 1011 of person 1010 to the torso center of gravity 1012. A connection section 1014 shows the connection result from the face center of gravity 1011 of person 1010 to the torso center of gravity 1022 of person 1020.

[0055] In S903 of FIG. 9, the inter-part connection unit 602 evaluates each connection result connected in S901 and S902 and determines one connection source. One method for evaluating the connection results is, for example, to select a connection destination whose distance on the image plane between the connection source and the connection destination is closest. Another method is to compare the depth information of the connection source with the depth information of the connection destination and select the connection destination whose depth difference is smallest. In the example of FIG. 10, the distance on the image plane between the face centroid 1011 and the torso centroid 1012 is shorter than the distance on the image plane between the face centroid 1011 and the torso centroid 1022. As a result, the inter-part connection unit 602 selects connection unit 1013 and determines to connect the face centroid 1011 and the torso centroid 1012 of the person 1010. After S903, the inter-part connection process ends.

[0056] In this embodiment, an example has been described in which there are two people in the image, and two sets of two points, the face center of gravity and the torso center of gravity, are used. This example is not limiting, and application is possible even when the number of people increases, and the type of body part, such as the shoulder center of gravity or the arm center of gravity, and the number of detections are not important. Regarding the processing from S901 to S903 in Fig. 9, the connection destination and connection source are changed and the processing is repeated for the number of detections, so that connection processing between body parts can be performed for all detected body parts and all people. In the present embodiment, an example has been shown in which connection processing between parts is performed after sequentially searching for the connection source and the connection destination. However, the present invention is not limited to this example, and connection processing may be performed by feature extraction processing using CNN with the connection source and the connection destination as input.

[0057] According to this embodiment, in a situation where multiple subjects are detected in an image, the multiple detected parts that are separated from each other can be connected, and the non-detection area can be considered as part of the detected subject based on the connection result. Even if the user accidentally specifies a non-detection area between multiple parts, the imaging device can determine that the user specified that area when the specified area is included in the extended area.

[0058] [Embodiment 3] A third embodiment of the present invention will be described with reference to Figures 11 and 12. The functional configuration of subject detection unit 161 in this embodiment is the same as that in Figure 6, but at least one part (first part) detected by multiple part detection unit 601 has area information and coordinate information within the image. Furthermore, a second part does not have area information within the image, but only has coordinate information.

[0059] The processing of subject detection unit 161 will be described with reference to the flowchart in Fig. 11 and Fig. 12. Fig. 12 is a schematic diagram illustrating the detection processing of the joints of a person, which is a specific subject. The first part detected in this example is the person's face, which has information on area 1301 and coordinates 1302. The second parts, which are chest 1303, left hand 1304, right hand 1305, left foot 1306, and right foot 1307, each have only coordinate information.

[0060] First, in S1201 of Fig. 11, an input image is supplied from the imaging control unit 143 to the subject detection unit 161. In S1202, the multiple body part detection unit 601 detects at least two body parts of the subject from the image input from the imaging control unit 143. The first body part has area information and coordinate information within the image. The second body part has only coordinate information within the image. Specifically, in Fig. 13(A), the face, which is the first body part, has information on area 1301 and coordinates 1302. The chest (1303) and limbs (1304 to 1307) each have only coordinate information.

[0061] In S1203, inter-part connection unit 602 performs processing to connect parts of the same subject among the multiple parts detected in S1202. In the example of Fig. 13(A), the point of coordinate 1302 of the face center of gravity and the coordinate point of chest 1303 are connected by connection unit 1308. The following connection processing is performed in a similar manner. A process of connecting the coordinate points of the chest 1303 and the coordinate points of the left hand 1304 by a connecting section 1309. A process of connecting the coordinate points of the chest 1303 and the coordinate points of the right hand 1305 by the connecting unit 1310. A process of connecting the coordinate points of the chest 1303 and the coordinate points of the left foot 1306 by the connecting portion 1311. A process of connecting the coordinate points of the chest 1303 and the coordinate points of the right foot 1307 by the connecting portion 1312.

[0062] In S1204 of Fig. 11, the priority part determination unit 603 determines a priority part from among the multiple parts detected in S1202. One of the parts having area information is determined as the priority part. In the example of Fig. 13(A), the face having information on area 1301 and coordinates 1302 is determined as the priority part. When there are multiple parts having area information, the priority part can be determined according to designation by user operation, as in S403 of Fig. 4. Furthermore, a part that is close in distance on the image plane can be determined as the priority part according to coordinate designation by a touch operation or eye-gaze input performed by the user via the operation unit 156.

[0063] In S1205, the priority region expansion unit 604 performs region expansion processing from the priority region determined in S1204 in the direction of the connection made in S1203. In the example of Fig. 13(B) , the region is expanded from the face region 1301, which is the priority region, in the direction of the connection portion 1308. In other words, an expanded region 1313 adjacent to the region 1301 is acquired.

[0064] In S1206, the priority portion expansion unit 604 performs further expansion processing of the region from the region expanded in S1205 in the direction of the connection in S1203. In the example of Fig. 13(B) , the region is expanded from the expanded region 1313 in the direction of the connection part 1309. In other words, an expanded region 1314 adjacent to the region 1313 is obtained.

[0065] In S1207, a determination is made as to whether the region expansion processes of S1205 and S1206 have been performed for all parts that do not have region information. If it is determined that the region expansion processes have been completed, the process proceeds to S1208. If it is determined that the region expansion processes have not been completed, the process returns to S1205 and the region expansion processes are repeated.

[0066] In S1208, the extension range adjustment unit 605 performs processing to adjust the extension range of the region extended in S1205 and S1206, and then ends the subject detection processing.

[0067] According to this embodiment, even when a part that does not have area information, such as an animal joint, is detected, the area can be expanded from the priority part based on the connection result. For example, if the user mistakenly specifies an area, the imaging device can determine that the user specified the area if the specified area is included in the expanded area.

[0068] [Modification of the third embodiment] A modified example of the third embodiment will be described with reference to the flowchart of Fig. 13 and Fig. 14. In this modified example, an example of processing using the reliability of connections between parts is shown. The configuration of the subject detection unit in this modified example is the same as that in Fig. 6. Fig. 14 is a schematic diagram showing the detection of two people 1510, 1520 in an image, and shows an example of processing to connect the face, chest, and limbs of each subject in the image. The processing of S1201 to 1203 in Fig. 13 has already been explained using Fig. 11, and the processing proceeds to S1401 after S1203.

[0069] In S1401, the inter-part connection unit 602 calculates the reliability of the connection between the parts connected in S1203. For example, the reliability can be calculated based on the continuity of a straight line from the connection source to the connection destination. For example, a method of calculating the continuity uses a depth map that represents the distribution of depth values ​​within an image. Depth values ​​on the path connecting the connection source to the connection destination can be obtained from the depth map, and the reliability can be calculated based on the variance of the obtained depth values. Regarding the depth map, for example, a method of pupil-splitting light from a subject to generate multiple viewpoint images (parallax images), calculating the amount of parallax, and acquiring depth distribution information of the subject can be used. A pupil-splitting image sensor includes multiple microlenses and multiple photoelectric conversion units corresponding to each microlens, and each photoelectric conversion unit can output a signal of a different viewpoint image. The depth distribution information of the subject includes data representing the distance from the image capture unit to the subject (subject distance) as an absolute value, and data indicating the relative distance relationship in the image data (depth of the image) (distribution data of the amount of parallax, etc.). The depth direction corresponds to the direction of depth relative to the imaging device. Image data from multiple viewpoints can also be obtained by a multi-lens camera having multiple imaging means.

[0070] The reliability calculation process will be described in detail with reference to Fig. 14(A). In Fig. 14(A), a linear connection 1513 shows the result of connecting from the chest 1511 to the left hand 1512 of a person 1510 on the left side of the image. Furthermore, a connection 1523 shows the result of connecting from the chest 1521 to the right hand 1522 of a person 1520 on the right side of the image. The continuity of the line at connection 1523 of person 1520 is reduced due to occlusion occurring midway due to the left hand of person 1510. A lower value is calculated as the reliability of the connection at connection 1523 compared to connection 1513 where no occlusion has occurred.

[0071] In S1402 of Fig. 13, the priority region determination unit 603 determines a priority region from among the multiple regions detected in S1202. For example, assume that a face is to be determined as the priority region. In Fig. 14, two people 1510 and 1520 are present in the image, so the faces of each person are determined as the two priority regions.

[0072] In S1403, the priority region expansion unit 604 performs region expansion processing from the priority region determined in S1402 in the direction connected in S1203. In the example of Fig. 14(B), expansion regions 1514 and 1524 are acquired from the face detection regions of persons 1510 and 1520, respectively.

[0073] In S1404, the priority portion expansion unit 604 performs further expansion processing of the region from the region expanded in S1403 in the direction connected in S1203. In the example of Fig. 14(B), expansion processing from region 1514 to region 1515 and expansion processing from region 1524 to region 1525 are performed.

[0074] In S1405, a determination is made as to whether the region expansion processes of S1403 and S1404 have been performed for all parts that do not have region information. If it is determined that the region expansion processes have been completed, the process proceeds to S1406. If it is determined that the region expansion processes have not been completed, the process returns to S1403 and the region expansion processes are repeated.

[0075] In S1406, the extension range adjustment unit 605 performs processing to adjust the extension range of the region extended in S1403 and S1404. In S1407, the priority portion determination unit 603 determines the priority portion of the overlapping region. If it is determined that overlap occurs in the regions extended in S1403 and S1404, the priority portion determination unit 603 uses the reliability of the connection calculated in S1401 to determine the priority portion. In the example of FIG. 14(B), region 1515 and region 1525 overlap. Since the reliability of connection portion 1523 is lower than that of connection portion 1513, the overlapping portion is determined to be the region of person 1510. At this time, it is assumed that the user has specified the subject at coordinate 1501 in FIG. 14(B) by touch operation, eye-gaze input, or the like via operation unit 156. Since the position of the coordinates 1501 is within the overlapping area of ​​the area 1515 and the area 1525, it is determined that the user has designated the person 1510. After S1407, the series of processes ends.

[0076] According to this embodiment, in a situation where multiple subjects are detected overlapping in an image (occlusion), even if the user specifies a location that does not have area information, such as a joint, it is possible to specify a subject with a higher priority.

[0077] In conventional technology, if a user specifies a non-detected portion between detected subject parts, there is a possibility that the subject frame will not be displayed. In contrast, in the above embodiment, it is possible to treat a portion that cannot be detected by the imaging device as an extended region of the detected region. If a user erroneously specifies an area that is not a detection region, the imaging device can determine that a detection region has been specified. For example, it is possible to control the display of a subject frame (such as a focus detection frame or tracking frame) that corresponds to an area including the extended region.

[0078] The technical scope of the present invention is not limited to the above-described embodiments. There are also embodiments in which the priority part is changed depending on the position designated by a user operation or the size of the detected part. For example, assume that the face and torso of a subject are detected at a predetermined size or larger. If the size of the detected part is smaller than a predetermined size (threshold), the face is determined as the priority part rather than the torso. Furthermore, if the size of the detected part is equal to or larger than a predetermined size (threshold), the face or torso is determined as the priority part depending on the position designated by a user operation. Although the preferred embodiments of the present invention have been described above, the present invention is not limited to the above-described embodiments, and various modifications and changes are possible within the scope of the gist of the present invention.

[0079] [Other embodiments] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions. [Explanation of symbols]

[0080] 100 Imaging device 141 Image sensor 143 Imaging control unit 151 CPU 152 Image processing section 161 Subject detection unit

Claims

1. An image processing device that acquires an image and detects a subject, a detection means for detecting a face and a body part of the same subject in the image; and a determination means for determining that the face has been designated when the user designates the area between the face and the torso.

1. An image processing device comprising:

2. When the determining means determines that the face has been designated, the determining means performs focus detection processing on the area corresponding to the face.

2. The image processing device according to claim 1, wherein:

3. The area between the face and the torso is an area not detected by the detection means.

3. The image processing device according to claim 1, wherein the image processing device is a computer.

4. When the determining means determines that the face has been designated, the determining means performs AE processing or image processing based on the area corresponding to the face.

4. The image processing device according to claim 1, wherein the image processing device is a computer.

5. an expanding means for expanding the area corresponding to the face detected by the detecting means in a direction toward the area corresponding to the trunk detected by the detecting means; The determining means determines that the face has been specified when the area expanded by the expanding means is specified.

5. The image processing device according to claim 1, wherein the image processing device is a computer.

6. the determining means determines that the face has been designated when a user designates a predetermined range of area surrounding the face area detected by the detecting means; The predetermined range is wider in a direction toward the body than in other directions.

6. The image processing device according to claim 1, wherein the image processing device is a computer.

7. An image processing device that acquires an image and detects a subject, a detection means for detecting a face and a body part of the same subject in the image; a display control means for displaying a frame in an area corresponding to the face displayed on the display means when the user specifies the area between the face and the torso; 1. An image processing device comprising:

8. and a determining means for determining, when a user specifies an area between the face and the body, that the face has been specified and performing focus detection processing on the area corresponding to the face.

8. The image processing device according to claim 7,

9. The area between the face and the torso is an area not detected by the detection means.

9. The image processing device according to claim 7, wherein the image processing device is a computer.

10. When the determining means determines that the face has been designated, the determining means performs AE processing or image processing based on the area corresponding to the face.

9. The image processing device according to claim 8,

11. an expanding means for expanding the area corresponding to the face detected by the detecting means in a direction toward the area corresponding to the trunk detected by the detecting means; The determining means determines that the face has been specified when the area expanded by the expanding means is specified.

9. The image processing device according to claim 8,

12. the determining means determines that the face has been designated when a user designates a predetermined range of area surrounding the face area detected by the detecting means; The predetermined range is wider in a direction toward the body than in other directions.

12. The image processing device according to claim 11.

13. An image processing device that acquires an image and detects a subject, a detection means for detecting joints relating to a plurality of subjects in the image; a connecting means for connecting the joints detected by the detecting means so that the joints are related to the same subject; and a determining means for determining, when the areas corresponding to the subjects connected by the connecting means overlap, that the overlapping area is the area of ​​the subject with higher reliability.

1. An image processing device comprising:

14. The determining means determines the reliability based on the linearity of the connection when the joints of the subjects are connected by the connecting means in the overlapping region or the variance of the depth values ​​of the joints.

14. The image processing device according to claim 13.

15. an imaging means for capturing an image; An image processing device according to any one of claims 1 to 14; a display means for displaying the image captured by the imaging means; An imaging device characterized by:

16. A control method executed in an image processing device that acquires an image and detects a subject, comprising: a detection step of detecting a face and a body part of the same subject in the image; a determining step of determining that the face has been designated when the user designates a region between the face and the torso. A control method comprising:

17. A control method executed in an image processing device that acquires an image and detects a subject, comprising: a detection step of detecting a face and a body part of the same subject in the image; a display control step of displaying a frame in an area corresponding to the face displayed on the display means when the user specifies the area between the face and the torso. A control method comprising:

18. A control method executed in an image processing device that acquires an image and detects a subject, comprising: a detection step of detecting joints related to a plurality of subjects in the image; a connecting step of connecting the joints detected in the detecting step so that the joints are related to the same subject; and a determining step of determining, when the areas corresponding to the subjects connected in the connecting step overlap, that the overlapping area is the area of ​​the subject with higher reliability. A control method comprising:

19. A program that causes a computer of an image processing apparatus to execute each step according to any one of claims 16 to 18.

Citation Information

Patent Citations

  • A library seat management system based on the visual Internet of Things

    CN109902628A

  • Image pickup device, image pickup method, and program

    JP2011045014A

  • Image processing device and image processing method

    JP2012015889A

  • Imaging apparatus and image reproduction apparatus

    JP2013012982A

  • Tracking controller

    JP2019201387A