Image processing device, control method for image processing device, and program
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- CANON KK
- Filing Date
- 2024-07-16
- Publication Date
- 2026-08-03
AI Technical Summary
【0009】 本発明によれば、被写体の姿勢を推定する際に、その推定結果が誤推定となるのを低減することができる。
Smart Images

Figure 0007899261000001 
Figure 0007899261000002 
Figure 0007899261000003
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing apparatus, a control method for an image processing apparatus, and a program.
Background Art
[0002] In the field of computer vision, there is an object detection technique for detecting an object in an image and displaying a detection frame surrounding the object as a detection result. As an application example of the object detection technique, for example, there is a technique for detecting key points (feature points) such as joints of a person existing in an image and estimating the posture of the person based on the detection result. This posture estimation technique includes a top-down type (top-down method) and a bottom-up type (bottom-up method). In the top-down type, a person in an image is detected, and the posture of the person is estimated based on the detection of key points defined in advance for the person. In the bottom-up type, a plurality of key points in an image are detected, and the key points are connected by a straight line, that is, joined together, to estimate the posture of the person. The top-down type has higher accuracy in posture estimation than the bottom-up type, but the calculation cost at the time of posture estimation tends to be higher than that of the bottom-up type. On the other hand, the bottom-up type has a smaller amount of calculation at the time of posture estimation than the top-down type, but false detection of key points and false connection between key points tend to occur. Also, in the bottom-up type, false detection of key points and false connection between key points may occur, and the estimated posture of the person may be within a range where it is not unnatural, that is, a range where the posture can occur. In this case, the estimated posture of the person becomes an incorrect estimation result. Non-Patent Document 1 discloses a bottom-up type posture estimation technique. Patent Document 1 also discloses an apparatus that performs a detection process for detecting a plurality of types of body parts of a subject in an image and selects any one of a plurality of determination methods for determining the action of the subject based on the result of the detection process. In the apparatus described in Patent Document 1, when selecting any one of the plurality of determination methods, the positional relationship of two or more types of body parts among the plurality of types of body parts is used to determine the action of the subject. Then, the action of the subject is determined according to the selected determination method. [Prior art documents] [Patent Documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2023-68992 [Non-patent literature]
[0004] [Non-Patent Document 1] Alejandro NewelL, Zhiao Huang, Jia Deng, 2017, “Associative Embedding: End-to-End Learning for Joint Detection and Grouping”, ADVANCES IN NEURAL INFORMATION PROCESSING SYSTEMS 30, 2278-2288 [Overview of the project] [Problems that the invention aims to solve]
[0005] However, in the technology described in Non-Patent Document 1, namely the bottom-up pose estimation technology, as mentioned above, if misdetection of keypoints or misconnection of keypoints occurs, the estimated pose of the person will be an incorrect estimation result.
[0006] The present invention has been made in view of the above-mentioned problems. The object of the present invention is to provide an image processing apparatus, a control method for the image processing apparatus, and a program that can reduce the occurrence of erroneous estimations when estimating the posture of a subject. [Means for solving the problem]
[0007] To achieve the above objective, the image processing apparatus of the present invention includes: detection means for detecting a plurality of key points for each of a plurality of subjects contained in an image; extraction means for extracting body parts of the plurality of subjects as a detection range for the subjects in a manner different from that of the detection means; first determination means for determining whether the subject from which the key points were detected and the subject from which the detection range was extracted are the same subject; and second determination means for determining the posture of a corresponding subject from the plurality of key points detected by the detection means based on the determination result of the first determination means. A reliability determination means that, depending on the positional relationship between the key point detected by the detection means and determined to be the same subject, and the detection range determined to be the same as the key point, reduces the reliability of at least one of the detection range and the key point, or reduces the reliability of the subject posture determination using the key point determined to be the same; It is characterized by having the following features.
[0008] Furthermore, the image processing apparatus of the present invention includes: detection means for detecting a plurality of key points on a subject included in an image; connecting means for connecting the key points detected by the detection means with straight lines; extraction means for extracting a detection range that surrounds a part of the body of a subject included in the image that is subject to determination for which posture or movement is to be determined, and indicates the detection range of the subject; and determination means for determining whether the subject from which the key points were detected and the subject from which the detection range was extracted are the same subject. A reliability determination means that, depending on the positional relationship between the key point detected by the detection means and determined to be the same subject, and the detection range determined to be the same as the key point, reduces the reliability of at least one of the detection range and the key point, or reduces the reliability of the subject posture determination using the key point determined to be the same; The connection means is characterized in that, if the determination means determines that the subject is the same, the connection means determines, at least, the key point to be connected from among the key points detected by the detection means, according to the positional relationship between the key point detected by the detection means and the detection range extracted by the extraction means. [Effects of the Invention]
[0009] According to the present invention, when estimating the posture of a subject, it is possible to reduce the likelihood of the estimation result being incorrect. [Brief explanation of the drawing]
[0010] [Figure 1] This block diagram shows an example of the hardware configuration of an imaging device according to the first embodiment. [Figure 2] This is a flowchart illustrating the posture estimation process using a bottom-up approach. [Figure 3] This is a flowchart illustrating the posture estimation process using a top-down approach. [Figure 4] This is a flowchart showing the attitude estimation process performed by the imaging device according to the first embodiment. [Figure 5] This is an image in which keypoints and detection frames are superimposed on the input image that has been input from the imaging control unit to the object detection unit. [Figure 6] This is an image in which keypoints and detection frames are superimposed on the input image that has been input from the imaging control unit to the object detection unit. [Figure 7] This is a flowchart showing the attitude estimation process performed by the imaging device according to the second embodiment. [Figure 8] This is an image in which keypoints and detection frames are superimposed on the input image that has been input to the object detection unit from the imaging control unit of the imaging apparatus according to the third embodiment. [Figure 9] This is an image in which keypoints and detection frames are superimposed on the input image that has been input to the object detection unit from the imaging control unit of the imaging apparatus according to the third embodiment. [Figure 10] This is a flowchart showing the attitude estimation process performed by the imaging device according to the fifth embodiment. [Figure 11] This is an image in which keypoints and detection frames are superimposed on the input image that has been input to the object detection unit from the imaging control unit of the imaging apparatus according to the fifth embodiment. [Figure 12] This is an image in which keypoints and detection frames are superimposed on the input image that has been input to the object detection unit from the imaging control unit of the imaging apparatus according to the fifth embodiment. [Figure 13] This is an image in which keypoints and detection frames are superimposed on the input image that has been input to the object detection unit from the imaging control unit of the imaging apparatus according to the fifth embodiment. [Figure 14] This is an image in which keypoints and detection frames are superimposed on the input image that has been input to the object detection unit from the imaging control unit of the imaging apparatus according to the fifth embodiment. [Modes for carrying out the invention]
[0011] Hereinafter, each embodiment of the present invention will be described in detail while referring to the drawings. However, the configurations described in the following embodiments are merely examples, and the scope of the present invention is not limited by the configurations described in each embodiment. For example, each part constituting the present invention can be replaced with any configuration that can exhibit the same function. Also, any component may be added. Further, any two or more configurations (features) among the embodiments can be combined.
[0012] <First Embodiment> Hereinafter, the first embodiment will be described with reference to FIGS. 1 to 6. FIG. 1 is a block diagram showing an example of the hardware configuration of an imaging device according to the first embodiment. The imaging device 100 shown in FIG. 1 is, in this embodiment, a digital still camera, a video camera, or the like to which an image processing device is applied, but is not limited thereto. The imaging device 100 includes a lens unit 101, an aperture control unit 105, a zoom control unit 113, a focus control unit 133, an imaging element 141, an imaging signal processing unit 142, and an imaging control unit 143. The imaging device 100 also includes a monitor display 150, a CPU (Central Processing Unit) 151, an image processing unit 152, an image compression / decompression unit 153, a RAM (Random Access Memory) 154, and a flash memory 155. The imaging device 100 further includes an operation switch 156, an image recording medium 157, a power management unit 158, a battery 159, a position and orientation change acquisition unit 161, an object detection unit 162, and a defocus calculation unit 163. These hardware components of the imaging device 100 are connected to each other via a bus 160 so as to be communicable. The lens unit 101 includes a fixed single-group lens 102, an aperture 103, an aperture motor (AM) 104, a zoom lens 111, a zoom motor (ZM) 112, a fixed three-group lens 121, a focus lens 131, and a focus motor (FM) 132.
[0013] The CPU 151 is a computer that controls the operation of each hardware component. The aperture control unit 105 drives the aperture 103 via the aperture motor 104. Thereby, the aperture diameter of the aperture 103 can be adjusted to control the amount of light during shooting. The zoom control unit 113 drives the zoom lens 111 via the zoom motor 112. Thereby, the focal length can be changed. The focus control unit 133 determines the driving amount for driving the focus motor 132 based on the amount of deviation (defocus amount) in the focus direction of the lens unit 101. Also, the focus control unit 133 drives the focus lens 131 via the focus motor 132. Thereby, the focus adjustment state can be controlled. By moving the focus lens 131, AF (autofocus) control becomes possible. Note that the focus lens 131 is a focus adjustment lens and is configured as a single lens in FIG. 1, but is usually composed of a plurality of lenses. An object image is formed on the imaging element 141 via the lens unit 101. This object image is converted into an electrical signal by the imaging element 141. The imaging element 141 is a photoelectric conversion element. The imaging element 141 has light receiving elements arranged in m (where "m" is an integer) pixels in the horizontal direction and n (where "n" is an integer) pixels in the vertical direction. The image formed and photoelectrically converted on the imaging element 141 is arranged as an image signal (image data) by the imaging signal processing unit 142. Thereby, an image is acquired on the imaging surface of the imaging element 141. Thus, in this embodiment, the imaging element 141 and the like constitute imaging means for shooting an object and acquiring an image.
[0014] Image data is output from the imaging signal processing unit 142. The image data is transmitted to the imaging control unit 143 and temporarily stored in the RAM 154. The image data stored in the RAM 154 is compressed by the image compression / decompression unit 153 and then recorded on the image recording medium 157. In parallel with this recording, the image data stored in the RAM 154 is transmitted to the image processing unit 152. The image processing unit 152 processes the image signal and performs operations such as reducing or enlarging the size of the image data and calculating the similarity between image data. The image data processed to the optimal size by the image processing unit 152 is displayed as an image on the monitor display 150. The monitor display 150 can also display a preview image and a through image, and can superimpose the object detection results of the object detection unit 162 onto the image data. In addition, the imaging device 100 can use the RAM 154 as a ring buffer. This allows for buffering, for example, multiple image data captured within a predetermined period, the detection results from the object detection unit 162 corresponding to each image data, and the position and orientation changes of the imaging device 100 acquired by the position and orientation change acquisition unit 161.
[0015] The operation switch 156 is an input interface including, for example, a touch panel or buttons. This allows for operations such as selecting various function icons displayed on the monitor display 150. The CPU 151 can determine the storage time of the image sensor 141 based on instructions from the operator input via the operation switch 156 and the magnitude of the pixel signals of image data temporarily stored in the RAM 154. The CPU 151 can also determine the gain setting value when outputting from the image sensor 141 to the image signal processing unit 142. The image processing control unit 143 receives instructions from the CPU 151 regarding the storage time and gain setting value and controls the image sensor 141. The object detection unit 162 uses the image signal to determine the region in the image where a predetermined subject exists. This region may be output as rectangular information, or as a subject region map where the pixel values represent the "likelihood of the subject existing". The focus control unit 133 can perform AF control for a specific subject region. The aperture control unit 105 can perform exposure control using the brightness value of a specific subject region. The image processing unit 152 can perform gamma correction, white balance processing, and other operations based on the subject area.
[0016] The battery 159 is managed by the power management unit 158 and supplies power to each piece of hardware of the imaging device 100. The flash memory 155 stores control programs necessary for the operation of the imaging device 100, as well as parameters used for the operation of each part. The control programs also include programs that cause the computer to execute each piece of hardware of the imaging device 100, i.e., each part and each means (control method for the image processing device). When the imaging device 100 is started by user operation, that is, when it transitions from a power-off state to a power-on state, the control programs and parameters stored in the flash memory 155 are loaded into a part of the RAM 154. The CPU 151 controls the operation of each piece of hardware according to the control programs and parameters loaded into the RAM 154. The position and attitude change acquisition unit 161 is composed of position and attitude detection sensors such as a gyroscope, accelerometer, and electronic compass. The position and attitude change acquisition unit 161 measures the position and attitude changes of the imaging device 100 relative to the shooting scene. The position and orientation change information measured by the position and orientation change acquisition unit 161 is stored in the RAM 154. The defocus calculation unit 163 calculates the amount of defocus for any region in the image. The amount of defocus may be output as a single point, or it may be calculated at equal intervals across the entire image and output as a defocus map. The amount of defocus is then stored in the RAM 154 and can be accessed by the image processing unit 152.
[0017] In this embodiment, an image acquired by the imaging means is input from the imaging control unit 143 to the object detection unit 162. The object detection unit 162, under control from the CPU 151, detects a subject in this input image and estimates the pose of the subject. In this embodiment, a bottom-up method is used for pose estimation. In the bottom-up method, if the input image contains multiple subjects (people), the pose estimation of these subjects is performed simultaneously. Figure 2 is a flowchart of the pose estimation process when using the bottom-up method. As shown in Figure 2, in step S201, the object detection unit 162 simultaneously estimates the pose of all subjects included in the input image. In this embodiment, a neural network is used for pose estimation. The neural network simultaneously determines the position of the subject's keypoints and the output result of tags used to determine if they are the same person.
[0018] Furthermore, there is a top-down method for pose estimation. In the top-down method, if the input image contains multiple subjects (people), the region of each subject is detected, and then the pose of each subject is estimated. Figure 3 is a flowchart of the pose estimation process when using the top-down method. As shown in Figure 3, in step S301, the object detection unit 162 detects the region of subjects contained in the input image. In step S302, the object detection unit 162 estimates the pose of the subjects within the region detected in step S301. In step S303, the object detection unit 162 determines whether the estimation in step S302 has been completed for the number of people detected in step S301. If the determination in step S303 is that the estimation has been completed, the process ends. On the other hand, if the determination in step S303 is that the estimation has not been completed, the process returns to step S302, and the subsequent steps are executed in order.
[0019] Figure 4 is a flowchart showing the posture estimation process performed by the imaging device according to the first embodiment. The detailed processing in step S201 of the flowchart shown in Figure 2 is shown in the flowcharts of steps S401 and S402 in Figure 4. As shown in Figure 4, in step S401, the object detection unit 162 uses, for example, a neural network to detect multiple keypoints and tags for each keypoint in the subject included in the input image from the imaging control unit 143 (detection step). In this embodiment, the subject is a person (human), but it is not limited to this, and may be other animals, for example. If the subject is a person, at least one part from among the following is detected (extracted) as a keypoint: pupil, ear, top of head, neck, shoulder, elbow, wrist, waist, knee, and ankle. These keypoints are feature points that can contribute to estimating the posture of a person. Each keypoint includes location information indicating which part of the person it is and a likelihood indicating the accuracy of the location information. The tags are used to determine the same person for each keypoint. Each tag includes a classification (person identification information) indicating which person each key point belongs to, and a likelihood indicating the accuracy of that classification. Thus, in this embodiment, the object detection unit 162 also functions as a detection means for detecting key points and tags.
[0020] In step S402, the object detection unit 162 connects the keypoints detected in step S401 with straight lines (connection step). Thus, in this embodiment, the object detection unit 162 also functions as a connection means for connecting keypoints with straight lines. Then, the object detection unit 162 estimates the posture of a person based on the connection result of connecting the keypoints with straight lines. Thus, in this embodiment, the object detection unit 162 also functions as an estimation means for estimating the posture of a person to be judged. In step S402, the connection is usually determined according to the position between the keypoints and the likelihood of the tags. In this implementation, the connection was made using tag information, but the connection may also be determined by segmentation of the subject. Note that if the likelihood of classification by tags is above a predetermined threshold and the maximum value is used, the connection between the keypoints is determined at the time of keypoint detection in step S401, so step S402 can be effectively omitted. In such cases, steps S401 and S402 in Figure 4 may be represented as a single step, as in step S201 in Figure 2.
[0021] Step S403 is executed in parallel with steps S401 and S402. In step S403, the object detection unit 162 extracts a detection frame for a person included in the input image from the imaging control unit 143, for example, using a neural network (extraction step). The detection frame indicates the detection range for a person by enclosing the body parts of all persons whose posture is to be determined in the input image with rectangles. Preferably, the range enclosed by this detection frame is such that the body parts can be sufficiently identified. For example, it is preferable that the detection frame encloses the whole body, the entire upper body, the entire head, etc., of the person whose posture is to be determined. In addition, in step S403, along with the extraction of the detection frame, the center of the detection frame and the vertical and horizontal lengths of the detection frame are also extracted. Thus, in this embodiment, the object detection unit 162 also functions as an extraction means for extracting detection frames. Furthermore, for example, if there are detection frames for the face, head, upper body, and whole body, the determination of whether these detection frames belong to the same person is made based on the degree of overlap between the detection frames and the distance between them.
[0022] In step S404, the object detection unit 162 determines, for example, that the person is running, based on the person's posture estimation result in step S402. In this embodiment, the positional relationship between the keypoints connected in step S402 (keypoint group) and the detection frame extracted in step S403 is compared to determine the reliability of the posture.
[0023] In step S405, the object detection unit 162 issues instructions to the CPU 151, for example, to switch the focus range of the imaging device 100, in accordance with the determination result of step S404.
[0024] Step S406 is executed. In step S406, the object detection unit 162 determines whether the person whose keypoints were detected in step S401 and the person whose detection frame was extracted in step S403 are the same person (determination step). This determination is made, for example, based on the degree of overlap between the region formed by multiple keypoints and the detection frame, or the distance between the keypoints and the detection frame. As a method of comparison, the smallest circle or rectangle containing the keypoints corresponding to the top of the head and neck is formed and compared with the detection frame of the head, or a rectangle connecting the keypoints corresponding to the left and right shoulders and waist is compared with the upper body frame. Alternatively, the distance between the keypoints and the detection frame may be used for comparison. If it is determined that they are the same person, the unit checks whether there is any inconsistency between the keypoints determined to be the same and the detection frame, as shown below.
[0025] Figure 5 shows an image in which keypoints and detection frames are superimposed on the input image received from the imaging control unit to the object detection unit. In image 500 shown in Figure 5, the solid line 501 outside it represents the range of the image 500. This image 500 contains person A and person B. Person A is located in the lower left of image 500, and only the upper part from the chest up is visible. Person B is located in the center of image 500, and the entire body is visible. The detection frame 502 is the detection result of person A's head, and is indicated by a dotted line surrounding person A's head. Keypoints detected include keypoint KP1 corresponding to the top of the head, keypoint KP2 corresponding to the neck, keypoint KP3 corresponding to the left shoulder, and keypoint KP4 corresponding to the right shoulder. In addition, keypoint KP5 corresponding to the left hip, keypoint KP6 corresponding to the left knee, and keypoint KP7 corresponding to the left ankle are also detected. Furthermore, keypoint KP8 corresponding to the right hip, keypoint KP9 corresponding to the right knee, and keypoint KP10 corresponding to the right ankle are also detected. Of keypoints KP1 to KP10, keypoints KP1 to KP4 are keypoints of person A. Also, keypoints KP5 to KP10 are originally keypoints of person B, but are detected as keypoints of the hip, knee, and foot of person A, which are partially obscured in the image. In this way, multiple keypoints are determined to belong to the same subject by the object detection unit 162 (keypoint identity determination means) (keypoint identity determination process).
[0026] Here, "identity" refers to the same person. Keypoints KP1 to KP10 form a keypoint group KPG1 connected by a dashed line. Therefore, the object detection unit 162 detects that person A is in a fallen position based on the keypoint group KPG1. Note that keypoints are not limited to keypoints KP1 to KP10; other keypoints may also be detected.
[0027] For example, suppose the key point KP1 at the top of the head is at ±30 degrees from the vertical axis in image 500 relative to the key point KP2 at the neck, and the center of the head detection frame 502 is within two dimensions of the size of the detection frame 502 from the bottom edge of image 500. In this case, it is highly likely that person A is not bending their body or neck, their upper body is cut off, and the key points at the waist, knees, and ankles are not visible. When there is such an inconsistency in the relationship between the detection frame 502 and the key point detection positions (key points KP5 to 10), the object detection unit 162 (reliability determination means) lowers the confidence level of the posture and makes a determination (reliability determination process). When lowering the confidence level of the posture, the following three processes are appropriately selected.
[0028] The first process involves reducing the confidence level of key points KP5 to KP10. A specific example would be to exclude some key points included in the key point group KPG1, namely key points KP5 to KP10 (waist, knee, and ankle), or to reduce their detected likelihood before processing them in the next step.
[0029] The second process involves reducing the confidence level of detection frame 502. Specifically, for example, detection frame 502 is not used in subsequent processes.
[0030] The third process is to reduce the confidence level of the posture estimation based on the keypoint group KPG1. Specifically, for example, if the posture is initially estimated to be lying down based on the keypoint group KPG1, but the positional relationship between the detection frame 502 and the keypoint KP1 at the top of the head suggests an upright posture, then the estimation result that the person is lying down is incorrect is considered incorrect. By adding this confidence reduction process (hereinafter referred to as the "confidence reduction process"), it is possible to reduce the number of cases where the judgment in step S404 is incorrect, that is, where person A is judged to be in a lying-down posture. It is also possible to prevent the processing in step S405 from being incorrect based on the judgment result of step S404.
[0031] Figure 6 shows an image in which keypoints and detection frames are superimposed on the input image received from the imaging control unit to the object detection unit. In Figure 6, keypoints for the shoulders and waist have been reduced to reduce computational load. In image 600 shown in Figure 6, the solid line 601 outside it represents the range of image 600. This image 600 contains person A and person B. Person A is located in the lower left of image 600, and the upper part from the waist up is visible. Person B is located slightly to the right of the center of image 600, and the entire body is visible. The detection frame 602 is the detection result for the entire upper body of person A, and is shown by a dotted line enclosing the entire upper body of person A. Keypoints detected include keypoint KP1 corresponding to the top of the head and keypoint KP2 corresponding to the neck. In addition, keypoints KP21 corresponding to the left elbow, keypoint KP22 corresponding to the left wrist, keypoint KP23 corresponding to the right elbow, and keypoint KP24 corresponding to the right wrist have also been detected. Furthermore, keypoints KP6 corresponding to the left knee, KP7 corresponding to the left ankle, KP9 corresponding to the right knee, and KP10 corresponding to the right ankle have also been detected. Of these keypoints, keypoints KP1, KP2, and KP21-KP24 are keypoints of person A. In addition, keypoints KP6, KP7, KP9, and KP10 are actually located at the position of person B, but are detected as keypoints of the knees and feet of person A, which are partially obscured. As a result, these keypoints form a keypoint group KPG2 connected by dashed lines. Therefore, the object detection unit 162 detects that person A is lying down or in a sleeping position based on the keypoint group KPG2.
[0032] For example, suppose the keypoint KP1 at the top of the head is at ±30 degrees from the vertical axis in image 500 relative to the keypoint KP2 at the neck, and the vertical length of the upper body detection frame 602 is more than twice its horizontal length (aspect ratio of 2). In this case, it is highly likely that person A is actually in a standing position, and the positions of the keypoints KP6, KP, KP9, and KP10 at the knees and ankles are not higher than the keypoints KP21 and KP23 at the elbows. Note that in the case of a lying position, the ratio of the vertical length to the horizontal length of the upper body detection frame 602 will be smaller than the ratio of the vertical length to the horizontal length of the detection frame 602 in the standing position described above. Also, the positional relationship between the keypoint KP1 at the top of the head and the keypoint KP2 at the neck will differ between a lying position and a standing position. If there is an inconsistency in the relationship between the detection frame 602 and the keypoint detection positions (keypoints KP6, KP, KP9, KP10), the judgment process in step S404 and the confidence reduction process in step S406 are executed. The confidence reduction process reduces the likelihood of a false judgment in step S404, i.e., the judgment that person A is lying down or in a sleeping position. This allows, for example, the judgment that person A is standing. In addition, although the number of keypoints detected used for posture determination is 10 in both Figures 5 and 6, it is not limited to this, and can be applied to posture determination even if there are fewer than 10 or more than 10 keypoints.
[0033] As described above, in this embodiment, a confidence reduction process can be added based on information such as the position and likelihood of keypoints, and the position, size, and likelihood of the detection frame. This prevents errors in the processing in step S404 and step S405. In this embodiment, the orientation of the head was not considered, based on the position of the person's eyes, ears, etc., but this is not limited to this. For example, considering the orientation of the head may make it easier to determine whether the person is in an impossible posture. Also, depending on the situation, such as a specific event or sport, the normal posture may be restricted, so it may be possible to determine whether the posture is impossible. Furthermore, if there is an inconsistency between the detection frame and the keypoint detection results, the confidence of the detection frame may be reduced. Whether the keypoint or the detection frame is more reliable depends on the likelihood of each detection result, the performance of the detector, the scene, etc.
[0034] <Second Embodiment> The second embodiment will now be described with reference to Figure 7, focusing on the differences from the previously described embodiment, and omitting explanations of similar matters. Figure 7 is a flowchart showing the posture estimation process performed by the imaging device according to the second embodiment. The target of the posture estimation process is, for example, the image 500 shown in Figure 5. As mentioned above, for example, suppose the key point KP1 at the top of the head is at ±30 degrees from the vertical axis in the image 500 relative to the key point KP2 at the neck, and the center of the head detection frame 502 is within two of the size of the detection frame 502 from the bottom edge of the image 500. In this case, it is highly likely that the upper body of person A is cut off, and the key points of the waist, knees, and ankles are not captured. Therefore, in this embodiment, as shown in Figure 7, in step S704 before step S705, the reliability of the key points of person A's waist, knees, and feet is reduced for the determination.
[0035] In the flowchart shown in Figure 7, steps S701, S702, and 703 are executed. Steps S701, S702, and 703 are the same as steps S401, S402, and 403 in the flowchart shown in Figure 4, respectively. After steps S702 and 703 are executed, the process proceeds to step S704. In step S704, the object detection unit 162 determines the posture by excluding key points KP5 to 10 of the waist, knees, and feet. This prevents the determination result in step S705 and the processing result in step S706 from being incorrect. Steps S705 and S706 are the same as steps S404 and S405 in the flowchart shown in Figure 4. In this embodiment, posture determination was performed by excluding some key points, but this is not limited to this. For example, posture can be determined by taking into account the reduction in likelihood when lowering the confidence level of the key points. This prevents the judgment result in step S705 and the processing result in step S706 from being incorrect. Furthermore, the same effect as with image 500 in Figure 5 can be obtained for image 600 shown in Figure 6.
[0036] <Third Embodiment> The third embodiment will now be described with reference to Figure 8, focusing on the differences from the previously described embodiment, and omitting explanations of similar matters. This embodiment is the same as the first embodiment, except for the example concerning the reliability of keypoints. Figures 8 and 9 are images in which keypoints and detection frames are superimposed on the input image input to the object detection unit from the imaging control unit of the imaging device according to the third embodiment. In image 800 shown in Figure 8, the solid line 801 on the outside represents the range of the image 800. This image 800 includes person A and person B. Person A is located on the left side of image 800. Person B is located on the right side of image 800. Person A is also located behind person B, and a part of their body, namely person A's left arm and person B's right arm in Figure 8, overlaps. In this overlapping area, the keypoint and posture determination are expected to be lower than usual (the detection result may also be lower, but this is a rule-based control that reduces the reliability). Therefore, it is preferable to either lower the reliability of the keypoint-based determination, as in the first embodiment, or lower the reliability of the keypoint, as in the second embodiment. This prevents the posture determination result and the processing based on the posture determination result from being incorrect.
[0037] Detection frame 802 represents the detection result for the entire upper body of person A, and is indicated by a dotted line enclosing the entire upper body of person A. Detection frame 812 represents the detection result for the entire upper body of person B, and is indicated by a dotted line enclosing the entire upper body of person B. In addition, key points KP1 corresponding to the top of the head, key point KP2 corresponding to the neck, and key point KP3 corresponding to the left shoulder have been detected. Key points KP21 corresponding to the left elbow, key point KP22 corresponding to the left wrist, key point KP23 corresponding to the right elbow, and key point KP24 corresponding to the right wrist have also been detected. Furthermore, key points KP31 corresponding to the left hip, key point KP32 corresponding to the left knee, and key point KP33 corresponding to the left ankle have also been detected. In addition, key points KP34 corresponding to the right hip, key point KP35 corresponding to the right knee, and key point KP36 corresponding to the right ankle have also been detected. These key points are connected by a dashed line to form key point group KPG3. In this embodiment, in step S404 of the flowchart shown in Figure 4, the object detection unit 162 compares the detection frame 802 surrounding the upper body of person A with the detection frame 812 adjacent to detection frame 802 and surrounding the upper body of person B. Then, it reduces the reliability of the detection frame with fewer keypoints.
[0038] Furthermore, in image 800' shown in Figure 9, even though the positional relationship between person A and person B is the same as in image 800 shown in Figure 8, some of person A's key points may not be detected. In such cases, in step S404 of the flowchart shown in Figure 4, the object detection unit 162 compares the detection frame 802 surrounding the upper body of person A with the detection frame 812 surrounding the upper body of person B, which is adjacent to person A and has been determined not to be person A. If the overlap exceeds a predetermined level, the reliability of the posture determination result is reduced. In this embodiment, the application example is shown with person A positioned behind person B, but for example, person A may be positioned in front of person B. Also, the amount of reduction in reliability may be changed (adjusted) depending on the distance and degree of overlap between detection frame 802 and detection frame 812. Furthermore, as shown in image 800' in Figure 9, the criterion for reducing the reliability of posture determination may be that the number of detected key points is less than a predetermined number.
[0039] <Fourth Embodiment> The fourth embodiment will be described below, focusing on the differences from the previously described embodiments, and similar matters will be omitted. In this embodiment, the examples regarding the reliability of key points are the same as in the third embodiment, and the flowchart is the same as in the second embodiment. In this embodiment, it is assumed that the detection frame 802 surrounding the upper body of person A and the detection frame 812 surrounding the upper body of person B overlap, similar to image 800 shown in Figure 8 and image 800' shown in Figure 9. When the degree of overlap between the detection frames is a predetermined degree, the reliability of key points KP3, KP21, and KP22 located on the side of detection frame 812 adjacent to detection frame 802 among the key points included in the key point group is reduced. In this embodiment, in step S704 of the flowchart shown in Figure 7, the aforementioned key points with low reliability are excluded, and in step S705, the posture is determined. In this embodiment, the amount of confidence reduction may be changed or the key points that reduce confidence may be changed depending on the distance and degree of overlap between detection frame 802 and detection frame 812.
[0040] <Fifth Embodiment> The fifth embodiment will be described below with reference to Figures 10 to 14, focusing on the differences from the previously described embodiments, and similar matters will be omitted. In this embodiment, the connection of key points is determined according to the positional relationship between detection frames such as faces, heads, and upper bodies, which are determined to be the same subject, and key points such as the top of the head and joints. Figure 10 is a flowchart of the posture estimation process performed by the imaging device according to the fifth embodiment. In the flowchart shown in Figure 10, steps S1001 and S1003 are executed in parallel. Steps S1001 and S1003 are the same as steps S401 and S403 in the flowchart shown in Figure 4, respectively. After the execution of steps S1001 and S1003, the process proceeds to step S1002. In step S1002, the object detection unit 162 uses detection frames determined to be the same subject to connect key points. For example, the object detection unit 162 determines the key point connection range, that is, the key points to be connected, based on the position and shape of the detection frames for the head and upper body. One method for determining the connection range is to decide whether or not to connect based on the distance information between keypoints. Another method is to create a cost function based on the likelihood of the keypoints, the tag ID, and the likelihood information, and connect the combination that minimizes the cost function.
[0041] In this embodiment, the range of positions for the top of the head, neck, shoulders, waist, and knees, as well as the key points to be connected, are determined according to the position and shape of the detection frame surrounding the head and the detection frame surrounding the upper body. Then, as a cost function, the cost within the range is set to 0 and the cost outside the range is set to ∞ to prevent the connection of key points outside the range. This improves the accuracy of the connections between key points. Even if a key point is included within the range, the cost can be changed depending on the position of that key point. In this embodiment, although the key points of the top of the head and neck are within the range of the detection frame surrounding the head, taking into account the error in key point detection, the predetermined range can be set to, for example, 1.3 times the size of the detection frame from the center of the detection frame. If the key points of the top of the head and neck are outside the range, the connection of those key points is not performed. In addition, in this embodiment, the positions of the top of the head and neck can be predicted according to the position and shape of the detection frame surrounding the head and the detection frame surrounding the upper body. Therefore, the range may be further restricted or the cost function may be changed. Furthermore, in this embodiment, the range for key points of the shoulders, waist, and knees can also be defined according to the position and shape of the detection frame surrounding the head and the detection frame surrounding the upper body. In addition, the orientation of the body can be predicted based on the number and position of the detected pupils. This prediction result can be used to connect the key points.
[0042] After step S1002 is executed, the process proceeds to steps S1004 and S1005 in order. Steps S1004 and S1005 are the same as steps S404 and S405 in the flowchart shown in Figure 4, respectively.
[0043] Figures 11 to 14 are images in which keypoints and detection frames are superimposed on the input image input from the imaging control unit to the object detection unit of the imaging device according to the fifth embodiment. In image 1100 shown in Figure 11, the solid line 1101 on the outside represents the range of the image 1100. This image 1100 contains person A. Person A is located in the center of image 1100 and is shown in a standing position with their entire body visible. Image 1100 also contains detection frames 1102 and 1103. Detection frame 1102 is shown as a dotted line and surrounds the head of person A. Detection frame 1103 is shown as a dotted line and surrounds the upper body of person A. Also, as shown in Figure 11, when person A is in a straight posture, detection frame 1103 becomes a vertically elongated rectangle in image 1100. Furthermore, all the detected keypoints are keypoints of person A and are connected by dashed lines to form the keypoint group KPG3. In this case, the shoulders (keypoints KP3 and KP4) are located between the vicinity of detection frame 1102 and the center of the long side of detection frame 1103. The waist (keypoints KP31 and KP34) are located near the bottom edge of detection frame 1103. The knees (keypoints KP32 and KP35) are located near the bottom edge of detection frame 1103 or outside the bottom edge.
[0044] In image 1200 shown in Figure 12, the solid line 1201 outside the image represents the area of image 1200. This image 1200 contains person A. Person A is located at the bottom of image 1200 and is either lying down or asleep, with their entire body visible. Image 1200 also contains detection frames 1202 and 1203. Detection frame 1202 is shown by a dotted line and surrounds person A's head. Detection frame 1203 is shown by a dotted line and surrounds person A's upper body. Also, as shown in Figure 12, if person A is in an upright posture, detection frame 1203 will be a horizontally elongated rectangle in image 1200. In this case as well, similar to Figure 11, the shoulders are located between the vicinity of detection frame 1202 and the center of the longer side of detection frame 1203. The waist is located near the right edge of detection frame 1203. Furthermore, the knee is located near the right end or outside the right end of the detection frame 1203. In this embodiment, misconnections can be reduced by providing a range of key points and a range of connections between key points that can be assumed from the position of the detection frame. Note that the number of detection frames is not limited to two, but may be three or more, for example.
[0045] In image 1300 shown in Figure 13, the solid line 1301 outside the image represents the area of image 1300. Image 1300 contains person A and person B. Person A is located in the center of image 1300, standing upright, and their entire upper body is visible. Person B is behind person A, and only their head is visible. Image 1300 also contains detection frames 1302 and 1303. Detection frame 1302 is indicated by a dotted line and surrounds the head of person A. Detection frame 1303 is indicated by a dotted line and surrounds the upper body of person A. Keypoints detected include keypoint KP41 corresponding to the top of the head, keypoint KP2 corresponding to the neck, keypoint KP3 corresponding to the left shoulder, and keypoint KP4 corresponding to the right shoulder. In addition, keypoints KP21 corresponding to the left elbow, KP22 corresponding to the left wrist, KP23 corresponding to the right elbow, and KP24 corresponding to the right wrist have also been detected. Furthermore, keypoints KP31 corresponding to the left hip and KP34 corresponding to the right hip have also been detected. Of these keypoints, keypoint KP41 is a keypoint of person B, and the remaining keypoints are keypoints of person A. Keypoint KP41 is actually a keypoint of person B, but it has been detected as a keypoint of person A. As a result, these keypoints form a keypoint group KPG4 connected by a dashed line. By limiting the range of keypoint KP41a to the vicinity of the detection frame, it is possible to reduce the misconnection of keypoint KP41, which is actually a keypoint of person B, with a keypoint of person A, as shown in Figure 13.
[0046] In image 1400 shown in Figure 14, the solid line 1401 outside the image represents the area of image 1400. This image 1400 contains person A. Person A is located in the center of image 1400, running, and their entire body is visible in a posture with their upper body leaning forward. Image 1400 also contains detection frames 1402 and 1403. Detection frame 1402 is shown by a dotted line and surrounds person A's head. Detection frame 1403 is shown by a dotted line and surrounds person A's upper body. When person A's upper body is leaning forward, the difference (ratio) between the vertical length and horizontal length of detection frame 1403 is smaller compared to when person A is standing. In this state, within detection frame 1403, the waist (keypoints KP31 and KP34) is located diagonally opposite the position of detection frame 1402. Therefore, by defining a range for the waistline and reflecting the location of keypoints in the cost function, it is possible to reduce misconnections between keypoints.
[0047] While preferred embodiments of the present invention have been described above, the present invention is not limited to the embodiments described above, and various modifications and changes are possible within the scope of its gist. The present invention can also be realized by supplying a program that implements one or more of the functions of the embodiments described above to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. Furthermore, the present invention can also be realized by a circuit (e.g., ASIC) that implements one or more functions. In addition, the imaging device 100 is a device to which an image processing device is applied, and the pose estimation process is performed within the device, but is not limited to this. For example, the imaging device 100 may be connected to an information processing device (server) in a communicative manner, and the information processing device may be a device to which an image processing device is applied. In this case, the information processing device performs the pose estimation process. The information processing device is not particularly limited, and for example, a desktop or notebook personal computer, a tablet terminal, a smartphone, etc., can be used.
[0048] Each embodiment of the disclosure includes the following configurations, methods, and programs. (Configuration 1) A detection means for detecting multiple key points for each of the multiple subjects contained in the image, An extraction means for extracting body parts of the multiple subjects as the detection range of the subjects in a manner different from that of the detection means, A first determination means for determining whether the subject from which the keypoint was detected and the subject from which the detection range was extracted are the same subject, Based on the determination result of the first determination means, a second determination means determines the posture of the corresponding subject from a plurality of key points detected by the detection means, A reliability determination means that, depending on the positional relationship between the key point detected by the detection means and determined to be the same subject, and the detection range determined to be the same as the key point, reduces the reliability of at least one of the detection range and the key point, or reduces the reliability of the subject posture determination using the key point determined to be the same; An image processing apparatus characterized by comprising: (Configuration 2) The image processing apparatus according to Configuration 1, further comprising a control means for controlling the imaging of the imaging means for imaging the plurality of subjects according to the posture of the subjects determined by the second determination means. (Configuration 3) The reliability determination means is characterized in that, depending on the positional relationship between a key point determined to be the same subject and a detection range that is adjacent to and determined to be different from the detection range determined to be the same as the key point, the reliability of at least one of the detection range and the key point, or the reliability of the subject pose estimation using the key point determined to be the same. Configuration 1 or 2 The image processing device described above. (Configuration 4) The subject is a human being, If the detection range includes the upper body of the person, the reliability of at least one of the detection ranges determined to be the same subject is reduced, or the reliability of the determination is reduced, depending on the aspect ratio of the detection range and the positional relationship of the keypoint. One of configurations 1 through 3 The image processing device described above. (Configuration 5) The subject is a human being, If the detection range includes the human head, then, depending on the positional relationship between the detection range and the key point indicating the head, The connection between the aforementioned keypoints is determined. Characterized by One of configurations 1 through 4 The image processing device described above. (Configuration 6) When a plurality of detection ranges are extracted by the extraction means, and an overlapping portion is formed where the detection ranges partially overlap with detection ranges that are determined not to be the same subject, the reliability of the keypoints located within the overlapping portion is reduced. One of configurations 1 through 5 The image processing device described above. (Configuration 7) It includes identity determination means for determining whether multiple key points correspond to the same subject, The identity determination means is characterized by making a determination using the location and likelihood of the detected key point, tag information representing the classification of the subject, and the likelihood of the tag information. One of configurations 1 through 6 The image processing device described above. (Configuration 8) The first determination means is characterized in that it determines whether or not the subject is the same subject based on the degree of overlap or distance between the subject from which the key point was detected and the subject from which the detection range was extracted, or the degree of overlap or distance between the region formed by a plurality of key points determined to be the same key point and the detection range. One of configurations 1 through 7 The image processing device described above. (Configuration 9) The configuration is characterized by comprising a connecting means for connecting the key points detected by the detection means with a straight line. One of configurations 1 through 8 The image processing device described above. (Configuration 10) A detection means for detecting multiple key points in a subject included in an image, A connecting means for connecting the key points detected by the detection means with a straight line, Extraction means for extracting a detection range that surrounds the body parts of a subject that are subject to determination and whose posture or movement is to be determined, and which indicates the detection range of the subject, A determination means for determining whether the subject from which the keypoint was detected and the subject from which the detection range was extracted are the same subject, A reliability determination means that, depending on the positional relationship between the key point detected by the detection means and determined to be the same subject, and the detection range determined to be the same as the key point, reduces the reliability of at least one of the detection range and the key point, or reduces the reliability of the subject posture determination using the key point determined to be the same; Equipped with, The image processing apparatus is characterized in that, if the determination by the determination means determines that the subjects are the same, the connection means determines, at least, the key point to be connected from among the key points detected by the detection means, according to the positional relationship between the key point detected by the detection means and the detection range extracted by the extraction means. (Configuration 11) The connection means is characterized in that it determines the key point to be connected to according to the size of the detection range. Configuration 10 The image processing device described above. (Configuration 12) The determination means is characterized in that it determines whether or not the subject is the same subject based on the degree of overlap between the subject in which the key point was detected and the subject from which the detection range was extracted, or the distance between the subject in which the key point was detected and the subject from which the detection range was extracted. Configuration 10 or 11 The image processing device described above. (Configuration 13) The detection means is characterized in that it detects at least one of the subject's eyes, ears, top of head, neck, shoulders, elbows, wrists, waist, knees, and ankles as the key point. One of configurations 1 through 12 The image processing device described above. (Configuration 14) An imaging device characterized by comprising imaging means for photographing a subject and acquiring the image. One of configurations 1 through 13 The image processing device described above. (Configuration 15) A detection means for detecting multiple key points for each of the multiple subjects included in the image, An extraction means for extracting at least one part of the subject's body from the image as the detection range of the subject, in a manner different from that of the detection means, A determination means for determining the posture of a subject based on a plurality of key points detected by the detection means and the detection range of the subject extracted by the extraction means, A reliability determination means that, depending on the positional relationship between the key point detected by the detection means and determined to be the same subject, and the detection range determined to be the same as the key point, reduces the reliability of at least one of the detection range and the key point, or reduces the reliability of the subject posture determination using the key point determined to be the same; An image processing apparatus characterized by comprising: (Method 1) Computer A method for controlling an image processing device, A detection process that detects multiple keypoints for each of the multiple subjects contained in the image, An extraction step in which body parts of the multiple subjects are extracted as the detection range of the subjects in a manner different from the method in the detection step, A first determination step of determining whether the subject from which the keypoint was detected and the subject from which the detection range was extracted are the same subject, Based on the determination result of the first determination step, a second determination step is performed to determine the posture of the corresponding subject from a plurality of key points detected in the detection step, A reliability determination step which reduces the reliability of at least one of the detection range and the key point, or reduces the reliability of the posture determination of the subject using the key point determined to be the same, depending on the positional relationship between the key point detected in the detection step and the detection range determined to be the same as the key point, A control method for an image processing apparatus, characterized by having the following features. (Method 2) Computer A method for controlling an image processing device, A detection process that detects multiple key points in a subject contained in an image, A connection step involves connecting the key points detected in the detection step with a straight line, Extraction step of extracting a detection range that surrounds the body parts of the subject to be determined, which are included in the aforementioned image and whose posture or movement is to be determined, and which indicates the detection range of the subject, A determination step to determine whether the subject from which the keypoint was detected and the subject from which the detection range was extracted are the same subject, A reliability determination step which reduces the reliability of at least one of the detection range and the key point, or reduces the reliability of the posture determination of the subject using the key point determined to be the same, depending on the positional relationship between the key point detected in the detection step and the detection range determined to be the same as the key point, It has, A control method for an image processing apparatus, characterized in that, if the determination in the determination step determines that the subjects are the same, the connection step determines, at least, the key point to be connected from among the key points detected in the detection step, according to the positional relationship between the key point detected in the detection step and the detection range extracted in the extraction step. (Method 3) Computer A method for controlling an image processing device, A detection process that detects multiple keypoints for each of the multiple subjects contained in the image, An extraction step in which, in a method different from the method in the detection step, at least one part of the subject's body is extracted from the image as the detection range of the subject, A determination step is performed to determine the posture of a subject based on the multiple key points detected in the detection step and the detection range of the subject extracted in the extraction step. A reliability determination step which reduces the reliability of at least one of the detection range and the key point, or reduces the reliability of the posture determination of the subject using the key point determined to be the same, depending on the positional relationship between the key point detected in the detection step and the detection range determined to be the same as the key point, A control method for an image processing apparatus, characterized by having the following features. (Program 1) Configurations 1 to 15 A program to cause a computer to execute each of the means of the image processing apparatus described in any one of the following. [Explanation of symbols]
[0049] 100 Imaging device 141 Image sensor 162 Object detection unit 502 detection frame KP1-KP10 Key Points
Claims
1. A detection means for detecting multiple key points for each of the multiple subjects contained in an image, An extraction means for extracting body parts of the multiple subjects as the detection range of the subjects in a manner different from that of the detection means, A first determination means for determining whether the subject from which the keypoint was detected and the subject from which the detection range was extracted are the same subject, Based on the determination result of the first determination means, a second determination means determines the posture of the corresponding subject from a plurality of key points detected by the detection means, An image processing apparatus comprising: a reliability determination means that, depending on the positional relationship between the key point detected by the detection means and determined to be the same subject, and the detection range determined to be the same as the key point, reduces the reliability of at least one of the detection range and the key point, or reduces the reliability of the subject posture determination using the key point determined to be the same.
2. The image processing apparatus according to claim 1, further comprising a control means for controlling the imaging of the imaging means for imaging the plurality of subjects according to the posture of the subjects determined by the second determination means.
3. The image processing apparatus according to claim 1, wherein the reliability determination means reduces the reliability of at least one of the detection range and the key point, or reduces the reliability of the pose estimation of the subject using the key point determined to be the same, according to the positional relationship between the key point determined to be the same as the detection range and the detection range that is adjacent to but determined to be different from the same as the key point.
4. The subject in question is a human being. The image processing apparatus according to claim 1, characterized in that, if the detection range includes the upper body of the person, the reliability of at least one of the detection ranges determined to be the same subject is reduced, or the reliability of the determination is reduced, according to the aspect ratio of the detection range and the positional relationship of the keypoint.
5. The image processing apparatus according to claim 1, wherein the subject is a human being, and if the detection range includes the head of the human being, the connection between the key points is determined according to the positional relationship between the detection range and the key points indicating the head.
6. The image processing apparatus according to claim 1, characterized in that, when a plurality of detection ranges are extracted by the extraction means and an overlapping portion is formed where the detection ranges partially overlap with detection ranges that are determined not to be the same subject, the reliability of the keypoints located within the overlapping portion is reduced.
7. It includes identity determination means for determining whether multiple key points correspond to the same subject, The image processing apparatus according to claim 1, characterized in that the identity determination means determines identity using the location and likelihood of the detected key point, tag information representing the classification of the subject, and the likelihood of the tag information.
8. The image processing apparatus according to claim 1, characterized in that the first determination means determines whether or not the subject is the same subject based on the degree or distance of overlap between the subject from which the keypoint was detected and the subject from which the detection range was extracted, or the degree or distance of overlap between the region formed by a plurality of keypoints determined to be the same keypoint and the detection range.
9. The image processing apparatus according to claim 1, further comprising a connecting means for connecting the keypoints detected by the detection means with a straight line.
10. A detection means for detecting multiple key points in a subject contained in an image, A connecting means for connecting the key points detected by the detection means with a straight line, Extraction means for extracting a detection range that surrounds the body parts of a subject that are subject to determination and whose posture or movement is to be determined, and which indicates the detection range of the subject, A determination means for determining whether the subject from which the keypoint was detected and the subject from which the detection range was extracted are the same subject, The system includes a reliability determination means that, depending on the positional relationship between the key point detected by the detection means and determined to be the same subject, and the detection range determined to be the same as the key point, reduces the reliability of at least one of the detection range and the key point, or reduces the reliability of the subject posture determination using the key point determined to be the same. The image processing apparatus is characterized in that, if the determination by the determination means determines that the subjects are the same, the connection means determines, at least, the key point to be connected from among the key points detected by the detection means, according to the positional relationship between the key point detected by the detection means and the detection range extracted by the extraction means.
11. The image processing apparatus according to claim 10, characterized in that the connection means determines the key point to be connected to according to the size of the detection range.
12. The image processing apparatus according to claim 10, characterized in that the determination means determines whether or not the subject is the same subject based on the degree of overlap between the subject in which the key point was detected and the subject from which the detection range was extracted, or the distance between the subject in which the key point was detected and the subject from which the detection range was extracted.
13. The image processing apparatus according to claim 1 or 10, characterized in that the detection means detects at least one of the subject's eyes, ears, top of head, neck, shoulders, elbows, wrists, waist, knees, and ankles as the key point.
14. The image processing apparatus according to claim 1 or 10, characterized in that it is an imaging apparatus equipped with imaging means for photographing a subject and acquiring the image.
15. A detection means for detecting multiple key points for each of the multiple subjects contained in an image, An extraction means for extracting at least one part of the subject's body from the image as the detection range of the subject, in a manner different from that of the detection means, A determination means for determining the posture of a subject based on a plurality of key points detected by the detection means and the detection range of the subject extracted by the extraction means, An image processing apparatus comprising: a reliability determination means that, depending on the positional relationship between the key point detected by the detection means and determined to be the same subject, and the detection range determined to be the same as the key point, reduces the reliability of at least one of the detection range and the key point, or reduces the reliability of the subject posture determination using the key point determined to be the same.
16. A method for a computer to control an image processing device, A detection process that detects multiple keypoints for each of the multiple subjects contained in the image, An extraction step in which body parts of the multiple subjects are extracted as the detection range of the subjects in a manner different from the method in the detection step, A first determination step of determining whether the subject from which the keypoint was detected and the subject from which the detection range was extracted are the same subject, Based on the determination result of the first determination step, a second determination step is performed to determine the posture of the corresponding subject from the plurality of key points detected in the detection step, A control method for an image processing apparatus, comprising: a reliability determination step, which reduces the reliability of at least one of the detection range and the key point, or reduces the reliability of the pose determination of the subject using the key point determined to be the same, according to the positional relationship between the key point detected in the detection step and determined to be the same subject, and the detection range determined to be the same as the key point.
17. A method for a computer to control an image processing device, A detection process that detects multiple key points in a subject contained in an image, A connection step involves connecting the key points detected in the detection step with a straight line, Extraction step of extracting a detection range that surrounds the body parts of the subject to be determined, which are included in the aforementioned image and whose posture or movement is to be determined, and which indicates the detection range of the subject, A determination step to determine whether the subject from which the keypoint was detected and the subject from which the detection range was extracted are the same subject, The system includes a reliability determination step which, depending on the positional relationship between the key point detected in the detection step and determined to be the same subject, and the detection range determined to be the same as the key point, reduces the reliability of at least one of the detection range and the key point, or reduces the reliability of the subject posture determination using the key point determined to be the same. A control method for an image processing apparatus, characterized in that, if the determination in the determination step determines that the subjects are the same, the connection step determines, at least, the key point to be connected from among the key points detected in the detection step, according to the positional relationship between the key point detected in the detection step and the detection range extracted in the extraction step.
18. A method for a computer to control an image processing device, A detection process that detects multiple keypoints for each of the multiple subjects contained in the image, An extraction step, in a method different from the method used in the detection step, in which at least one part of the subject's body is extracted from the image as the detection range of the subject, A determination step is performed to determine the posture of a subject based on the multiple key points detected in the detection step and the detection range of the subject extracted in the extraction step. A control method for an image processing apparatus, comprising: a reliability determination step, which reduces the reliability of at least one of the detection range and the key point, or reduces the reliability of the pose determination of the subject using the key point determined to be the same, according to the positional relationship between the key point detected in the detection step and determined to be the same subject, and the detection range determined to be the same as the key point.
19. A program for causing a computer to execute each of the means of the image processing apparatus according to claim 1, 10, or 15.