Image processing apparatus, image processing method, and program

The image processing apparatus improves template image preparation by using multiple cameras to detect and calculate quality values for human body key points, addressing inefficiencies in existing technologies and enhancing detection accuracy.

JP7708226B2Active Publication Date: 2025-07-15NEC CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023580043
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-02-14
Publication Date
2025-07-15
Estimated Expiration
2042-02-14

AI Technical Summary

Technical Problem

Existing image processing technologies face challenges in preparing high-quality template images for human body detection, leading to reduced detection accuracy and inefficient work processes.

Method used

An image processing apparatus and method that utilize multiple cameras to detect key points of human bodies, identify the same human body across images, calculate a quality value based on the number of detected key points, and output locations or partial images with quality values above a threshold, facilitating the selection of high-quality template images.

Benefits of technology

Enhances the preparation of high-quality template images by identifying and selecting areas with sufficient key points across multiple images, improving detection accuracy and efficiency in human body image processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007708226000001
    Figure 0007708226000001
  • Figure 0007708226000002
    Figure 0007708226000002
  • Figure 0007708226000003
    Figure 0007708226000003
Patent Text Reader

Abstract

The present invention provides an image processing device (10) having: a skeletal structure detecting unit (11) that performs processing for detecting key points of human bodies included in each of a plurality of images generated by imaging the same place with a plurality of cameras; an identifying unit (12) that identifies the same human bodies included in the plurality of images generated by the plurality of cameras; a quality value calculating unit (13) that, for each human body, calculates a quality value of the key points detected from the plurality of images generated by the plurality of cameras; and an output unit (14) that outputs information indicating a location where a human body, the quality value of which is no less than a threshold, is imaged, or a partial image in which said location is cut out of the image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing apparatus, an image processing method, and a program.

Background Art

[0002] Technologies related to the present invention are disclosed in Patent Documents 1 to 4 and Non-Patent Document 1.

[0003] Patent Document 1 discloses a technology for calculating feature amounts of a plurality of key points of a human body included in an image, and searching for an image including a human body with a similar pose or a similar movement based on the calculated feature amounts, or classifying together those with similar poses or movements. Further, Non-Patent Document 1 discloses a technology related to human skeleton estimation.

[0004] Patent Document 2 discloses a technology for extracting skeleton points (joint positions) from each of the images captured by a plurality of cameras and pairing the skeleton points indicating the positions of the same joints of the same person extracted from the plurality of images.

[0005] Patent Document 3 discloses a technology for photographing the same subject from a plurality of directions with a plurality of cameras.

[0006] Patent Document 4 discloses a technology for extracting skeleton points corresponding to an object to be detected (e.g., a person) from an image, and when the number of skeleton points having a reliability equal to or higher than a threshold among the extracted skeleton points is equal to or higher than a threshold, determining that the object is the object to be detected.

Prior Art Documents

Patent Documents

[0007]

Patent Document 1

Patent Document 2

Patent Document 3

Patent Document 4

Non-Patent Document

[0008]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0009] According to the technology disclosed in the above-mentioned Patent Document 1, by registering in advance an image including a human body in a desired posture or a desired movement as a template image, it is possible to detect a human body in a desired posture or a desired movement from the images to be processed. As a result of studying the technology disclosed in such Patent Document 1, the present inventor newly found that the detection accuracy deteriorates unless an image of a certain quality is registered as a template image, and that there is room for improvement in the workability of the work of preparing such a template image.

[0010] Since any of the above-mentioned Patent Documents 1 to 4 and Non-Patent Document 1 does not disclose problems regarding template images and solutions therefor, there is a problem that the above problems cannot be solved.

[0011] An example of the object of the present invention is to provide an image processing apparatus, an image processing method, and a program that solve the problem of workability of the work of preparing a template image of a certain quality in view of the above-mentioned problems.

Means for Solving the Problems

[0012] According to one aspect of the present invention, skeleton structure detection means for performing a process of detecting key points of a human body included in each of a plurality of images generated by photographing the same location with a plurality of cameras; identifying means for identifying the same human body included in the plurality of images generated by the plurality of cameras; quality value calculation means for calculating a quality value of the key points detected from the plurality of images generated by the plurality of cameras for each human body; output means for outputting information indicating a location where a human body with a quality value equal to or greater than a threshold value appears, or a partial image obtained by cutting out the location from the image; An image processing apparatus having the above is provided.

[0013] Also, according to one aspect of the present invention, one or more computers perform a process of detecting key points of a human body included in each of a plurality of images generated by photographing the same location with a plurality of cameras, identify the same human body included in the plurality of images generated by the plurality of cameras, calculate a quality value of the key points detected from the plurality of images generated by the plurality of cameras for each human body, output information indicating a location where a human body with a quality value equal to or greater than a threshold value appears, or a partial image obtained by cutting out the location from the image, An image processing method is provided.

[0014] Also, according to one aspect of the present invention, a computer is provided with skeleton structure detection means for performing a process of detecting key points of a human body included in each of a plurality of images generated by photographing the same location with a plurality of cameras, identifying means for identifying the same human body included in the plurality of images generated by the plurality of cameras, quality value calculation means for calculating a quality value of the key points detected from the plurality of images generated by the plurality of cameras for each human body, Information indicating a location where a human body with the quality value equal to or higher than the threshold value appears, or output means for outputting a partial image obtained by cutting out the location from the image A program that functions as such is provided.

Effect of the Invention

[0015] According to one aspect of the present invention, an image processing apparatus, an image processing method, and a program for solving the problem of workability in preparing a template image of a certain quality can be obtained.

Brief Description of the Drawings

[0016] The above-described object, as well as other objects, features, and advantages, will become more apparent from the following Suitable described embodiments and the accompanying drawings below.

[0017]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Embodiments for Carrying Out the Invention

[0018] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings, the same components are denoted by the same reference numerals, and the description will be omitted as appropriate.

[0019] <First Embodiment> FIG. 1 is a functional block diagram showing an overview of an image processing apparatus 10 according to the first embodiment. As shown in FIG. 1, the image processing apparatus 10 includes a skeleton structure detection unit 11, an identification unit 12, a quality value calculation unit 13, and an output unit 14. The skeleton structure detection unit 11 performs a process of detecting key points of a human body included in each of a plurality of images generated by photographing the same location with a plurality of cameras. The identification unit 12 identifies the same human body included in a plurality of images generated by a plurality of cameras. The quality value calculation unit 13 calculates the quality value of the key points detected from a plurality of images generated by a plurality of cameras for each human body. The output unit 14 outputs information indicating a location where a human body with a quality value equal to or higher than a threshold value appears, or a partial image obtained by cutting out the location from the image.

[0020] According to this image processing apparatus 10, it is possible to solve the problem of workability in the work of preparing a template image of a certain quality.

[0021] <Second Embodiment> 「Overview」 The image processing apparatus 10 detects the key points of the human body included in each of a plurality of images generated by photographing the same location with a plurality of cameras. Next, when the image processing apparatus 10 identifies the same human body included in the plurality of images generated by the plurality of cameras, for each human body, it calculates the quality value of the detected key points based on the value obtained by adding up the number of key points detected from each of the plurality of images generated by the plurality of cameras. Then, the image processing apparatus 10 outputs information indicating a location where a human body with the above quality value equal to or higher than the threshold value appears, or a partial image obtained by cutting out the location from the image.

[0022] The user can prepare a template image of a certain quality by selecting a template image from among the locations where a human body with the above quality value equal to or higher than the threshold value appears.

[0023] "Hardware Configuration" Next, an example of the hardware configuration of the image processing apparatus 10 will be described. The image processing apparatus 10 may be communicably connected to the plurality of cameras. Each functional unit of the image processing apparatus 10 is realized by an arbitrary combination of hardware and software centered around the CPU (Central Processing Unit) of an arbitrary computer, a memory, a program loaded into the memory, a storage unit such as a hard disk for storing the program (which can store programs downloaded from storage media such as CDs (Compact Discs) or servers on the Internet in addition to programs stored in advance at the stage of shipping the apparatus), and a network connection interface. And it is understood by those skilled in the art that there are various variations in the realization method and apparatus.

[0024] FIG. 2 is a block diagram illustrating the hardware configuration of the image processing apparatus 10. As shown in FIG. 2, the image processing apparatus 10 includes a processor 1A, a memory 2A, an input / output interface 3A, a peripheral circuit 4A, and a bus 5A. The peripheral circuit 4A includes various modules. The image processing apparatus 10 may not have the peripheral circuit 4A. Note that the image processing apparatus 10 may be configured by a plurality of physically and / or logically separated apparatuses. In this case, each of the plurality of apparatuses can include the above hardware configuration.

[0025] The bus 5A is a data transmission path for the processor 1A, the memory 2A, the peripheral circuit 4A, and the input / output interface 3A to transmit and receive data to and from each other. The processor 1A is an arithmetic processing device such as a CPU or a GPU (Graphics Processing Unit). The memory 2A is a memory such as a RAM (Random Access Memory) or a ROM (Read Only Memory). The input / output interface 3A includes an interface for acquiring information from an input device, an external device, an external server, an external sensor, a camera, etc., and an interface for outputting information to an output device, an external device, an external server, etc. The input device is, for example, a keyboard, a mouse, a microphone, a physical button, a touch panel, etc. The output device is, for example, a display, a speaker, a printer, a mailer, etc. The processor 1A can issue commands to each module and perform operations based on the operation results thereof.

[0026] "Functional Configuration" FIG. 1 is a functional block diagram showing an overview of the image processing apparatus 10 according to the second embodiment. As shown in FIG. 1, the image processing apparatus 10 includes a skeleton structure detection unit 11, a specifying unit 12, a quality value calculation unit 13, and an output unit 14.

[0027] The skeleton structure detection unit 11 performs a process of detecting key points of a human body included in each of a plurality of images generated by photographing the same location with a plurality of cameras (two or more cameras).

[0028] A plurality of cameras are installed at different positions from each other and simultaneously photograph the same place from different angles. The place to be photographed is not limited. For example, the place to be photographed may be inside a vehicle such as a bus or a train, inside a building or near an entrance / exit, inside an outdoor facility such as a park or near an entrance / exit, or outdoors such as at an intersection.

[0029] The "image" is the image that serves as the basis for the template image. The template image is an image that is pre-registered in the technology disclosed in Patent Document 1 described above, and is an image including a human body in a desired posture or a desired movement (a posture or movement that the user wants to detect). The image may be a moving image composed of a plurality of frame images, or a still image composed of a single image.

[0030] The skeleton structure detection unit 11 detects N (N is an integer of 2 or more) key points of the human body included in the image. When the moving image is the object to be processed, the skeleton structure detection unit 11 performs the process of detecting key points for each frame image. The process by the skeleton structure detection unit 11 is realized using the technology disclosed in Patent Document 1. Although the details are omitted, in the technology disclosed in Patent Document 1, the detection of the skeleton structure is performed using a skeleton estimation technology such as OpenPose disclosed in Non-Patent Document 1. The skeleton structure detected by the technology is composed of "key points" which are characteristic points such as joints, and "bones (bone links)" indicating the links between the key points.

[0031] FIG. 3 shows the skeleton structure of the human body model 300 detected by the skeleton structure detection unit 11, and FIGS. 4 and 5 show detection examples of the skeleton structure. The skeleton structure detection unit 11 uses a skeleton estimation technology such as OpenPose to detect the skeleton structure of a human body model (2D skeleton model) 300 as shown in FIG. 3 from a 2D image. The human body model 300 is a 2D model composed of key points such as joints of a person and bones connecting the key points.

[0032] The skeletal structure detection unit 11 extracts, for example, feature points that can be key points from an image, and detects N key points of a human body with reference to information obtained by machine learning on the images of the key points. The N key points to be detected are determined in advance. The number of key points to be detected (i.e., the value of N) and which parts of the human body are to be detected as key points can vary, and any variations can be adopted.

[0033] Hereinafter, as shown in FIG. 3, it is assumed that the head A1, neck A2, right shoulder A31, left shoulder A32, right elbow A41, left elbow A42, right hand A51, left hand A52, right hip A61, left hip A62, right knee A71, left knee A72, right foot A81, and left foot A82 are defined as N key points (N = 14) to be detected. In the human body model 300 shown in FIG. 3, as the bones of a person connecting these key points, a bone B1 connecting the head A1 and the neck A2, bones B21 and B22 connecting the neck A2 with the right shoulder A31 and the left shoulder A32 respectively, bones B31 and B32 connecting the right shoulder A31 and the left shoulder A32 with the right elbow A41 and the left elbow A42 respectively, bones B41 and B42 connecting the right elbow A41 and the left elbow A42 with the right hand A51 and the left hand A52 respectively, bones B51 and B52 connecting the neck A2 with the right hip A61 and the left hip A62 respectively, bones B61 and B62 connecting the right hip A61 and the left hip A62 with the right knee A71 and the left knee A72 respectively, and bones B71 and B72 connecting the right knee A71 and the left knee A72 with the right foot A81 and the left foot A82 respectively are further defined.

[0034] FIG. 4 shows an example of detecting a person in an upright state. In FIG. 4, an upright person is imaged from the front, and the bones B1, B51 and B52, B61 and B62, B71 and B72 viewed from the front are detected without overlapping each other, and the bones B61 and B71 of the right foot are slightly bent more than the bones B62 and B72 of the left foot.

[0035] Figure 5 shows an example of detecting a person in a crouched state. In Figure 5, the crouched person is imaged from the right side, and the bones B1, B51, and B52, B61 and B62, B71 and B72 of the right leg and left leg are detected respectively. The bones B61 and B71 of the right leg and the bones B62 and B72 of the left leg are greatly bent and overlapping.

[0036] Returning to Figure 1, the specific part 12 identifies the same human body included in a plurality of images generated by a plurality of cameras. The same human body is the human body of the same person. As described above, the plurality of images generated by the plurality of cameras are generated by simultaneously photographing the same location with the plurality of cameras. Therefore, the same person may appear across a plurality of images.

[0037] There are various means for identifying the same human body appearing across a plurality of images. For example, face recognition technology or the like may be used to identify the same person appearing across a plurality of images, and the human bodies detected at the positions within each of the plurality of images in which the same person appears may be identified as the same human body.

[0038] In addition, when the image is a moving image, further, in the same manner as above, or in combination with person tracking technology or the like, the same human body appearing across a plurality of frame images in one moving image can be identified.

[0039] The quality value calculation unit 13 calculates the quality value of the key points detected from a plurality of images generated by a plurality of cameras for each human body. In addition, the quality value calculation unit 13 determines whether the quality value of the detected key points is equal to or greater than a threshold value for each detected human body. Then, the quality value calculation unit 13 identifies the location within the image in which the human body with the quality value of the detected key points equal to or greater than the threshold value appears according to the determination result. These processes will be described in detail below.

[0040] - Process of calculating the quality value of the detected key points - The quality value calculation unit 13 calculates a quality value for each human body. For example, when the human body of person A appears in the first image and the second image, the quality value calculation unit 13 does not calculate the quality value for the human body of person A in the first image and the human body of person A in the second image separately, but calculates one quality value corresponding to the human body of person A.

[0041] As shown in FIG. 6, when the image is a still image, the quality value of the human body of person A is calculated based on a plurality of still images.

[0042] As shown in FIG. 7, when the image is a moving image, the quality value calculation unit 13 identifies a plurality of frame images captured at the same timing from among the plurality of moving images based on the time stamps attached to the moving images. Then, the quality value calculation unit 13 calculates the above quality value for each combination of the plurality of frame images captured at the same timing.

[0043] The "quality value of the detected keypoint" is a value indicating how good the quality of the detected keypoint is, and can be calculated based on various data. In the present embodiment, the quality value calculation unit 13 calculates the quality value based on the value obtained by adding up the number of keypoints detected from each of the plurality of images. The larger the value obtained by adding up the number of keypoints detected from each of the plurality of images, the higher the quality value calculated by the quality value calculation unit 13. For example, the quality value calculation unit 13 may use the value obtained by adding up the number of keypoints detected from each of the plurality of images as the quality value, or may calculate the quality value as the value obtained by normalizing the added-up value according to a predetermined rule.

[0044] Here, the above quality value will be described using a specific example. For simplicity of explanation, it is assumed that two images (the first and second images) generated by photographing the same location with two cameras are processed. For example, it is assumed that K1 (K1 is an integer less than or equal to N) key points are detected from the body of person A shown in the first image, and K2 (K2 is an integer less than or equal to N) key points are detected from the body of person A shown in the second image. In this case, the quality value calculation unit 13 calculates the quality value of the key points detected from the body of person A based on (K1 + K2).

[0045] - Process of identifying a location within an image in which a body with a quality value of the detected key points equal to or greater than a threshold value is shown Based on the calculation result of the process of calculating the above-described quality value, the quality value calculation unit 13 identifies a location within the image in which a body with a quality value of the detected key points equal to or greater than the threshold value is shown. The quality value calculation unit 13 determines, for each detected body, whether the quality value of the detected key points is equal to or greater than the threshold value. Then, the quality value calculation unit 13 identifies the location within the image in which the body with a quality value equal to or greater than the threshold value is shown according to the determination result.

[0046] When the image is a still image, the "location within the image in which a body with a quality value equal to or greater than the threshold value is shown" is a partial area within one still image. In this case, for each still image, for example, in terms of the coordinates of the coordinate system set for the still image, the location within the image in which a body with a quality value of the detected key points equal to or greater than the threshold value is shown is indicated.

[0047] On the other hand, when the image is a moving image, the "location within the image in which a body with a quality value equal to or greater than the threshold value is shown" is a partial area within each of some of the frame images constituting the moving image. In this case, for each moving image, for example, with information indicating some of the frame images among the plurality of frame images (frame identification information, elapsed time from the start, etc.) and the coordinates of the coordinate system set for the frame image, the location within the image in which a body with a quality value of the detected key points equal to or greater than the threshold value is shown is indicated.

[0048] In addition, when the image is a moving image, it is preferable to specify "the part where the human body of the same person appears continuously and the human body appears in each of a plurality of frame images that satisfy the condition that the quality value of the key points detected from the human body is equal to or greater than the threshold value".

[0049] As described above, when the image is a moving image, the specifying unit 12 can specify the human body of the same person that appears across a plurality of frame images. Based on the result of the specification, the quality value calculation unit 13 can specify a plurality of frame images in which the human body of the same person appears continuously.

[0050] Next, the condition that "the quality value of the key points detected from the human body is equal to or greater than the threshold value" will be described. This condition may require that all of the specified plurality of frame images satisfy this condition. That is, in the plurality of frame images specified by the quality value calculation unit 13, the human body of the same person may appear continuously, and the quality value of the key points detected from the human body in all of the frame images may be equal to or greater than the threshold value.

[0051] Alternatively, the above condition may require that at least a part of the specified plurality of frame images satisfy the above condition. That is, in the plurality of frame images specified by the quality value calculation unit 13, the human body of the same person may appear continuously, and the quality value of the key points detected from the human body in at least a part of the frame images may be equal to or greater than the threshold value. In this case, as a condition for the plurality of frame images specified by the quality value calculation unit 13, it is also possible to add, for example, "the number of consecutive frame images in which a human body with a quality value less than the threshold value appears is Q or less". By adding such an additional condition, it is possible to suppress the inconvenience that a human body with a low quality value appears continuously for a predetermined number of frames or more in the plurality of frame images specified by the quality value calculation unit 13.

[0052] The output unit 14 outputs information indicating a location where a human body (a human body whose quality value of the detected keypoints is equal to or higher than a threshold value) with a quality value equal to or higher than the threshold value appears, or a partial image obtained by cutting out the location from the image. When the image is a moving image, the output unit 14 may output information indicating a location where the human body of the same person continuously appears and satisfies the condition that "the quality value of the keypoints detected from the human body is equal to or higher than the threshold value" in each of a plurality of frame images, or a partial image obtained by cutting out the location from the image.

[0053] Note that when the output unit 14 outputs a partial image, the image processing apparatus 10 may include a processing unit that cuts out a location where a human body with a quality value equal to or higher than the threshold value appears from the image to generate a partial image. Then, the output unit 14 can output the partial image generated by the processing unit.

[0054] In addition, the output unit 14 may output partial images cut out from each of a plurality of images generated by a plurality of cameras, after associating with each other those related to the same human body. Also, the output unit 14 may output information indicating a location where a human body with a quality value equal to or higher than the threshold value appears in each of a plurality of images generated by a plurality of cameras, after associating with each other the information related to the same human body. Further, the output unit 14 may output information indicating that the image includes a human body with a quality value equal to or higher than the threshold value.

[0055] The above-mentioned "location where a human body with a quality value equal to or higher than the threshold value appears in the image" serves as a candidate for the template image. Based on the above information or the above partial image, the user can view the location where a human body with a quality value equal to or higher than the threshold value appears, and select, as the template image, a location including a human body with a desired posture or a desired movement from among them.

[0056] Fig. 8 schematically shows an example of the information output by the output unit 14. In the example shown in Fig. 8, the human body identification information for identifying a plurality of detected human bodies from each other and the attribute information of each human body are displayed in association with each other. And, as an example of the attribute information, a quality value, the number of detected key points, information indicating a location within the image (information indicating the location where the above-described human body appears), and the shooting date and time of the image are displayed. The number of detected key points is a value obtained by adding up the number of key points detected from each of the plurality of images. The attribute information may also include information indicating the installation position (shooting position) of the camera that captured the image (e.g., the rear inside of the No. 102 bus, the entrance of XX Park, etc.) and the attribute information of the person calculated by image analysis (e.g., gender, age group, body type, etc.).

[0057] Next, an example of the processing flow of the image processing apparatus 10 will be described using the flowchart of Fig. 9.

[0058] When the image processing apparatus 10 acquires a plurality of images generated by photographing the same location with a plurality of cameras (S10), it performs a process of detecting the key points of the human body included in each of the plurality of images (S11). Next, the image processing apparatus 10 identifies the same human body included in the plurality of images generated by the plurality of cameras (S12). Note that the processing order of S11 and S12 may be reversed, or these two processes may be performed in parallel.

[0059] Next, the image processing apparatus 10 calculates the quality value of the key points detected from the plurality of images generated by the plurality of cameras for each human body (S13). In the second embodiment, the image processing apparatus 10 calculates the quality value based on the value obtained by adding up the number of key points detected from each of the plurality of images generated by the plurality of cameras. The higher the added value, the higher the quality value calculated by the image processing apparatus 10.

[0060] Next, it is determined whether the quality value of the keypoints detected for each person is equal to or greater than the threshold value (S14). Next, the image processing apparatus 10 specifies a location within the image in which a person whose detected keypoint quality value is equal to or greater than the threshold value appears, according to the determination result of S14 (S15). Then, the image processing apparatus 10 outputs information indicating the location in which a person whose quality value is equal to or greater than the threshold value appears, or a partial image obtained by cutting out the location from the image (S16). For example, the image processing apparatus 10 may output partial images cut out from each of a plurality of images generated by a plurality of cameras, after associating with each other those related to the same person. Further, the image processing apparatus 10 may output information indicating the locations in which a person whose quality value is equal to or greater than the threshold value appears in each of the plurality of images generated by the plurality of cameras, after associating with each other the information related to the same person.

[0061] "Operation and Effect" According to the image processing apparatus 10 of the second embodiment, the same operation and effect as those of the first embodiment are achieved. Further, according to the image processing apparatus 10 of the second embodiment, it is possible to present to the user, as a candidate for the template image, a location in which a person appears whose value obtained by adding up the number of keypoints detected from each of the plurality of images generated by the plurality of cameras is large. The user can easily prepare a template image that satisfies a certain quality by selecting a template image from among the presented template image candidates.

[0062] In addition, as shown in FIG. 10, there may be a case where some key points of the human body P are not detected because they are hidden by an obstacle Q or other parts of the human body P itself. An image of a human body with many undetected key points is not preferable as a template image. However, as shown in FIG. 11, when the undetected key points are detected in an image generated by another camera, the shortage can be compensated by the feature amounts of the key points detected from the other images. In this way, although a single image is not preferable as a template image, a combination of a plurality of images taken at the same timing may be preferable as a template image. By calculating the quality value of the key points detected from a plurality of images generated by a plurality of cameras for each human body, and selecting candidates for the template image based on the quality value, it is possible to select, as candidates for the template image, an image of a human body that is preferable as a template image when a plurality of images taken at the same timing as described above are combined.

[0063] <Third Embodiment> In the image processing apparatus 10 of the third embodiment, the method of calculating the quality value is different from that of the first and second embodiments.

[0064] The quality value calculation unit 13 calculates a quality value based on the number of key points detected in at least one of the plurality of images generated by a plurality of cameras among the plurality of key points to be detected (the above-described N key points), or the number of key points not detected in any of the plurality of images generated by a plurality of cameras among the plurality of key points to be detected.

[0065] The quality value calculation unit 13 calculates a higher quality value as the number of key points detected in at least one of the plurality of images generated by a plurality of cameras among the plurality of key points to be detected is larger. For example, the quality value calculation unit 13 may use the number of key points detected in at least one of the plurality of images generated by a plurality of cameras among the plurality of key points to be detected as the quality value, or may calculate, as the quality value, a value obtained by normalizing the number according to a predetermined rule.

[0066] In addition, the quality value calculation unit 13 calculates a higher quality value as the number of keypoints that are not detected in any of the plurality of images generated by the plurality of cameras among the plurality of keypoints to be detected is smaller. For example, the quality value calculation unit 13 may use, as the quality value, a number obtained by subtracting the number of keypoints that are not detected in any of the plurality of images generated by the plurality of cameras among the plurality of keypoints to be detected from a predetermined value, or may calculate, as the quality value, a value obtained by normalizing the number according to a predetermined rule.

[0067] Here, the above quality value will be described using a specific example. For simplicity of explanation, it is assumed that two images (first and second images) generated by photographing the same location with two cameras are processed. Also, the plurality of keypoints to be detected are five keypoints C1 to C5. It is assumed that keypoints C1 to C3 are detected from the first image and keypoints C2 to C4 are detected from the second image. In this case, the keypoints that are detected in at least one of the plurality of images generated by the plurality of cameras among the plurality of keypoints to be detected are keypoints C1 to C4, and the number thereof is "4". And the keypoint that is not detected in any of the plurality of images generated by the plurality of cameras among the plurality of keypoints to be detected is keypoint C5, and the number thereof is "1". The quality value calculation unit 13 calculates the quality value of the keypoints detected from the human body based on such numbers.

[0068] In addition, the quality value calculation unit 13 may calculate the quality value by combining the method described in the second embodiment with a method based on the number of keypoints detected in at least one of the plurality of images generated by a plurality of cameras among the plurality of keypoints to be detected, or the number of keypoints not detected in any of the plurality of images generated by a plurality of cameras among the plurality of keypoints to be detected. For example, the quality value calculation unit 13 calculates a first quality value by normalizing the quality value calculated by the method described in the second embodiment according to a predetermined rule, and calculates the number of keypoints detected in at least one of the plurality of images generated by a plurality of cameras among the plurality of keypoints to be detected, or a quality value calculated by a method based on the number of keypoints not detected in any of the plurality of images generated by a plurality of cameras among the plurality of keypoints to be detected is normalized according to a predetermined rule to calculate a second quality value. Then, the quality value calculation unit 13 may calculate a statistical value (average value, maximum value, minimum value, median value, mode value, weighted average value, etc.) of the first quality value and the second quality value as the quality value of the human body.

[0069] Other configurations of the image processing apparatus 10 according to the third embodiment are the same as those of the first and second embodiments.

[0070] According to the image processing apparatus 10 of the third embodiment, the same operational effects as those of the first and second embodiments are achieved. Further, according to the image processing apparatus 10 of the third embodiment, a portion of the human body where a large number of keypoints detected in at least one of the N keypoints to be detected is imaged can be presented to the user as a candidate for the template image. The user can easily prepare a template image that satisfies a certain quality in terms of the number of keypoints detected in at least one image by selecting a template image from among the presented template image candidates.

[0071] <Fourth Embodiment> The image processing apparatus 10 according to the fourth embodiment differs from the first to third embodiments in the way of calculating the quality value.

[0072] The quality value calculation unit 13 calculates the partial quality value of the keypoints detected from each of the plurality of images generated by the plurality of cameras for each image, and integrates the partial quality values for each image to calculate the quality value for each person. As shown in FIG. 12, when the image is a still image, the quality value calculation unit 13 calculates the partial quality value for each person detected from each of the plurality of images. Then, the quality value calculation unit 13 integrates the partial quality values of the bodies of the same person to calculate the quality value of the body of that person.

[0073] As shown in FIG. 13, when the image is a moving image, the quality value calculation unit 13 identifies a plurality of frame images captured at the same timing from among the plurality of moving images based on the time stamps attached to the moving images. Then, the quality value calculation unit 13 integrates the partial quality values of the bodies of the same person detected from each of the plurality of frame images for each combination of the plurality of frame images captured at the same timing, and calculates the quality value of the body of that person.

[0074] The "partial quality value of the detected keypoints" is a value indicating how good the quality of the detected keypoints is, and can be calculated based on various data. In the present embodiment, the quality value calculation unit 13 calculates the partial quality value based on the confidence level of the detection result of the keypoints. In the following embodiments, an example of calculating the above partial quality value based on data other than the confidence level of the detection result of the keypoints will be described. The method for calculating the confidence level is not particularly limited. For example, in a skeleton estimation technique such as OpenPose, the score output associated with each detected keypoint may be used as the confidence level of each keypoint.

[0075] The quality value calculation unit 13 calculates a higher partial quality value as the confidence level of the detection result of the key points is higher. For example, the quality value calculation unit 13 may calculate a statistical value (average value, maximum value, minimum value, median value, mode value, weighted average value, etc.) of the confidence levels of each of the N key points detected from the human body as the partial quality value of the human body. When some of the N key points are not detected, the confidence level of the undetected key point may be a fixed value such as "0". This fixed value is set to be lower than the confidence level of the detected key points.

[0076] Note that when the image is a still image, the quality value calculation unit 13 calculates the partial quality value for each human body detected from the still image. On the other hand, when the image is a moving image, the quality value calculation unit 13 calculates the partial quality value for each human body detected from each of the plurality of frame images.

[0077] Next, a process of calculating the quality value by integrating the partial quality values of the key points detected from each of the plurality of images generated by the plurality of cameras will be described. The quality value calculation unit 13 can calculate a statistical value (average value, maximum value, minimum value, median value, mode value, weighted average value, etc.) of the partial quality values of the key points detected from each of the plurality of images generated by the plurality of cameras as the quality value of the human body.

[0078] In addition, the quality value calculation unit 13 may calculate the quality value by combining at least one of the methods described in the second and third embodiments with the method based on the confidence level of the detection result of the key points. For example, the quality value calculation unit 13 performs at least one of the process of normalizing the quality value calculated by the method described in the second embodiment with a predetermined rule to calculate the first quality value, and the process of normalizing the quality value calculated by the method described in the third embodiment with a predetermined rule to calculate the second quality value. Further, the quality value calculation unit 13 normalizes the quality value calculated by the method based on the confidence level of the detection result of the key points with a predetermined rule to calculate the third quality value. Then, the quality value calculation unit 13 may calculate a statistical value (average value, maximum value, minimum value, median value, mode value, weighted average value, etc.) of at least one of the first and second quality values and the third quality value as the quality value of the human body.

[0079] Other configurations of the image processing apparatus 10 according to the fourth embodiment are the same as those of the first to third embodiments.

[0080] According to the image processing apparatus 10 of the fourth embodiment, the same operational effects as those of the first to third embodiments are achieved. Further, according to the image processing apparatus 10 of the fourth embodiment, it is possible to present to the user, as candidates for the template image, portions where a human body with a high confidence level in the detection result of the keypoints is captured. The user can easily prepare a template image whose confidence level in the detection result of the keypoints satisfies a certain quality by selecting a template image from among the presented template image candidates.

[0081] <Fifth Embodiment> In the image processing apparatus 10 according to the fifth embodiment, the method of calculating the quality value is different from that of the first to fourth embodiments.

[0082] The quality value calculation unit 13 calculates, for each image, the partial quality value of the keypoints detected from each of the plurality of images generated by the plurality of cameras, and integrates the partial quality values for each image to calculate the quality value for each human body. Then, the quality value calculation unit 13 calculates the partial quality value of a human body with a relatively large number of detected keypoints to be higher than the partial quality value of a human body with a relatively small number of detected keypoints. For example, the quality value calculation unit 13 may use the number of detected keypoints as the partial quality value. Alternatively, weight points may be set for each of the plurality of keypoints. Higher weight points are set for relatively important keypoints. Then, the quality value calculation unit 13 may calculate, as the partial quality value, the value obtained by adding up the weight points of each of the detected keypoints.

[0083] In addition, the quality value calculation unit 13 may calculate a quality value by combining at least one of the methods described in the second to fourth embodiments and a method based on the number of the above key points. For example, the quality value calculation unit 13 performs a process of calculating a first quality value by normalizing the quality value calculated by the method described in the second embodiment according to a predetermined rule, a process of calculating a second quality value by normalizing the quality value calculated by the method described in the third embodiment according to a predetermined rule, and a process of calculating a third quality value by normalizing the quality value calculated by the method described in the fourth embodiment according to a predetermined rule. At least one of these processes is performed. Further, the quality value calculation unit 13 calculates a fourth quality value by normalizing the quality value calculated by the method based on the number of the above key points according to a predetermined rule. Then, the quality value calculation unit 13 may calculate at least one of the first to third quality values and a statistical value (average value, maximum value, minimum value, median value, mode value, weighted average value, etc.) of the fourth quality value as the quality value of the human body.

[0084] Other configurations of the image processing apparatus 10 according to the fifth embodiment are the same as those in the first to fourth embodiments.

[0085] According to the image processing apparatus 10 of the fifth embodiment, the same operational effects as those in the first to fourth embodiments are achieved. Further, according to the image processing apparatus 10 of the fifth embodiment, a location where a human body with many detected key points appears can be presented to the user as a candidate for the template image. The user can easily prepare a template image whose number of detected key points satisfies a certain quality by selecting a template image from among the presented template image candidates.

[0086] <Sixth Embodiment> The image processing apparatus 10 according to the sixth embodiment differs from the first to fifth embodiments in the way of calculating the quality value.

[0087] The quality value calculation unit 13 calculates the partial quality value of the keypoints detected from each of the plurality of images generated by the plurality of cameras for each image, and integrates the partial quality values for each image to calculate the quality value for each person. Then, the quality value calculation unit 13 calculates the partial quality value based on the degree of overlap with other persons. Note that the state where "the body of person A overlaps with the body of person B" includes the state where the body of person A is partially or completely hidden by the body of person B, the state where the body of person A partially or completely hides the body of person B, and the state where both occur. Hereinafter, the calculation method will be specifically described.

[0088] - First method - The quality value calculation unit 13 calculates the partial quality value of a person whose body does not overlap with other persons to be higher than the partial quality value of a person whose body overlaps with other persons. For example, a rule is created in advance and stored in the image processing apparatus 10, where the partial quality value of a person whose body does not overlap with other persons is set as X1, and the partial quality value of a person whose body overlaps with other persons is set as X2. Note that X1 > X2. Then, the quality value calculation unit 13 calculates the partial quality value of a person whose body does not overlap with other persons as X1 and the partial quality value of a person whose body overlaps with other persons as X2 based on the rule.

[0089] Whether or not it overlaps with other persons may be specified based on the degree of overlap of the human body model 300 (see FIG. 3) detected by the skeleton structure detection unit 11, or may be specified based on the degree of overlap of the bodies shown in the image.

[0090] For example, when the distance between predetermined keypoints (e.g., head A1) of two persons in the image is equal to or less than a threshold value, it may be determined that the two persons overlap. In this case, the threshold value may be a variable value that changes according to the size of the detected person in the image. The larger the size of the detected person in the image, the larger the threshold value. Note that instead of the size of the person in the image, the length of a predetermined bone (e.g., bone B1 connecting head A1 and neck A2) or the size of the face in the image may be adopted.

[0091] In addition, if any bone of a certain human body intersects with any bone of another human body, the two human bodies may be determined to overlap each other.

[0092] -Second method- The quality value calculation unit 13 calculates the partial quality value of a human body that does not overlap with other human bodies to be higher than the partial quality value of a human body that overlaps with other human bodies, and among the human bodies that overlap with other human bodies, calculates the partial quality value of the human body located on the front side to be higher than the partial quality value of the human body located on the rear side.

[0093] That is, the quality value calculation unit 13 calculates the partial quality value of a human body that does not overlap with other human bodies to be the highest, calculates the partial quality value of a human body that overlaps with other human bodies but is located on the front side to be the next highest, and calculates the partial quality value of a human body that overlaps with other human bodies and is located on the rear side to be the lowest.

[0094] For example, let the partial quality value of a human body that does not overlap with other human bodies be X1, and the partial quality value X 21 of a human body that overlaps with other human bodies and is located on the front side, and the partial quality value X 22 of a human body that overlaps with other human bodies and is located on the rear side. A rule is created in advance and stored in the image processing apparatus 10. Note that X1>X 21 >X 22 is satisfied. Then, based on the rule, the quality value calculation unit 13 calculates the partial quality value of a human body that does not overlap with other human bodies to be X1, calculates the partial quality value of a human body that overlaps with other human bodies and is located on the front side to be X 21 and calculates the partial quality value of a human body that overlaps with other human bodies and is located on the rear side to be X 22

[0095] ​Whether a person is in front of or behind another person may be specified based on the degree of concealment or loss of the human body model 300 (see FIG. 3) detected by the skeletal structure detection unit 11, or may be specified based on the degree of concealment of the body shown in the image. For example, among two overlapping human bodies, if all N key points of one are detected and only a part of the N key points of the other are detected, it can be determined that the human body with all N key points detected is located in the front, and the other human body is located in the back.

[0096] In addition, the quality value calculation unit 13 may calculate the quality value by combining at least one of the methods described in the second to fifth embodiments and the method based on the degree of overlap of the above human bodies. For example, the quality value calculation unit 13 performs a process of normalizing the quality value calculated by the method described in the second embodiment according to a predetermined rule to calculate a first quality value, a process of normalizing the quality value calculated by the method described in the third embodiment according to a predetermined rule to calculate a second quality value, a process of normalizing the quality value calculated by the method described in the fourth embodiment according to a predetermined rule to calculate a third quality value, and a process of normalizing the quality value calculated by the method described in the fifth embodiment according to a predetermined rule to calculate a fourth quality value. At least one of the processes is performed. Further, the quality value calculation unit 13 normalizes the quality value calculated by the method based on the degree of overlap of the above human bodies according to a predetermined rule to calculate a fifth quality value. Then, the quality value calculation unit 13 may calculate at least one of the first to fourth quality values and a statistical value (average value, maximum value, minimum value, median value, mode value, weighted average value, etc.) of the fifth quality value as the quality value of the human body.

[0097] Other configurations of the image processing apparatus 10 according to the sixth embodiment are the same as those of the first to fifth embodiments.

[0098] According to the image processing apparatus 10 of the sixth embodiment, the same operational effects as those of the first to fifth embodiments are achieved. Further, according to the image processing apparatus 10 of the sixth embodiment, a location where a human body that does not overlap with other human bodies appears, or a location where a human body that overlaps with other human bodies but is located in the front appears can be presented to the user as a candidate for the template image. By selecting a template image from among the candidates for the template image presented in this way, the user can easily prepare a template image whose degree of overlap with other human bodies satisfies a certain quality.

[0099] <Seventh Embodiment> The image processing apparatus 10 of the seventh embodiment differs from the first to sixth embodiments in the way of calculating the quality value.

[0100] First, the skeleton structure detection unit 11 performs processing to detect a person area in the image and detect key points within the detected person area. That is, the skeleton structure detection unit 11 does not target all areas in the image for the processing of detecting key points, but only the detected person area for the processing of detecting key points. The details of the processing for detecting a person area in the image are not particularly limited and may be realized using an object detection technique such as YOLO.

[0101] The quality value calculation unit 13 calculates a partial quality value for each key point detected from each of the plurality of images generated by the plurality of cameras for each image, and integrates the partial quality values for each image to calculate a quality value for each human body. Then, the quality value calculation unit 13 calculates a partial quality value based on the confidence level of the detection result of the person area. The method for calculating the confidence level of the detection result of the person area is not particularly limited. For example, in an object detection technique such as YOLO, the score (which may also be referred to as a reliability, etc.) output in association with the detected object area may be used as the confidence level for each person area.

[0102] The quality value calculation unit 13 calculates a higher partial quality value as the confidence level of the detection result of the person area is higher. For example, the quality value calculation unit 13 may calculate the confidence level of the detection result of the person area as the partial quality value.

[0103] In addition, the quality value calculation unit 13 may calculate a quality value by combining at least one of the methods described in the second to sixth embodiments and a method based on the confidence level of the detection result of the human region. For example, the quality value calculation unit 13 performs a process of calculating a first quality value by normalizing the quality value calculated by the method described in the second embodiment according to a predetermined rule, a process of calculating a second quality value by normalizing the quality value calculated by the method described in the third embodiment according to a predetermined rule, a process of calculating a third quality value by normalizing the quality value calculated by the method described in the fourth embodiment according to a predetermined rule, a process of calculating a fourth quality value by normalizing the quality value calculated by the method described in the fifth embodiment according to a predetermined rule, and a process of calculating a fifth quality value by normalizing the quality value calculated by the method described in the sixth embodiment according to a predetermined rule, and performs at least one of these processes. Further, the quality value calculation unit 13 calculates a sixth quality value by normalizing the quality value calculated by the method based on the confidence level of the detection result of the human region according to a predetermined rule. Then, the quality value calculation unit 13 may calculate a statistical value (average value, maximum value, minimum value, median value, mode value, weighted average value, etc.) of at least one of the first to fifth quality values and the sixth quality value as the quality value of the human body.

[0104] Other configurations of the image processing apparatus 10 according to the seventh embodiment are the same as those of the first to sixth embodiments.

[0105] According to the image processing apparatus 10 of the seventh embodiment, the same operational effects as those of the first to sixth embodiments are achieved. Further, according to the image processing apparatus 10 of the seventh embodiment, a location where a person is captured with high confidence can be presented to the user as a candidate for the template image. The user can easily prepare a template image whose detection result of the human region satisfies a certain quality by selecting a template image from among the candidates for the template image presented in this way.

[0106] <Eighth Embodiment> The image processing apparatus 10 according to the eighth embodiment differs from the first to seventh embodiments in the way of calculating the quality value.

[0107] The quality value calculation unit 13 calculates the partial quality values of the keypoints detected from each of the plurality of images generated by the plurality of cameras for each image, and integrates the partial quality values for each image to calculate the quality value for each person. Then, the quality value calculation unit 13 calculates the partial quality value based on the size of the person in the image. The quality value calculation unit 13 calculates the partial quality value of a relatively large person to be higher than the partial quality value of a relatively small person. The size of the person in the image may be indicated by the size (area, etc.) of the person area shown in the seventh embodiment, or may be indicated by the length of a predetermined bone (e.g., bone B1), or may be indicated by the length between a predetermined two keypoints (e.g., keypoints A31 and A32), or may be indicated by other methods.

[0108] In addition, the quality value calculation unit 13 may calculate the quality value by combining at least one of the methods described in the second to seventh embodiments and the method based on the size of the above-mentioned person. For example, the quality value calculation unit 13 performs a process of calculating a first quality value by normalizing the quality value calculated by the method described in the second embodiment according to a predetermined rule, a process of calculating a second quality value by normalizing the quality value calculated by the method described in the third embodiment according to a predetermined rule, a process of calculating a third quality value by normalizing the quality value calculated by the method described in the fourth embodiment according to a predetermined rule, a process of calculating a fourth quality value by normalizing the quality value calculated by the method described in the fifth embodiment according to a predetermined rule, a process of calculating a fifth quality value by normalizing the quality value calculated by the method described in the sixth embodiment according to a predetermined rule, and a process of calculating a sixth quality value by normalizing the quality value calculated by the method described in the seventh embodiment according to a predetermined rule, and performs at least one of these processes. Further, the quality value calculation unit 13 calculates a seventh quality value by normalizing the quality value calculated by the method based on the size of the above-mentioned person according to a predetermined rule. Then, the quality value calculation unit 13 may calculate at least one of the first to sixth quality values and the statistical value (average value, maximum value, minimum value, median value, mode value, weighted average value, etc.) of the seventh quality value as the quality value of the person.

[0109] Other configurations of the image processing apparatus 10 according to the eighth embodiment are the same as those in the first to seventh embodiments.

[0110] According to the image processing apparatus 10 of the eighth embodiment, the same operational effects as those of the first to seventh embodiments are achieved. Further, according to the image processing apparatus 10 of the eighth embodiment, a portion where a human body is imaged to a certain extent can be presented to the user as a candidate for the template image. By selecting a template image from among the presented template image candidates in this way, the user can easily prepare a template image in which the size of the human body satisfies a certain quality.

[0111] <Ninth Embodiment> The image processing apparatus 10 of the ninth embodiment is different from the first to eighth embodiments in the process of selecting a portion to be a candidate for the template image.

[0112] The quality value calculation unit 13 specifies a portion where a human body is imaged, where the quality value is equal to or greater than a threshold value and the number of keypoints detected from each of the plurality of images generated by the plurality of cameras is equal to or greater than a lower limit value. Then, the output unit 14 outputs information indicating a portion where a human body is imaged, where the quality value is equal to or greater than a threshold value and the number of keypoints detected from each of the plurality of images generated by the plurality of cameras is equal to or greater than a lower limit value, or a partial image obtained by cutting out the portion from the image.

[0113] The other configuration of the image processing apparatus 10 of the ninth embodiment is the same as that of the first to eighth embodiments.

[0114] According to the image processing apparatus 10 of the ninth embodiment, the same operational effects as those of the first to eighth embodiments are achieved. Further, according to the image processing apparatus 10 of the ninth embodiment, a portion where a human body is imaged, where the above-described quality value is equal to or greater than a threshold value and keypoints equal to or greater than a lower limit value are detected in each of the plurality of images generated by the plurality of cameras, can be presented to the user as a candidate for the template image. By selecting a template image from among the presented template image candidates in this way, the user can easily prepare a template image in which the above-described quality value is equal to or greater than a threshold value and the number of keypoints detected in each of the plurality of images satisfies a certain quality.

[0115] <Modification Example> In the above embodiment, when the image is a moving image, the "portion where a person whose quality value is equal to or greater than the threshold value appears" is a partial area within each of some of the frame images constituting the moving image. Then, the output unit 14 outputs information indicating such a portion or a partial image obtained by cutting out such a portion from the image. This is a configuration assuming that a single frame image may include a plurality of persons.

[0116] As a modification example, when the image is a moving image, the portion where a person whose quality value is equal to or greater than the threshold value appears may be a part of the plurality of frame images constituting the moving image. Then, the output unit 14 may output information indicating a part of such a plurality of frame images or a partial image obtained by cutting out some of the frame images from the image. Further, the frame image itself in which a person whose quality value is equal to or greater than the threshold value appears may be output as a candidate for the template image. This is a configuration assuming that a single frame image may include only one person whose quality value is equal to or greater than the threshold value.

[0117] As described above, the embodiments of the present invention have been described with reference to the drawings, but these are examples of the present invention, and various configurations other than the above can also be adopted.

[0118] Also, in the plurality of flowcharts used in the above description, a plurality of steps (processes) are described in order, but the execution order of the steps executed in each embodiment is not limited to the described order. In each embodiment, the order of the illustrated steps can be changed within a range that does not substantially affect the content. Also, the above-described embodiments can be combined within a range where the contents do not conflict.

[0119] Some or all of the above embodiments can also be described as follows, but are not limited thereto. 1. Skeleton structure detection means for performing a process of detecting key points of a person included in each of a plurality of images generated by photographing the same location with a plurality of cameras; Identifying means for identifying the same human body included in the plurality of images generated by the plurality of cameras; Quality value calculation means for calculating, for each human body, a quality value of the keypoints detected from the plurality of images generated by the plurality of cameras; Output means for outputting information indicating a location where a human body with a quality value equal to or higher than a threshold value appears, or a partial image obtained by cutting out the location from the image; An image processing apparatus having the above. 2. The image processing apparatus according to 1, wherein the quality value calculation means calculates the quality value based on a value obtained by adding up the number of keypoints detected from each of the plurality of images generated by the plurality of cameras. 3. The image processing apparatus according to 1 or 2, wherein the quality value calculation means calculates the quality value based on the number of keypoints detected in at least one of the plurality of images generated by the plurality of cameras among the plurality of keypoints to be detected, or the number of keypoints not detected in any of the plurality of images generated by the plurality of cameras among the plurality of keypoints to be detected. 4. The image processing apparatus according to any one of 1 to 3, wherein the quality value calculation means calculates a partial quality value of the keypoints detected from each of the plurality of images generated by the plurality of cameras for each image, and integrates the partial quality values for each image to calculate the quality value. 5. The image processing apparatus according to 4, wherein the quality value calculation means calculates the partial quality value based on the confidence level of the detection result of the keypoints. 6. The skeleton structure detection means performs a process of detecting a person area in the image and detecting the keypoints within the detected person area. 7. The image processing apparatus according to 4 or 5, wherein the quality value calculation means calculates the partial quality value based on the confidence level of the detection result of the person area. 8. The image processing apparatus according to any one of 4 to 6, wherein the quality value calculation means calculates the partial quality value based on the degree of overlap with other human bodies. 8. The image processing apparatus according to claim 7, wherein the quality value calculation means calculates the partial quality value of a human body not overlapping with other human bodies to be higher than the partial quality value of a human body overlapping with other human bodies. 9. The image processing apparatus according to claim 8, wherein the quality value calculation means calculates the partial quality value of a human body located on the front side among human bodies overlapping with other human bodies to be higher than the partial quality value of a human body located on the rear side. 10. The image processing apparatus according to any one of claims 4 to 9, wherein the quality value calculation means calculates the partial quality value of a human body with a relatively large number of the detected keypoints to be higher than the partial quality value of a human body with a relatively small number of the detected keypoints. 11. The image processing apparatus according to any one of claims 4 to 10, wherein the quality value calculation means calculates the partial quality value based on the size of the human body in the image. 12. One or more computers perform a process of detecting keypoints of a human body included in each of a plurality of images generated by photographing the same location with a plurality of cameras, identify the same human body included in the plurality of images generated by the plurality of cameras, calculate the quality value of the keypoints detected from the plurality of images generated by the plurality of cameras for each human body, output information indicating a location where a human body with a quality value equal to or higher than a threshold value appears, or a partial image obtained by cutting out the location from the image. Image processing method. 13. A computer as a skeleton structure detection means for performing a process of detecting keypoints of a human body included in each of a plurality of images generated by photographing the same location with a plurality of cameras, as an identification means for identifying the same human body included in the plurality of images generated by the plurality of cameras, as a quality value calculation means for calculating the quality value of the keypoints detected from the plurality of images generated by the plurality of cameras for each human body, as an output means for outputting information indicating a location where a human body with a quality value equal to or higher than a threshold value appears, or a partial image obtained by cutting out the location from the image. Program for causing the computer to function as such.

Description of Symbols

[0120] 10 Image processing apparatus 11 Skeleton structure detection unit 12 Specifying unit 13 Quality value calculation unit 14 Output unit 1A Processor 2A Memory 3A Input / output I / F 4A Peripheral circuit 5A Bus

Claims

1. Skeletal structure detection means for performing a process of detecting key points of a human body included in each of a plurality of images generated by photographing the same location with a plurality of cameras; Identification means for identifying the same human body included in the plurality of images generated by the plurality of cameras; Quality value calculation means for calculating a quality value of the key points detected from the plurality of images generated by the plurality of cameras for each human body; Output means for outputting information indicating a location where a human body with a quality value equal to or higher than a threshold value appears, or a partial image obtained by cutting out the location from the image; having; The quality value calculation means calculates a partial quality value of the key points detected from each of the plurality of images generated by the plurality of cameras for each image, and integrates the partial quality values for each image to calculate the quality value. An image processing apparatus.

2. The quality value calculation means is the confidence level of the detection result of the key points, the confidence level of the detection result of the person area in the image, the degree of overlap with other human bodies, the size of the human body on the image, The image processing apparatus according to claim 1, wherein the partial quality value is calculated based on at least one of them.

3. The image processing apparatus according to claim 2, wherein the quality value calculation means calculates the partial quality value of a human body that does not overlap with other human bodies to be higher than the partial quality value of a human body that overlaps with other human bodies.

4. The image processing apparatus according to claim 3, wherein the quality value calculation means calculates the partial quality value of a human body located in the front among the human bodies that overlap with other human bodies to be higher than the partial quality value of a human body located in the rear.

5. The image processing apparatus according to any one of claims 1 to 4, wherein the quality value calculation means calculates the partial quality value of a human body with a relatively large number of detected key points to be higher than the partial quality value of a human body with a relatively small number of detected key points.

6. The image processing apparatus according to any one of claims 1 to 5, wherein the quality value calculation means calculates the quality value based on a value obtained by adding up the number of key points detected from each of the plurality of images generated by the plurality of cameras.

7. The quality value calculation means calculates the quality value based on the number of the keypoints detected in at least one of the plurality of images generated by the plurality of the cameras among the plurality of keypoints to be detected, or the number of the keypoints not detected in any of the plurality of images generated by the plurality of the cameras among the plurality of keypoints to be detected. The image processing apparatus according to any one of claims 1 to 6.

8. One or more computers perform a process of detecting keypoints of a human body included in each of a plurality of images generated by photographing the same location with a plurality of cameras, identify the same human body included in the plurality of images generated by the plurality of the cameras, calculate, for each human body, a quality value of the keypoints detected from the plurality of images generated by the plurality of the cameras, output information indicating a location where a human body with a quality value equal to or higher than a threshold value appears, or a partial image obtained by cutting out the location from the image, In the process of calculating the quality value, a partial quality value of the keypoints detected from each of the plurality of images generated by the plurality of the cameras is calculated for each image, and the partial quality values for each image are integrated to calculate the quality value. An image processing method.

9. A computer is caused to function as skeleton structure detection means for performing a process of detecting keypoints of a human body included in each of a plurality of images generated by photographing the same location with a plurality of cameras, identifying means for identifying the same human body included in the plurality of images generated by the plurality of the cameras, quality value calculation means for calculating, for each human body, a quality value of the keypoints detected from the plurality of images generated by the plurality of the cameras, output means for outputting information indicating a location where a human body with a quality value equal to or higher than a threshold value appears, or a partial image obtained by cutting out the location from the image, and the quality value calculation means calculates a partial quality value of the keypoints detected from each of the plurality of images generated by the plurality of the cameras for each image, and integrates the partial quality values for each image to calculate the quality value. A program. ​

Citation Information

Patent Citations

  • External parameter estimation method, estimation device and estimation program of camera

    JP2019102877A

  • Information processing device, storage device, image processing device, image processing system, control method, and program

    JP2019103067A

  • Object determination apparatus

    JP2021056968A

  • Image processing device, image processing method, and non-transitory computer-readable medium having image processing program stored thereon

    WO2021084677A1

  • Image processing device, image processing method, and program

    WO2021250808A1