Image processing device, image processing method, and program
The image processing device integrates feature amounts from multiple human bodies to address hidden key points, enhancing the accuracy of image search and classification for human body postures or movements.
Patent Information
- Application Number
- JP2023559384
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-11-15
- Publication Date
- 2025-08-20
- Estimated Expiration
- 2041-11-15
AI Technical Summary
Existing image processing techniques for searching and classifying human body postures or movements are inaccurate when parts of the body are hidden by other objects, as key points cannot be detected, leading to poor accuracy in image search and classification.
An image processing device and method that integrates feature amounts from detected key points across multiple human bodies to calculate integrated feature amounts, even when key points are not detected in one body by using feature amounts from other bodies, thereby complementing missing features.
Improves the accuracy of image search and classification for human body postures or movements by ensuring complete feature calculation for all key points, even when some are hidden, by integrating features across multiple images or video frames.
Smart Images

Figure 0007726290000001 
Figure 0007726290000002 
Figure 0007726290000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image processing device, an image processing method, and a program. [Background technology]
[0002] Technologies related to the present invention are disclosed in Patent Document 1 and Non-Patent Document 1. Patent Document 1 discloses a technology for calculating feature amounts for each of multiple key points of a human body included in an image, searching for images containing human bodies with similar postures or movements based on the calculated feature amounts, and classifying images with similar postures or movements together. Non-Patent Document 1 also discloses a technology related to human skeleton estimation. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] International Publication No. 2021 / 084677 [Non-patent literature]
[0004] [Non-Patent Document 1] Zhe Cao, Tomas Simon, Shih-En Wei, Yaser Sheikh, "Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields", The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, P. 7291-7299 Summary of the Invention [Problem to be solved by the invention]
[0005] When the search and classification disclosed in Patent Document 1 are performed using images in which parts of the human body are hidden by other objects or other parts of the person, the accuracy is poor. This problem can be alleviated by using images in which parts of the human body are not hidden and all key points can be detected. However, it may be difficult to prepare such images.
[0006] An object of the present invention is to improve the accuracy of a technique for searching for images including human bodies with similar postures or movements, and for classifying images including human bodies with similar postures or movements together. [Means for solving the problem]
[0007] According to the present invention, a skeletal structure detection means for performing processing to detect a plurality of key points corresponding to a plurality of parts of a human body included in an image; a feature calculation means for calculating a feature of each of the detected key points; a processing means for integrating the feature amounts detected from the plurality of human bodies for each of the body parts to calculate an integrated feature amount for each of the body parts, and performing image search or image classification based on the integrated feature amount; and The processing means An image processing device is provided that, when the key point corresponding to a first part of the plurality of parts is not detected from a part of the plurality of human bodies, and the key point corresponding to the first part is detected from another part of the plurality of human bodies, calculates the integrated feature of the first part based on the feature of the key point corresponding to the first part detected from the other part.
[0008] Further, according to the present invention, The computer a skeletal structure detection step of detecting a plurality of key points corresponding to a plurality of body parts included in the image; a feature calculation step of calculating a feature of each of the detected key points; a processing step of integrating the feature amounts detected from each of the plurality of human bodies for each of the body parts to calculate an integrated feature amount for each of the body parts, and performing image search or image classification based on the integrated feature amount; Run In the processing step, When the key point corresponding to a first part of the plurality of parts is not detected from one part of the plurality of human bodies, and the key point corresponding to the first part is detected from another part of the plurality of human bodies, an image processing method is provided for calculating the integrated feature of the first part based on the feature of the key point corresponding to the first part detected from the other part.
[0009] Further, according to the present invention, Computer, a skeletal structure detection means for detecting a plurality of key points corresponding to a plurality of parts of the human body included in the image; a feature calculation means for calculating a feature of each of the detected key points; a processing means for integrating the feature amounts detected from the plurality of human bodies for each of the body parts, calculating an integrated feature amount for each of the body parts, and performing image search or image classification based on the integrated feature amount; It functions as The processing means When the key point corresponding to a first part of the plurality of parts is not detected from a part of the plurality of human bodies, and the key point corresponding to the first part is detected from another part of the plurality of human bodies, a program is provided that calculates the integrated feature of the first part based on the feature of the key point corresponding to the first part detected from the other part. [Effects of the Invention]
[0010] According to the present invention, it is possible to improve the accuracy of techniques for searching for images containing human bodies with similar postures or movements, and for classifying images containing human bodies with similar postures or movements together. [Brief explanation of the drawings]
[0011] The above and other objects, features and advantages are described below. Suitable This will become more apparent from the following embodiments and the accompanying drawings.
[0012] [Figure 1] 10A and 10B are diagrams illustrating an example of a process for calculating an integrated feature amount from a still image according to the present embodiment. [Figure 2] FIG. 1 is a diagram illustrating an example of a hardware configuration of an image processing apparatus according to an embodiment of the present invention. [Figure 3] FIG. 1 is a diagram illustrating an example of a functional block diagram of an image processing apparatus according to an embodiment of the present invention. [Figure 4] 3A and 3B are diagrams illustrating an example of a skeletal structure of a human body model detected by the image processing apparatus of the present embodiment. [Figure 5] 3A and 3B are diagrams illustrating an example of a skeletal structure of a human body model detected by the image processing apparatus of the present embodiment. [Figure 6] 3A and 3B are diagrams illustrating an example of a skeletal structure of a human body model detected by the image processing apparatus of the present embodiment. [Figure 7] FIG. 10 is a diagram illustrating an example of feature amounts of key points calculated by the image processing apparatus of the present embodiment. [Figure 8] FIG. 10 is a diagram illustrating an example of feature amounts of key points calculated by the image processing apparatus of the present embodiment. [Figure 9] FIG. 10 is a diagram illustrating an example of feature amounts of key points calculated by the image processing apparatus of the present embodiment. [Figure 10] FIG. 10 is a diagram illustrating an example of a process for calculating an integrated feature from a moving image according to the present embodiment. [Figure 11] 10A and 10B are diagrams illustrating an example of processing for identifying a correspondence relationship between frame images according to the present embodiment. [Figure 12] FIG. 10 is a diagram illustrating an example of a process for calculating an integrated feature from a moving image according to the present embodiment. [Figure 13] 10 is a flowchart showing an example of a processing flow of the image processing apparatus of the present embodiment. [Figure 14]10 is a flowchart showing an example of a processing flow of the image processing apparatus of the present embodiment. [Figure 15] 10A and 10B are diagrams illustrating an example of a process for calculating an integrated feature amount from a still image according to the present embodiment. [Figure 16] 10A and 10B are diagrams illustrating an example of a process for calculating an integrated feature amount from a still image according to the present embodiment. [Figure 17] 10A and 10B are diagrams illustrating an example of a process for calculating an integrated feature amount from a still image according to the present embodiment. [Figure 18] 10A and 10B are diagrams illustrating an example of a process for calculating an integrated feature amount from a still image according to the present embodiment. [Figure 19] 10A and 10B are diagrams illustrating an example of a process for calculating an integrated feature from a moving image according to the present embodiment. [Figure 20] 10A and 10B are diagrams illustrating an example of a process for calculating an integrated feature from a moving image according to the present embodiment. [Figure 21] FIG. 1 is a diagram illustrating an example of a functional block diagram of an image processing apparatus according to an embodiment of the present invention. [Figure 22] FIG. 2 is a diagram schematically illustrating an example of information displayed by the image processing apparatus according to the present embodiment. [Figure 23] FIG. 2 is a diagram schematically illustrating an example of information displayed by the image processing apparatus according to the present embodiment. [Figure 24] 10 is a flowchart showing an example of a processing flow of the image processing apparatus of the present embodiment. [Figure 25] FIG. 1 is a diagram illustrating an example of a functional block diagram of an image processing apparatus according to an embodiment of the present invention. [Figure 26] FIG. 1 is a diagram illustrating an example of a functional block diagram of an image processing apparatus according to an embodiment of the present invention. [Figure 27] FIG. 2 is a diagram schematically illustrating an example of information displayed by the image processing apparatus according to the present embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0013] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings, like components are designated by like reference numerals, and the description thereof will be omitted as appropriate.
[0014] First Embodiment "overview" The image processing device of this embodiment detects key points corresponding to each part of a human body from each of multiple human bodies (hereinafter, "human body parts" may be simply referred to as "parts"), integrates the feature amounts of the key points for each part, and calculates an integrated feature amount for each part. The image processing device then performs image search and image classification based on the calculated integrated feature amount for each part. With this image processing device, if a certain key point is not detected from one human body, it can be complemented with the feature amount of that key point detected from another human body. Therefore, it is possible to calculate an integrated feature amount corresponding to each of all parts.
[0015] An example of the process for calculating integrated features will be described using Figure 1. The first still image shown in the figure is an image of a person washing their hands, captured from the left side of the person. In the first still image, part of the right side of the person's body is hidden and not visible. When processing to detect N key points on the human body is performed on such a first still image, some of the N key points, i.e., key points included in the unhidden parts, are detected, but the other part of the N key points, i.e., key points included in the hidden parts, are not detected. As a result, the features of some key points are missing.
[0016] Similarly, the second still image is an image of a person washing their hands, taken from the right side of the person. In the second still image, part of the left side of the person's body is hidden and not visible. When processing to detect N keypoints on the human body is performed on such a second still image, some of the N keypoints, i.e., keypoints in the unhidden parts, are detected, but the other part of the N keypoints, i.e., keypoints in the hidden parts, are not detected. As a result, the features of some keypoints are missing.
[0017] When the image processing device of this embodiment integrates the feature amounts of keypoints detected from the human body included in such a first still image with the feature amounts of keypoints detected from the human body included in the second still image, the feature amounts of keypoints not detected from the human body included in the first still image can be complemented with the feature amounts of keypoints detected from the human body included in the second still image. Similarly, the feature amounts of keypoints not detected from the human body included in the second still image can be complemented with the feature amounts of keypoints detected from the human body included in the first still image. As a result, integrated feature amounts corresponding to all N body parts can be calculated. Then, the integrated feature amounts corresponding to all N body parts can be used to search for images containing human bodies with similar postures or movements, or to classify images containing human bodies with similar postures or movements together, thereby improving accuracy.
[0018] "Hardware Configuration" Next, an example of the hardware configuration of an image processing device will be described. Each functional unit of the image processing device is realized by any combination of hardware and software, centered around a CPU (Central Processing Unit) of any computer, memory, programs loaded into the memory, a storage unit such as a hard disk that stores the programs (this can store programs that are pre-loaded when the device is shipped, as well as programs downloaded from storage media such as CDs (Compact Discs) or servers on the Internet), and a network connection interface. Those skilled in the art will understand that there are many variations in the implementation methods and devices.
[0019] FIG. 2 is a block diagram illustrating an example of the hardware configuration of an image processing device. As shown in FIG. 2, the image processing device has a processor 1A, a memory 2A, an input / output interface 3A, a peripheral circuit 4A, and a bus 5A. The peripheral circuit 4A includes various modules. The image processing device does not necessarily have to have the peripheral circuit 4A. Note that the image processing device may be composed of multiple devices that are physically and / or logically separated. In this case, each of the multiple devices can have the above hardware configuration.
[0020] The bus 5A is a data transmission path for the processor 1A, memory 2A, peripheral circuit 4A, and input / output interface 3A to transmit and receive data among them. The processor 1A is an arithmetic processing device such as a CPU or a GPU (Graphics Processing Unit). The memory 2A is a memory such as a RAM (Random Access Memory) or a ROM (Read Only Memory). The input / output interface 3A includes an interface for acquiring information from an input device, an external device, an external server, an external sensor, a camera, etc., and an interface for outputting information to an output device, an external device, an external server, etc. Examples of input devices include a keyboard, a mouse, a microphone, physical buttons, a touch panel, etc. Examples of output devices include a display, a speaker, a printer, a mailer, etc. The processor 1A can issue commands to each module and perform calculations based on the results of those calculations.
[0021] "Function Configuration" 3 shows an example of a functional block diagram of the image processing device 100 of this embodiment. The illustrated image processing device 100 includes a skeletal structure detection unit 101, a feature calculation unit 102, a processing unit 103, and a storage unit 104. Note that the image processing device 100 does not necessarily have to include the storage unit 104. In this case, an external device includes the storage unit 104. The storage unit 104 is configured to be accessible from the image processing device 100.
[0022] The skeletal structure detection unit 101 performs processing to detect N (N is an integer equal to or greater than 2) key points corresponding to each of multiple parts of the human body included in the image. The concept of image includes still images and videos. When a video is being processed, the skeletal structure detection unit 101 performs processing to detect key points for each frame image. This processing by the skeletal structure detection unit 101 is realized using the technology disclosed in Patent Document 1. Although details are omitted, the technology disclosed in Patent Document 1 detects the skeletal structure using a skeletal estimation technology such as OpenPose disclosed in Non-Patent Document 1. The skeletal structure detected by this technology is composed of "key points," which are characteristic points such as joints, and "bones (bone links)," which indicate the links between key points.
[0023] Fig. 4 shows the skeletal structure of a human body model 300 detected by the skeletal structure detection unit 101, and Fig. 5 and Fig. 6 show examples of detected skeletal structures. The skeletal structure detection unit 101 detects the skeletal structure of a human body model (two-dimensional skeletal model) 300 as shown in Fig. 4 from a two-dimensional image using a skeletal estimation technique such as OpenPose. The human body model 300 is a two-dimensional model made up of key points such as a person's joints and bones connecting each key point.
[0024] The skeletal structure detection unit 101, for example, extracts feature points that can be key points from an image, and detects N key points of the human body by referring to information obtained by machine learning of the image of the key points. The N key points to be detected are determined in advance. The number of key points to be detected (i.e., the number N) and which parts of the human body are to be detected as key points vary, and any variation can be adopted.
[0025] In the following, as shown in Figure 4, the head A1, neck A2, right shoulder A31, left shoulder A32, right elbow A41, left elbow A42, right hand A51, left hand A52, right waist A61, left waist A62, right knee A71, left knee A72, right foot A81, and left foot A82 are defined as N key points (N=14) to be detected. In the human body model 300 shown in FIG. 4, the following bones are further defined as the bones of the person connecting these key points: bone B1 connecting the head A1 and neck A2; bone B21 and bone B22 connecting the neck A2 and the right shoulder A31 and left shoulder A32, respectively; bone B31 and bone B32 connecting the right shoulder A31 and left shoulder A32 and the right elbow A41 and left elbow A42, respectively; bone B41 and bone B42 connecting the right elbow A41 and left elbow A42 and the right hand A51 and left hand A52, respectively; bone B51 and bone B52 connecting the neck A2 and the right hip A61 and left hip A62, respectively; bone B61 and bone B62 connecting the right hip A61 and left hip A62 and the right knee A71 and left knee A72, respectively; and bone B71 and bone B72 connecting the right knee A71 and left knee A72 and the right foot A81 and left foot A82, respectively.
[0026] FIG. 5 is an example of key points detected from a human body standing upright. In FIG. 5, the upright human body is imaged from the front, and all 14 key points are detected. FIG. 6 is an example of key points detected from a human body crouching. In FIG. 6, the crouching human body is imaged from the right side, and only some of the 14 key points are detected. Specifically, in FIG. 6, the head A1, neck A2, right shoulder A31, right elbow A41, right hand A51, right hip A61, right knee A71, and right foot A81 are detected, but the left shoulder A32, left elbow A42, left hand A52, left hip A62, left knee A72, and left foot A82 are not detected.
[0027] 3, the feature amount calculation unit 102 calculates the feature amount of the detected two-dimensional skeletal structure. For example, the feature amount calculation unit 102 calculates the feature amount of each of the detected key points.
[0028] Skeletal structure features indicate the characteristics of a person's skeleton and are used to classify and search a person's state (posture and movement) based on the person's skeleton. Typically, these features include multiple parameters. The features may be the features of the entire skeletal structure, the features of a portion of the skeletal structure, or multiple features for each part of the skeletal structure. The feature calculation method may be any method, such as machine learning or normalization, and normalization may involve finding a minimum or maximum value. Examples of feature values include features obtained by machine learning of the skeletal structure, the size of the skeletal structure on an image from the head to the feet, the relative positions of multiple key points in the vertical direction of a skeletal region containing the skeletal structure on an image, and the relative positions of multiple key points in the horizontal direction of the skeletal region. The size of the skeletal structure refers to the vertical height or area of the skeletal region containing the skeletal structure on an image. The vertical direction (height direction or vertical direction) refers to the up-down direction (Y-axis direction) in the image, for example, the direction perpendicular to the ground (reference plane). The left-right direction (horizontal direction) is the left-right direction in the image (X-axis direction), and is, for example, a direction parallel to the ground.
[0029] In order to perform the classification and search desired by the user, it is preferable to use features that are robust to the classification and search process. For example, if the user desires classification and search that are not dependent on the orientation or body shape of a person, features that are robust to the orientation and body shape of a person may be used. By learning the skeletons of people facing in various directions in the same posture or the skeletons of people with various body shapes in the same posture, or by extracting features only in the up-down direction of the skeleton, it is possible to obtain features that are not dependent on the orientation or body shape of a person.
[0030] The above processing by the feature amount calculation unit 102 is realized using the technology disclosed in Patent Document 1.
[0031] 7 shows an example of the feature amounts of each of a plurality of key points calculated by the feature amount calculation unit 102. Note that the feature amounts of the key points illustrated here are merely examples, and are not limited to these.
[0032] In this example, the feature values of keypoints indicate the relative positional relationships of multiple keypoints in the vertical direction of the skeletal region containing the skeletal structure on the image. Because the neck keypoint A2 is used as the reference point, the feature value of keypoint A2 is 0.0. The feature values of keypoint A31 on the right shoulder and keypoint A32 on the left shoulder, which are at the same height as the neck, are also 0.0. The feature value of keypoint A1 on the head, which is higher than the neck, is -0.2. The feature values of keypoints A51 on the right hand and A52 on the left hand, which are lower than the neck, are 0.4, and the feature values of keypoints A81 on the right foot and A82 on the left foot are 0.9. If the person raises their left hand from this position, as shown in Figure 8, the left hand will be higher than the reference point, and the feature value of keypoint A52 on the left hand will be -0.4. However, because normalization is performed using only the Y-axis coordinate, the feature values do not change even if the width of the skeletal structure changes, as shown in Figure 9, compared to Figure 7. That is, the feature amount (normalized value) in this example indicates the feature in the height direction (Y direction) of the skeletal structure (keypoint), and is not affected by changes in the lateral direction (X direction) of the skeletal structure.
[0033] Returning to FIG. 3, the processing unit 103 integrates the features of the key points detected from each of M (M is an integer equal to or greater than 2) human bodies for each body part to calculate an integrated feature for each body part. Then, the processing unit 103 performs image search or image classification based on the integrated feature for each body part. As described above, multiple key points correspond to multiple body parts, respectively. Therefore, performing processing "for each body part" is the same as performing processing "for each key point." For example, the "integrated feature for each body part" obtained by calculating for each body part is the same as the "integrated feature for each of N key points" obtained by calculating for each key point.
[0034] -Processing to calculate integrated features- When processing still images First, the user specifies M human bodies to be subjected to the process of calculating integrated features. For example, the user may specify M human bodies by specifying M still images each including one human body (specifying M still image files). The specification of M still images may be, for example, an operation of inputting M still images to the image processing device 100, or an operation of selecting M still images from a plurality of still images stored in the image processing device 100. In this case, the skeletal structure detection unit 101 described above performs a process of detecting N key points for each of the specified M still images. Note that all N key points may be detected in some cases, or only some of the N key points may be detected in other cases. The feature calculation unit 102 calculates the feature amount of each of the detected key points.
[0035] Alternatively, the user may specify M human bodies by specifying at least one still image (specifying at least one still image file) and M regions each containing one human body within the specified at least one still image. Note that multiple regions (i.e., multiple human bodies) may be specified within one still image. The process of specifying a portion of a region within a still image can be realized using any conventional technology. In this case, the skeletal structure detection unit 101 described above performs a process of detecting N key points for each of the specified M regions. Note that all N key points may be detected in some cases, and only some of the N key points may be detected in other cases. The feature calculation unit 102 calculates the feature amount of each detected key point.
[0036] After calculating the feature amounts of each of the M keypoints of the human body specified by the user, the processing unit 103 integrates them for each keypoint to calculate an integrated feature amount. The processing unit 103 sequentially selects, for example, one of the N keypoints and performs a process of calculating the integrated feature amount. Hereinafter, one of the N keypoints that is selected as the processing target will be referred to as the "first keypoint."
[0037] When the first keypoint is not detected from one part of the M human bodies but the first keypoint is detected from another part of the M human bodies, the processing unit 103 calculates an integrated feature of the first keypoint (synonymous with "integrated feature of the first part") based on the feature of the first keypoint detected from the other part. This processing makes it possible to integrate the feature of the keypoints calculated from each of the multiple human bodies by complementing each other's missing parts.
[0038] The detection state of the first keypoint is either (1) detected from only one of the M human bodies, (2) detected from multiple of the M human bodies, or (3) not detected from any of the M human bodies. The processing unit 103 can calculate the integrated feature by processing according to each detection state. This will be explained in detail below.
[0039] (1) Detection from only one of M human bodies If the first keypoint is detected from only one of the M human bodies, the processing unit 103 sets the feature of the first keypoint detected from that one human body as the integrated feature of the first keypoint.
[0040] (2) Detection from multiple bodies among M bodies When the first keypoints are detected from a plurality of the M human bodies, the processing unit 103 calculates the integrated feature amount of the first keypoints by one of the following calculation examples 1 to 4.
[0041] Calculation example 1 When the first keypoints are detected from multiple of the M human bodies, the processing unit 103 calculates the statistical value of the feature amounts of the first keypoints detected from the multiple human bodies as the integrated feature amount of the first keypoints. The statistical value is the average value, median value, mode value, maximum value, or minimum value.
[0042] Calculation example 2 When the first keypoint is detected from multiple of the M human bodies, the processing unit 103 determines the feature with the highest confidence among the feature of the first keypoint detected from the multiple human bodies as the integrated feature of the first keypoint. There are no particular limitations on the method for calculating the confidence. For example, in a skeleton estimation technology such as OpenPose, a score associated with each detected keypoint and output may be used as the confidence of each keypoint.
[0043] Calculation example 3 When the first keypoints are detected from multiple of the M human bodies, the processing unit 103 calculates a weighted average value of the feature amounts of the first keypoints detected from each of the multiple human bodies according to the confidence levels of the feature amounts of the first keypoints as the integrated feature amount of the first keypoints. There are no particular limitations on the method for calculating the confidence level. For example, in a skeleton estimation technology such as OpenPose, a score associated with each detected keypoint and output may be used as the confidence level of each keypoint.
[0044] Calculation example 4 The user specifies in advance the priority of each of the specified M human bodies. The specified content is input to the image processing device 100. Then, when the first keypoint is detected from more than one of the M human bodies, the processing unit 103 sets the feature amount of the first keypoint detected from the human body with the highest priority among the multiple human bodies in which the first keypoint is detected as the integrated feature amount of the first keypoint.
[0045] (3) Not detected in any of the M bodies If the first keypoint is not detected from any of the M human bodies, the processing unit 103 does not calculate the integrated feature of the first keypoint.
[0046] When processing video First, the user specifies M human bodies to be subjected to the process of calculating integrated features. For example, the user may specify M human bodies by specifying M videos each including one human body (specifying M video files). The specification of M videos may be, for example, an operation of inputting M videos to the image processing device 100, or an operation of selecting M videos from a plurality of videos stored in the image processing device 100. In this case, the skeletal structure detection unit 101 described above performs a process of detecting N key points for frame images of each of the specified M videos. Note that all N key points may be detected in some cases, and only some of the N key points may be detected in other cases. The feature calculation unit 102 calculates the feature amount of each of the detected key points.
[0047] Alternatively, the user may specify M human bodies by specifying at least one video (specifying at least one video file) and M scenes (a portion of a video, or a scene composed of a portion of frame images among a plurality of frame images included in a video) or M regions within the specified at least one video, each containing one human body. Note that multiple scenes or multiple regions (i.e., multiple human bodies) may be specified within a single video. The process of specifying a portion of a scene or a portion of a region within a video can be realized using any conventional technology. In this case, the skeletal structure detection unit 101 described above performs a process of detecting N key points for the frame images of each of the specified M scenes (or a portion of a frame image specified by the user). Note that all N key points may be detected, or only some of the N key points may be detected. The feature calculation unit 102 calculates the feature values of each of the detected key points.
[0048] After calculating the feature amounts of the key points of each of the M human bodies designated by the user, the processing unit 103 integrates them for each key point to calculate an integrated feature amount. The processing unit 103 identifies the correspondence between frame images in the M videos or M scenes, and integrates the feature amounts of the key points detected from each of the corresponding frame images for each key point. This will be described in more detail below with reference to FIGS. 10 to 12.
[0049] 10 shows two (M=2) moving images (scenes), each of which includes one human body and multiple frame images.
[0050] As shown in FIG. 11, the processing unit 103 associates frame images of a human body performing a predetermined movement in the first video with frame images of a human body performing a predetermined movement in the second video in a similar pose. In FIG. 11, corresponding frame images are connected by lines. As shown in the figure, one frame image of the first video may be associated with multiple frame images of the second video. Also, one frame image of the second video may be associated with multiple frame images of the first video. The identification of the correspondence can be achieved, for example, using a technique such as DTW (Dynamic Time Warping). In this case, the distance between feature quantities (Manhattan distance or Euclidean distance) can be used as the distance score required for identifying the correspondence. With this technique, the correspondence can be identified even when the first video and the second video have different durations (i.e., the number of frame images differs), as shown in FIG. 10.
[0051] In this case, as shown in Fig. 12, by calculating the feature quantities of N key points for each combination of corresponding multiple frame images, time series data of the integrated feature quantities of N key points can be obtained. 11 +F 21 is the frame image F of the first moving image in FIG. 11 The feature values of the detected human body key points and the frame image F of the second video are 21The means for integrating the feature values of the human body keypoints detected from the corresponding frame images is the same as the means for integrating the feature values of the human body keypoints detected from the still image described above.
[0052] -Image search processing- In the image search process, the processing unit 103 uses the integrated feature calculated based on the M number of human bodies specified by the user as described above as a query to search for still images including human bodies in a posture similar to that indicated by the integrated feature, videos including human bodies performing movements similar to those indicated by the time-series data of the integrated feature, etc. The search method can be realized by using the technology disclosed in Patent Document 1.
[0053] -Image classification processing- In the image classification process, the processing unit 103 treats the postures and movements indicated by the integrated feature values calculated based on the M number of human bodies designated by the user as one target for classification, and classifies images with similar postures and movements together. The classification method can be realized by using the technology disclosed in Patent Document 1.
[0054] -Other processing- The processing unit 103 may register the postures and movements indicated by the integrated features calculated based on the M number of human bodies specified by the user as described above in a database (storage unit 104) as a single processing target. The multiple postures and movements registered in the database may be, for example, targets to be matched with a query in the image search process described above, or may be targets for classification in the image classification process described above. For example, by photographing the same person from multiple angles with multiple cameras and specifying multiple human bodies of the same person included in the multiple images photographed by the multiple cameras as the M number of human bodies, integrated features that clearly represent the postures and movements of the bodies are calculated and registered in the database.
[0055] Next, an example of the processing flow of the image processing device 100 will be described with reference to the flowchart of FIG.
[0056] First, the image processing device 100 acquires at least one image (S10). Next, the image processing device 100 performs a process of detecting N key points from each of M human bodies included in at least one acquired image (S11). In some cases, all N key points are detected from each human body, and in other cases, only some of the N key points are detected.
[0057] Next, the image processing device 100 calculates the feature amounts of the detected keypoints for each human body (S12). Next, the image processing device 100 integrates the feature amounts of the keypoints detected from each of the M human bodies to calculate an integrated feature amount for each of the N keypoints (S13). Next, the image processing device 100 performs image search or image classification based on the integrated feature amount calculated in S13 (S14).
[0058] An example of the process of S13 will now be described in detail with reference to the flowchart of FIG.
[0059] The image processing device 100 selects one of the N keypoints as a processing target (S20). Hereinafter, the selected keypoint will be referred to as the first keypoint.
[0060] Thereafter, the image processing device 100 performs processing according to the number of human bodies from which the first keypoints have been detected. If the first keypoints have been detected from only one of the M human bodies ("one" in S21), the image processing device 100 outputs the feature amount of the first keypoints detected from that one human body as the integrated feature amount of the first keypoints (S23).
[0061] If the first keypoints are detected from multiple of the M human bodies ("multiple" in S21), the image processing device 100 outputs a value calculated by a calculation process based on the feature amounts of the first keypoints detected from the multiple human bodies as the integrated feature amount of the first keypoints (S24). The details of the calculation process are as described above.
[0062] If the first keypoint is not detected from any of the M human bodies ("0" in S21), the processing unit 103 does not calculate the integrated feature of the first keypoint, and outputs a message indicating that there is no combined feature (S22).
[0063] "Action and effect" In an image, a part of a human body may be hidden by another object or other part of the human body and therefore not be visible. When such an image is processed using the technology disclosed in Patent Document 1, the key points of the hidden part are not detected, and their feature values are not calculated. If search / classification is performed based only on the feature values of some of the detected key points, images containing human bodies with similar postures or movements of at least a part of the body may be searched, or images with similar postures or movements of at least a part of the body may be classified together. As a result, the accuracy of search and classification may decrease.
[0064] The image processing device 100 of this embodiment integrates the feature amounts of the key points detected from each of a plurality of human bodies, and calculates the integrated feature amount of each of the plurality of key points. 100 The image processing device performs image search and image classification based on the calculated integrated feature amount. 100 According to this method, the feature values of key points not detected from one human body can be complemented by the feature values of key points detected from other bodies. This makes it possible to calculate integrated feature values corresponding to each of all key points. Then, image search and classification can be performed based on the integrated feature values corresponding to each of all key points, thereby improving the accuracy.
[0065] In this embodiment, for example, N keypoints of multiple human bodies P as shown in FIGS. 15 and 16 can be integrated. The still image in FIG. 15 is an image of a person washing their hands taken from the left side of the person. In the first still image, the left side of the person's body is visible, but the right side of the body is hidden and not visible. As a result, keypoints included in the left part of the person's body are detected, but keypoints included in the right part are not detected. The still image in FIG. 16 is an image of a person washing their hands taken from the right side of the person. In the second still image, the right side of the person's body is visible, but the left side is hidden and not visible. As a result, keypoints included in the right part of the person's body are detected, but keypoints included in the left part are not detected. By integrating the feature amounts of the human body keypoints detected from these two still images, the missing parts of each are complemented, and integrated feature amounts corresponding to all N keypoints can be calculated.
[0066] Furthermore, in this embodiment, for example, N key points of multiple human bodies P as shown in FIGS. 17 and 18 can be integrated. The still image in FIG. 17 is an image of a person standing with their left hand on their hip, photographed from the front of the person. In the first still image, no hidden parts of the person's body are present. As a result, all N key points are detected from the human body P. The still image in FIG. 18 is an image of a person standing with their right hand raised, photographed from the front of the person. In the second still image, part of the left half of the person's body is hidden by a vehicle Q. As a result, key points included in the unhidden parts of the person's body are detected, but key points included in the hidden parts are not detected. By integrating the feature amounts of the human body key points detected from these two still images, the missing parts in the second still image are complemented by the first still image, and integrated feature amounts corresponding to all N key points can be calculated. In this example, for example, the method of Example 4 described above, i.e., calculation of integrated feature amounts based on the priority of each of the M human bodies, may be performed. For example, the user may assign a higher priority to the human body included in the second still image than to the human body included in the first still image. In this case, the feature of the portion that appears in both the first and second still images is the portion that appears in the second still image. As a result, the calculated N integrated features represent a posture in which the person is standing with their left hand on their hip, as in the first still image, and their right hand raised, as in the second still image.
[0067] In this embodiment, for example, N key points of a plurality of human bodies P as shown in Fig. 19 and Fig. 20 can be integrated. The video in Fig. 19 is an image of a person standing and raising his right hand, photographed from the front. 1In the video (1), a portion of the left half of the person is obscured by vehicle Q. As a result, key points in the unobscured portions of the person's body are detected, but key points in the obscured portions are not. The video in FIG. 20 is an image of a person standing with their hands on their hips, captured from the front. In the second video, no obscured portions of the person's body are detected. As a result, all N key points are detected from the person P. By integrating the features of the key points of the human bodies detected from these two videos, the missing portions in the first video can be supplemented with the second video, and integrated features corresponding to all N key points can be calculated. In this example, for example, the method described in Example 4 above, i.e., calculation of integrated features based on the priority of each of the M human bodies, may be performed. For example, the user may assign a higher priority to the human bodies included in the first video than to the human bodies included in the second video. In this case, the features of the portions appearing in both the first video and the second video are adopted from the portions appearing in the first video. In this case, the time-series data of the calculated N integrated features will show the movement of the person placing their left hand on their hip as in the second video and raising their right hand while standing as in the first video.
[0068] It should be noted that the M human bodies may belong to the same person or may belong to different people.
[0069] <Second embodiment> The image processing device 100 of this embodiment differs from the first embodiment in the details of the process of integrating key points detected from each of M human bodies to calculate an integrated feature. In the first embodiment, the integrated feature is calculated using a flow such as that shown in FIG. 14 . In this embodiment, the image processing device 100 integrates key points detected from each of M human bodies to calculate an integrated feature using a method specified by user input. This will be described in detail below.
[0070] 21 shows an example of a functional block diagram of the image processing device 100 of this embodiment. The illustrated image processing device 100 includes a skeletal structure detection unit 101, a feature calculation unit 102, a processing unit 103, a storage unit 104, and an input unit 106. Note that the image processing device 100 does not necessarily have to include the storage unit 104. In this case, an external device includes the storage unit 104. The storage unit 104 is configured to be accessible from the image processing device 100.
[0071] The input unit 106 receives a user input specifying a method for integrating feature quantities of key points detected from each of the M human bodies. The input unit 106 can receive the user input via any input device, such as a touch panel, a keyboard, a mouse, a physical button, a microphone, or a gesture input device.
[0072] The processing unit 103 integrates the feature amounts detected from each of the M human bodies for each keypoint using a method designated by a user input, and calculates an integrated feature amount for each of the N keypoints.
[0073] The input unit 106 and the processing unit 103 can execute either of the following processing examples 1 and 2.
[0074] - Processing example 1 - In this example, the input unit 106 inputs a keypoint from which a feature is to be adopted for each of the M human bodies. This is equivalent to an input specifying, for each keypoint, from which human body the feature of the keypoint detected should be adopted. Then, the processing unit 103 determines the feature of the first keypoint detected from the human body specified by the user input as the integrated feature of the first keypoint.
[0075] There are various means for accepting the user input. For example, as shown in Fig. 22, the input unit 106 may display a human body model in which N objects R corresponding to N key points are arranged at corresponding skeletal positions of the human body, and accept a user input for selecting an object corresponding to a key point for which the calculated feature amount is to be adopted or an object corresponding to a key point for which the calculated feature amount is not to be adopted, for each of the M human bodies.
[0076] In addition, the input unit 106 can also be used to input the head, neck, right shoulder, The names of body parts corresponding to each of a plurality of key points, such as left shoulder, right elbow, left elbow, right hand, left hand, right hip, left hip, right knee, left knee, right foot, left foot, etc., may be displayed, and a user input for selecting key points for which calculated feature quantities are to be adopted or not adopted may be accepted for each of the M human bodies. In this case, a UI (user interface) component such as a check box may be used.
[0077] Alternatively, the input unit 106 may display a human body model in which N objects R corresponding to N key points are arranged at corresponding skeletal positions of the human body, as shown in Fig. 23, and accept a user input for selecting at least a portion of the body in the human body model. Then, the input unit 106 may determine the key points present in the body part selected by the user input as key points for which calculated feature amounts are adopted or key points for which calculated feature amounts are not adopted. In the example shown in Fig. 23, at least a portion of the body is selected by a frame W. The user changes the position and size of the frame W and adjusts it so that the desired key points are included within the frame W.
[0078] Alternatively, the input unit 106 may display names of body parts such as upper body, lower body, right body, and left body, and accept a user input for selecting at least one of them. Then, the input unit 106 may determine key points present in the body part selected by the user input as key points for which calculated feature quantities are adopted or key points for which calculated feature quantities are not adopted. In this case, a UI (user interface) component such as a check box may be used.
[0079] - Processing example 2 - In this example, the input unit 106 accepts a user input specifying a weight for the feature amount calculated from each of the M human bodies for each keypoint for each of the M human bodies. Then, the processing unit 103 calculates a weighted average value of the feature amounts calculated from each of the M human bodies according to the weight specified by the user as the integrated feature amount for each keypoint.
[0080] There are various methods for specifying a weight for each key point. For example, the input unit 106 may receive an input for individually specifying a key point using the method described in processing example 1, and then further receive an input for specifying a weight for the specified key point. Alternatively, the input unit 106 may receive an input for specifying a body part using the method described in processing example 1, and then further receive an input for specifying a weight common to all key points included in the specified body part.
[0081] Next, an example of the processing flow of the image processing device 100 will be described using the flowchart in Fig. 24. Note that the processing order of each step can be changed as appropriate.
[0082] First, the image processing device 100 acquires at least one image (S30). Next, the image processing device 100 receives a user input specifying a method for integrating feature quantities of key points detected from M (M is an integer equal to or greater than 2) human bodies (S31).
[0083] Next, the image processing device 100 performs a process of detecting N key points from each of M human bodies included in at least one acquired image (S32). In some cases, all N key points are detected from each human body, and in other cases, only some of the N key points are detected.
[0084] Next, the image processing device 100 calculates the feature amounts of the detected keypoints for each human body (S33). Next, the image processing device 100 integrates the feature amounts of the keypoints detected from each of the M human bodies using the method specified in S31 to calculate an integrated feature amount for each of the N keypoints (S34). Next, the image processing device 100 performs image search or image classification based on the integrated feature amount calculated in S34 (S35).
[0085] Other configurations of the image processing device 100 of this embodiment are the same as those of the first embodiment.
[0086] The image processing device 100 of this embodiment achieves the same effects as those of the first embodiment. Furthermore, since the user can specify the integration method, it becomes possible to calculate an integrated feature amount that the user desires.
[0087] <Third embodiment> The image processing apparatus 100 of this embodiment has a function of outputting information for distinguishing between key points for which integrated features have been calculated and key points for which integrated features have not been calculated, as will be described in detail below.
[0088] 25 shows an example of a functional block diagram of the image processing device 100 of this embodiment. The image processing device 100 shown in the figure includes a skeletal structure detection unit 101, a feature calculation unit 102, a processing unit 103, a storage unit 104, and a display unit 105.
[0089] 26 shows another example of a functional block diagram of the image processing device 100 of this embodiment. The image processing device 100 shown in the figure includes a skeletal structure detection unit 101, a feature calculation unit 102, a processing unit 103, a storage unit 104, a display unit 105, and an input unit 106.
[0090] It should be noted that the image processing device 100 does not necessarily have to have the storage unit 104. In this case, an external device includes the storage unit 104. The storage unit 104 is configured to be accessible from the image processing device 100.
[0091] The display unit 105 displays information that distinguishes between keypoints that are not detected from any of the M human bodies specified by the user and for which no integrated feature has been calculated, and keypoints that are detected from at least one of the M human bodies and for which an integrated feature has been calculated.
[0092] For example, the display unit 105 may display a human body model in which N objects R corresponding to N key points are arranged at corresponding skeletal positions of the human body, as shown in Fig. 27, and may distinguishably display objects corresponding to key points for which integrated features have not been calculated and objects corresponding to key points detected from at least one of the M human bodies and for which integrated features have been calculated. A method for distinguishably displaying the objects may be, but is not limited to, filling in the objects as shown in Fig. 27. Other examples of the method include, for example, differentiating the colors of the objects, differentiating the shapes of the objects, or highlighting objects corresponding to key points for which integrated features have been calculated or key points for which integrated features have not been calculated by blinking or the like.
[0093] The display unit 105 may further display information that identifies key points detected from each of the M human bodies specified by the user and key points that were not detected. That is, the display unit 105 may further display information that identifies parts where key points were detected and parts where key points were not detected. This display can be achieved by a method similar to the method described using FIG. 27.
[0094] Other configurations of the image processing device 100 of this embodiment are the same as those of the first and second embodiments.
[0095] The image processing device 100 of this embodiment achieves the same effects as those of the first and second embodiments. Furthermore, the image processing device 100 of this embodiment allows the user to easily understand which of the N key points are covered by the specified M human bodies, based on the information displayed on the display unit 105. Furthermore, by using an image such as that shown in FIG. 27, the user can intuitively understand the above content. As a result, the user can understand what kind of human body should be added to generate integrated features of all N key points.
[0096] Although the embodiments of the present invention have been described above with reference to the drawings, these are merely examples of the present invention, and various other configurations may be adopted. The configurations of the above-described embodiments may be combined with each other, or some of the configurations may be replaced with other configurations. Furthermore, various modifications may be made to the configurations of the above-described embodiments without departing from the spirit of the invention. Furthermore, the configurations and processes disclosed in the above-described embodiments and modified examples may be combined with each other.
[0097] In addition, in the flowcharts used in the above description, multiple steps (processes) are described in order, but the order of execution of the steps performed in each embodiment is not limited to the order described. In each embodiment, the order of the steps shown in the drawings can be changed to the extent that the content is not affected. Furthermore, the above-mentioned embodiments can be combined to the extent that the content is not contradictory.
[0098] A part or all of the above-described embodiments can be described as, but not limited to, the following supplementary notes. 1. A skeletal structure detection means for detecting a plurality of key points corresponding to a plurality of parts of the human body included in the image; a feature calculation means for calculating a feature of each of the detected key points; a processing means for integrating the feature amounts detected from the plurality of human bodies for each of the body parts to calculate an integrated feature amount for each of the body parts, and performing image search or image classification based on the integrated feature amount; and The processing means When the key point corresponding to a first part of the plurality of parts is not detected from a part of the plurality of human bodies, and the key point corresponding to the first part is detected from another part of the plurality of human bodies, an image processing device calculates the integrated feature of the first part based on the feature of the key point corresponding to the first part detected from the other part. 2. The processing means 1. An image processing device according to claim 1, wherein, when the key point corresponding to the first part is detected from one of the multiple human bodies, the feature of the key point corresponding to the first part detected from the one human body is used as the integrated feature of the first part. 3. The processing means 3. The image processing device according to claim 1 or 2, wherein, when the keypoints corresponding to the first region are detected from multiple of the multiple human bodies, the statistical value of the feature amounts of the keypoints corresponding to the first region detected from the multiple human bodies is used as the integrated feature amount of the first region. 4. The processing means 3. The image processing device according to claim 1 or 2, wherein, when the keypoints corresponding to the first region are detected from multiple of the multiple human bodies, the feature with the highest degree of certainty among the feature amounts of the keypoints corresponding to the first region detected from the multiple human bodies is used as the integrated feature amount of the first region. 5. The processing means 3. The image processing device according to claim 1 or 2, wherein, when the keypoints corresponding to the first region are detected from multiple of the multiple human bodies, a weighted average value of the feature amounts of the keypoints corresponding to the first region, which is determined according to the confidence level of the feature amounts of the keypoints corresponding to the first region detected from each of the multiple human bodies, is used as the integrated feature amount of the first region. 6. An image processing device described in any one of 1 to 5, further comprising a display means for displaying information that identifies the body part for which the key point is not detected from any of the plurality of human bodies and the integrated feature is not calculated, and the body part for which the key point is detected from at least one of the plurality of human bodies and the integrated feature is calculated. 7. The display means 7. An image processing device according to claim 6, which displays a human body model in which a plurality of objects are arranged at the parts of the human body, and displays the objects corresponding to the parts for which the integrated features have been calculated and the objects corresponding to the parts for which the integrated features have not been calculated in a manner that allows them to be distinguished from each other. 8. The display means 8. The image processing device according to claim 6 or 7, further displaying information associated with each of the plurality of human bodies to identify the areas where the key points are detected and the areas where the key points are not detected. 9. The computer a skeletal structure detection step of detecting a plurality of key points corresponding to a plurality of body parts included in the image; a feature calculation step of calculating a feature of each of the detected key points; a processing step of integrating the feature amounts detected from each of the plurality of human bodies for each of the body parts to calculate an integrated feature amount for each of the body parts, and performing image search or image classification based on the integrated feature amount; Run In the processing step, An image processing method for calculating, when the keypoint corresponding to a first part of the plurality of parts is not detected from a part of the plurality of human bodies, and the keypoint corresponding to the first part is detected from another part of the plurality of human bodies, the integrated feature of the first part is calculated based on the feature of the keypoint corresponding to the first part detected from the other part. 10. Computer a skeletal structure detection means for detecting a plurality of key points corresponding to a plurality of parts of the human body included in the image; a feature calculation means for calculating a feature of each of the detected key points; a processing means for integrating the feature amounts detected from the plurality of human bodies for each of the body parts, calculating an integrated feature amount for each of the body parts, and performing image search or image classification based on the integrated feature amount; It functions as The processing means A program that, when the keypoint corresponding to a first part of the plurality of parts is not detected from a part of the plurality of human bodies, and the keypoint corresponding to the first part is detected from another part of the plurality of human bodies, calculates the integrated feature of the first part based on the feature of the keypoint corresponding to the first part detected from the other part. [Explanation of symbols]
[0099] 100 Image processing device 101 Skeletal structure detection unit 102 Feature calculation unit 103 Processing section 104 Storage section 105 Display section 106 Input section 1A processor 2A Memory 3A input / output I / F 4A peripheral circuit 5A Bus
Claims
1. a skeletal structure detection means for performing processing to detect a plurality of key points corresponding to a plurality of parts of a human body included in an image; a feature calculation means for calculating a feature of each of the detected key points; a processing means for integrating the feature amounts detected from the plurality of human bodies for each of the body parts to calculate an integrated feature amount for each of the body parts, and performing image search or image classification based on the integrated feature amount; and The processing means When the key point corresponding to a first part of the plurality of parts is not detected from a part of the plurality of human bodies, and the key point corresponding to the first part is detected from another part of the plurality of human bodies, an image processing device calculates the integrated feature of the first part based on the feature of the key point corresponding to the first part detected from the other part.
2. The processing means 2. The image processing device according to claim 1, wherein, when the key point corresponding to the first region is detected from one of the plurality of human bodies, the feature of the key point corresponding to the first region detected from the one human body is used as the integrated feature of the first region.
3. The processing means 3. The image processing device according to claim 1, wherein when the key points corresponding to the first region are detected from a plurality of the plurality of human bodies, a statistical value of the feature amounts of the key points corresponding to the first region detected from the plurality of human bodies is used as the integrated feature amount of the first region.
4. The processing means 3. The image processing device according to claim 1, wherein, when the key points corresponding to the first region are detected from a plurality of the plurality of human bodies, the feature with the highest degree of certainty among the feature amounts of the key points corresponding to the first region detected from the plurality of human bodies is set as the integrated feature amount of the first region.
5. The processing means 3. The image processing device according to claim 1, wherein, when the key points corresponding to the first region are detected from a plurality of the plurality of human bodies, a weighted average value of the feature amounts of the key points corresponding to the first region, which is determined according to a certainty of the feature amounts of the key points corresponding to the first region detected from each of the plurality of human bodies, is used as the integrated feature amount of the first region.
6. 6. The image processing device according to claim 1, further comprising: a display means for displaying information that identifies a region for which the key point has not been detected from any of the plurality of human bodies and the integrated feature has not been calculated, and a region for which the key point has been detected from at least one of the plurality of human bodies and the integrated feature has been calculated.
7. The display means 7. The image processing device according to claim 6, wherein a human body model is displayed in which a plurality of objects are arranged at the parts of the human body, and the objects corresponding to the parts for which the integrated feature has been calculated and the objects corresponding to the parts for which the integrated feature has not been calculated are displayed in a manner that allows them to be distinguished from each other.
8. The display means The image processing device according to claim 6 or 7, further displaying information associated with each of the plurality of human bodies, the information identifying the region where the key point is detected and the region where the key point is not detected.
9. The computer a skeletal structure detection step of detecting a plurality of key points corresponding to a plurality of body parts included in the image; a feature calculation step of calculating a feature of each of the detected key points; a processing step of integrating the feature amounts detected from each of the plurality of human bodies for each of the body parts to calculate an integrated feature amount for each of the body parts, and performing image search or image classification based on the integrated feature amount; Run In the processing step, An image processing method for calculating, when the key point corresponding to a first part of the plurality of parts is not detected from a part of the plurality of human bodies, and the key point corresponding to the first part is detected from another part of the plurality of human bodies, the integrated feature of the first part is calculated based on the feature of the key point corresponding to the first part detected from the other part.
10. Computer, a skeletal structure detection means for detecting a plurality of key points corresponding to a plurality of parts of the human body included in the image; a feature calculation means for calculating a feature of each of the detected key points; a processing means for integrating the feature amounts detected from the plurality of human bodies for each of the body parts, calculating an integrated feature amount for each of the body parts, and performing image search or image classification based on the integrated feature amount; It functions as The processing means A program that, when the key point corresponding to a first part of the plurality of parts is not detected from a part of the plurality of human bodies, and the key point corresponding to the first part is detected from another part of the plurality of human bodies, calculates the integrated feature of the first part based on the feature of the key point corresponding to the first part detected from the other part.
Citation Information
Patent Citations
Method for establishing action recognition library, electronic device, and storage medium
CN109308438A
Obtaining metrics for position using frames classified by associative memory
JP2016058078A
Image retrieving apparatus, image retrieving method, and setting screen used therefor
JP2019091138A
Object recognition device, object recognition method and object recognition program
JP2020135551A
Action analysis device and action analysis method
JP2020135747A