Posture Estimation Device, Posture Estimation Method, and Program
The posture estimation device and method address the lack of an optimal keypoint association algorithm by dynamically selecting an algorithm based on image-specific factors, achieving accurate pose estimation of persons in images.
Patent Information
- Application Number
- JP2024541918
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-01-21
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-01-21
AI Technical Summary
There is no optimal algorithm for key-point association in pose estimation for images containing multiple persons, as existing algorithms are not suitable for all situations.
A posture estimation device and method that dynamically selects a keypoint association algorithm based on selection factors such as density and resolution of persons in the image, allowing for accurate pose estimation by dividing keypoints into groups belonging to the same person.
The solution enables accurate pose estimation of persons in images by selecting the most appropriate keypoint association algorithm based on image-specific factors, improving the reliability and precision of pose estimation across varying image conditions.
Smart Images

Figure 0007697602000002 
Figure 0007697602000003 
Figure 0007697602000004
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to a technique for analyzing an image and estimating the pose of a person imaged in the image.
Background Art
[0002] There are various types of analyses performed on an image in which one or more persons are imaged. One of these analyses is pose estimation for estimating the pose of each person imaged in the image. The pose of a person can be estimated based on key points such as the joints of the body detected from the image.
[0003] When a plurality of persons are imaged in the image, pose estimation includes a process called "key-point association" that divides the key points into groups such that each group includes key points belonging to the same person. Patent Document 1 discloses one of the algorithms for key-point association.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] There are multiple algorithms for key-point association, and there is no optimal algorithm for all situations. The object of the present disclosure is to provide a novel technique for accurately estimating the pose of a person imaged in an image.
Means for Solving the Problems
[0006] The posture estimation device provided by the present disclosure includes at least one memory configured to store instructions and at least one processor. The processor is configured to obtain a target image in which one or more persons are imaged, detect key points from the target image, calculate one or more selection factors based on the key points, where the one or more selection factors include the density, resolution, or both of the persons in the target image, select an algorithm for key point association from a pre-defined algorithm for key point association based on the one or more selection factors, perform key point association on the key points using the selected algorithm to divide the key points into one or more key point groups, each including key points of the same person, and estimate the posture of the person corresponding to each key point group based on the key points included in the key point group.
[0007] The posture estimation method provided by the present disclosure is executed by one or more computers. The method includes obtaining a target image in which one or more persons are imaged, detecting key points from the target image, calculating one or more selection factors based on the key points, where the one or more selection factors include the density, resolution, or both of the persons in the target image, selecting an algorithm for key point association from a pre-defined algorithm for key point association based on the one or more selection factors, performing key point association on the key points using the selected algorithm to divide the key points into one or more key point groups, each including key points of the same person, and estimating the posture of the person corresponding to each key point group based on the key points included in the key point group.
[0008] The non - transitory computer - readable storage medium provided by the present disclosure stores a program. The program causes one or more computers to perform the steps of: obtaining a target image in which one or more persons are imaged; detecting key points from the target image; calculating one or more selection factors based on the key points, where the one or more selection factors include the density, resolution, or both of the persons in the target image; selecting an algorithm for keypoint association from a pre - defined algorithm for keypoint association based on the one or more selection factors; performing keypoint association on the key points using the selected algorithm to divide the key points into one or more keypoint groups, each of which includes key points of the same person; and estimating the pose of the person corresponding to each keypoint group based on the key points included in the keypoint group.
Advantages of the Invention
[0009] According to the present disclosure if , a novel technique for accurately estimating the pose of a person in an image is provided .
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Mode for Carrying Out the Invention
[0011] Embodiments according to the present disclosure will be described below with reference to the drawings. In each drawing, the same elements are denoted by the same reference numerals, and redundant descriptions are omitted as necessary. Also, predetermined information (for example, a predetermined value or a predetermined threshold) is stored in advance in a storage device accessible by a computer that uses the information, unless otherwise specified.
[0012] Embodiment 1 <Overview> FIG. 1 shows an overview of the posture estimation device 2000 according to Embodiment 1. Note that the overview shown in FIG. 1 shows an example of the operation of the posture estimation device 2000 in order to facilitate understanding of the posture estimation device 2000, and does not limit or narrow the range of operations that the posture estimation device 2000 can perform.
[0013] The pose estimation device 2000 acquires a target image 10 in which one or more persons are imaged, and estimates the pose of each person. For this purpose, the pose estimation device 2000 detects key points from the target image 10, and performs key point association on the detected key points. The key points can indicate feature points of a person's body such as joints. The key point association is a process of dividing the key points into groups so that each group includes key points belonging to the same person. Based on the key points identified as belonging to that person by the key point association, the pose of each person can be estimated.
[0014] There are multiple algorithms for key point association, and which algorithm is suitable for estimating the pose of the person imaged in the image depends on the image to be analyzed. Hereinafter, the algorithm for key point association is referred to as the "key point association algorithm". Therefore, the pose estimation device 2000 calculates factors related to the target image 10, and selects a key point association algorithm for the target image 10 from a plurality of pre-defined key point association algorithms. Hereinafter, this factor is referred to as the "selection factor". The selection factor may include the density, resolution, or both of the persons in the target image 10.
[0015] The pose estimation device 2000 executes the selected key point association algorithm on the key points detected from the target image 10 to obtain groups of key points (hereinafter referred to as key point groups), and each key point group includes key points estimated to belong to the same person. Then, for each key point, the pose estimation device 2000 identifies the pose of the person corresponding to the key point group based on the key points included in the key point group.
[0016] <Example of effects> There are various keypoint association algorithms, and there is no optimal algorithm for every situation. According to the pose estimation device 2000, the keypoint association algorithm applied to the target image 10 is not fixed, but is selected from a plurality of pre-defined keypoint association algorithms based on selection factors. The selection factors may include the density, resolution, or both of the people captured in the target image 10. In this way, the keypoint association algorithm applied to the target image 10 is appropriately selected based on the density, resolution, or both of the people captured in the target image 10. Therefore, the pose of the people in the target image 10 can be accurately estimated.
[0017] Hereinafter, a detailed description of the pose estimation device 2000 will be given.
[0018] <Example of functional configuration> FIG. 2 is a block diagram showing an example of the functional configuration of the pose estimation device 2000 according to Embodiment 1. The pose estimation device 2000 includes an acquisition unit 2020, a keypoint detection unit 2040, an algorithm selection unit 2060, a keypoint association unit 2080, and an estimation unit 2100. The acquisition unit 2020 acquires the target image 10. The keypoint detection unit 2040 detects keypoints from the target image 10. The algorithm selection unit 2060 calculates one or more selection factors and selects a keypoint association algorithm from the pre-defined ones based on the calculated selection factors. The keypoint association unit 2080 executes the selected keypoint association algorithm on the detected keypoints, thereby generating a keypoint group. For each keypoint group, the estimation unit 2100 estimates the pose of the person corresponding to the keypoint group based on the keypoints in the keypoint group.
[0019] <Example of hardware configuration> The posture estimation device 2000 may be implemented by one or more computers. Each of the one or more computers may be a dedicated computer manufactured to implement the posture estimation device 2000, or may be a general-purpose computer such as a personal computer (PC), a server machine, or a mobile device.
[0020] The posture estimation device 2000 may also be implemented by installing an application on a computer. The application is realized by a program for causing the computer to function as the posture estimation device 2000. That is, the program implements the functional units of the posture estimation device 2000.
[0021] FIG. 3 is a block diagram showing an example of the hardware configuration of a computer 1000 that realizes the posture estimation device 2000 according to Embodiment 1. In FIG. 3, the computer 1000 includes a bus 1020, a processor 1040, a memory 1060, a storage device 1080, an input / output (I / O) interface 1100, and a network interface 1120.
[0022] Bus 1020 is a data transmission path for the processor 1040, the memory 1060, the storage device 1080, the input / output interface 1100, and the network interface 1120 to transmit and receive data from each other. The processor 1040 is a processor such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or an FPGA (Field-Programmable Gate Array). The memory 1060 is a main memory element such as a RAM (Random Access Memory) or a ROM (Read Only Memory). The storage device 1080 is an auxiliary storage element such as a hard disk, an SSD (Solid State Drive), or a memory card. The input / output interface 1100 is an interface between the computer 1000 and peripheral devices such as a keyboard, a mouse, and a display device. The network interface 1120 is an interface between the computer 1000 and a network. The network may be a LAN (Local Area Network) or a WAN (Wide Area Network).
[0023] The hardware configuration of the computer 1000 is not limited to that shown in FIG. 3. For example, as described above, the posture estimation device 2000 may be implemented by a plurality of computers. In this case, those computers may be connected to each other via a network.
[0024] <Flow of processing> FIG. 4 is a flowchart showing an exemplary flow of processing executed by the posture estimation device 2000 according to Embodiment 1. The acquisition unit 2020 acquires the target image 10 (S102). The keypoint detection unit 2040 detects keypoints from the target image 10 (S104). The algorithm selection unit 2060 calculates one or more selection factors (S106). The algorithm selection unit 2060 selects a keypoint association algorithm from those predefined based on the calculated selection factors (S108). The keypoint association unit 2080 executes the selected algorithm on the detected keypoints to generate a keypoint group (S110). The estimation unit 2100 estimates the posture of the person for each keypoint group (S112).
[0025] <Acquisition of the target image 10: S102> The acquisition unit 2020 acquires the target image 10. There are various methods for acquiring the target image 10. In some embodiments, the target image 10 is pre-stored in a storage device so that the posture estimation device 2000 can acquire it. In this case, the acquisition unit 2020 may access the storage device to acquire the target image. In other embodiments, the target image 10 may be transmitted by another computer, such as a camera that generates the target image 10. In this case, the acquisition unit 2020 may acquire it by receiving the target image 10.
[0026] In some embodiments, the target image 10 may be one of a series of images, such as a video frame forming a video. In this case, the posture estimation device 2000 can acquire all or part of the series of images as the target image 10 and estimate the posture of each person for each target image 10.
[0027] <Detection of keypoints: S104> The keypoint detection unit 2040 detects keypoints from the target image 10 (S104). There are various methods for detecting human keypoints from an image, and the keypoint detection unit 2040 can use one of these methods to detect keypoints from the target image 10. The human keypoints may be one or more features of the human body, such as joints at the neck, shoulders, elbows, wrists, waist, knees, ankles, etc.
[0028] In some embodiments, the keypoint detection unit 2040 is configured to acquire an image as input and further has a machine learning-based model (e.g., a neural network) that is pre-trained to detect keypoints from the input image in response to the input of the image. Hereinafter, this model will be referred to as the "keypoint detection model".
[0029] The keypoint detection model acquires the target image 10 as input, extracts features from the target image 10, detects one or more keypoints from the target image 10 based on the extracted features, and can identify the class of each keypoint (e.g., the neck, the right shoulder, etc.) based on the extracted features. In this case, the keypoint detection model may include a first model pre-trained to extract features from the target image 10 and a second model pre-trained to detect and classify keypoints based on the features extracted from the target image 10. Each of the first model and the second model can be configured as a machine learning-based model such as a neural network. Note that there are various types of machine learning models that can detect keypoints from an input image and further classify them, and the keypoint detection model can be configured as one of such models.
[0030] <Calculation of selection factor: S106> To select a keypoint association algorithm suitable for the target image 10, the algorithm selection unit 2060 calculates a selection factor based on the detected keypoints (S106). As described above, the selection factor may include the density of people in the target image 10, the resolution of people in the target image 10, or both. Hereinafter, an example of a method for calculating these factors will be described.
[0031] 《Density of People》 When using the density of people as a selection factor, the algorithm selection unit 2060 calculates the density of people in the target image 10 based on the keypoints detected from the target image 10. The density of people can be measured using the keypoints of the right shoulder and the left shoulder.
[0032] FIG. 5 is a flowchart showing an example of a method for calculating the density of people in the target image 10. The algorithm selection unit 2060 extracts the keypoints representing the left shoulder or the right shoulder from all the detected keypoints (S202). Then, for each keypoint of the left shoulder, the algorithm selection unit 2060 searches for the nearest keypoint of the right shoulder and links them (S204). By this step, some keypoints of the right shoulder can be linked to a plurality of keypoints of the left shoulder.
[0033] For each MP point, the algorithm selection unit 2060 extracts the longest link and the shortest link having that MP point as one end, and when the length of the longest link exceeds a predetermined multiple (for example, 2 times) of the length of the shortest link, the longest link is removed (S206). By this step, some MP points may become non-MP points (that is, keypoints of the right shoulder linked only to a single keypoint of the left shoulder) as a result of removing the longest link.
[0034] Figure 6 shows an example of the case from step 202 to step 206. In this example, in step S202, three right shoulders 22-1 to 22-3 and five left shoulders 24-1 to 24-5 are detected. Next, in step S204, the left shoulders 24-1 to 24-5 are linked to the right shoulders 22-1, 22-2, 22-2, 22-3, and 22-3, respectively. In this case, the right shoulders 22-2 and 22-3 are MP points.
[0035] In step S206, for the right shoulder 22-3, it is determined that the length of the longest link between the right shoulder 22-3 and the left shoulder 24-5 exceeds a predetermined multiple of the length of the shortest link between the right shoulder 22-3 and the left shoulder 24-4. As a result, the link between the right shoulder 22-3 and the left shoulder 24-5 is removed. As a result of this removal, the right shoulder 22-3 becomes a non-MP point.
[0036] Step S206 is repeated until the number of MP points converges (for example, becomes constant). Hereinafter, this number of MP points is referred to as "NMP". Specifically, the algorithm selection unit 2060 determines whether NMP has converged (S208). If NMP has not converged (S208: NO), step S206 is executed again. On the other hand, if NMP has converged (S208: YES), the algorithm selection unit 2060 sets NMP as the density of people in the target image 10. Conceptually, the larger NMP is, the denser the people are in the target image 10.
[0037] In another aspect, the algorithm selection unit 2060 may calculate the density of people in the target image 10 based on NMP. For example, for calculating the density, a function that outputs a value proportional to the input value is defined in advance. In this case, the algorithm selection unit 2060 inputs NMP to this function, obtains an output value proportional to NMP, and can use this output value as the density of people in the target image 10.
[0038] 《Resolution of People》 When using the resolution of a person as a selection factor, the algorithm selection unit 2060 may calculate the resolution of the person based on the NMP described above. Specifically, the algorithm selection unit 2060 extracts links with a length less than a threshold value (for example, 1 / 25 of the width of the target image 10) that can be defined based on one of the dimensions of the target image from the MP points remaining after step S208 in FIG. 5. The links extracted here are referred to as "SL links".
[0039] The algorithm selection unit 2060 may calculate a value called RSL as follows.
Equation
[0040] Conceptually, the larger the RSL, the lower the resolution of the person in the target image 10. Therefore, in some embodiments, the algorithm selection unit 2060 may calculate the resolution of the person in the target image 10 as a value that increases as the RSL decreases. For example, for calculating the resolution, a function that outputs a value proportional to the reciprocal of the input value in advance is predefined . In this case, the algorithm selection unit 2060 may input the RSL into this function to obtain an output value proportional to 1 / RSL, and use this output value as the resolution of the person in the target image 10. Note that this function may be defined to output the maximum value when 0 is given as the input.
[0041] In other embodiments, the algorithm selection unit 2060 may use the RSL as a selection factor representing the resolution of the person in the target image 10. In this case, as will be described later, when determining whether the resolution of the person in the target image 10 is less than the threshold value, the algorithm selection unit 2060 may determine that the resolution of the person in the target image 10 is less than the threshold value when the RSL is greater than the threshold value.
[0042] <Keypoint Association Algorithm> The pre-defined keypoint association algorithm may include two or more of 1) the mid-point algorithm, 2) the direction map algorithm, and 3) the position map algorithm. Each algorithm will be described below.
[0043] <<Mid-Point Algorithm>> The mid-point algorithm detects a mid-point from the target image 10 in order to perform keypoint association. The mid-point is a point located in the middle of two keypoints. The details of the mid-point algorithm are disclosed in Patent Document 1.
[0044] The mid-point algorithm can be implemented using a machine learning model such as a neural network. Hereinafter, this machine learning model will be referred to as the "mid-point model". The mid-point model can be configured to obtain the target image 10 as input and can be pre-trained to output the mid-point in the target image 10 in response to the input of the input data.
[0045] The mid-point algorithm can input the target image 10 into the mid-point model in order to obtain the mid-point in the target image 10. Then, the mid-point algorithm divides the keypoints into keypoint groups based on the mid-point.
[0046] <<Direction Map Algorithm>> The orientation map algorithm generates an orientation map of the target image 10 and uses the orientation map for keypoint association to divide the keypoints into keypoint groups. The orientation map is a feature map extracted from the target image 10 and has the same size as the target image 10. The orientation map indicates the unit vector for each pixel in the region of the person in the target image 10 (hereinafter referred to as the person region). The unit vector corresponding to a pixel within the person region points from that pixel to a predefined reference point of that person region. The reference point of the person region may be a specific keypoint of the person corresponding to the person region (for example, the keypoint of the neck).
[0047] More specifically, the orientation map may include a set of two feature maps called the H-orientation map and the V-orientation map. In the H-orientation map, the pixels of the person region indicate the horizontal component (i.e., the x-component) of the corresponding unit vector. On the other hand, in the V-orientation map, the pixels of the person region indicate the vertical component (y-component) of the corresponding unit vector.
[0048] The orientation map algorithm can be implemented using a machine learning model such as a neural network. Hereinafter, this machine learning model is referred to as the "orientation map model". The orientation map model is configured to receive the target image 10 as input and can be pre-trained to output the orientation map of the target image 10 in response to the input of the input data. The orientation map algorithm can input the target image 10 into the orientation map model to obtain the orientation map. Then, the orientation map algorithm divides the keypoints into keypoint groups using the orientation map.
[0049] The orientation map algorithm calculates a score for each pair of a keypoint detected from the target image 10 and a person using the orientation map, and based on the calculated score, identifies which keypoint belongs to which person. Then, for each keypoint, the orientation map algorithm identifies that the keypoint belongs to the person corresponding to the maximum score among the scores related to that keypoint.
[0050] Suppose that the pair of keypoint K1 and person P1 has a score S1, the pair of keypoint K1 and person P2 has a score S2, and the pair of keypoint K1 and person P3 has a score S3. Also, assume that the maximum score among S1, S2, and S3 is S2. In this case, since the score of the pair of keypoint K1 and person P2 is the maximum, the direction map algorithm identifies that keypoint K1 belongs to person P2.
[0051] The score of the pair of a keypoint and a person can be calculated as the product of three factors OB, RoD, and D (i.e., S = OB * RoD * D). These three factors are calculated as follows. The direction map algorithm generates one or more intermediate points on the line between the keypoint and the reference point of the person. For each intermediate point, the direction map algorithm determines whether the intermediate point is located within the person region. The direction map algorithm calculates, as factor OB, what percentage of the intermediate points are located within the person region. For example, if two out of three intermediate points are located within the person region, factor OB is 2 / 3.
[0052] And for each intermediate point located within the person region, the direction map algorithm obtains the unit vector corresponding to that intermediate point from the direction map. Let the coordinates of the intermediate point on target image 10 be (x1, y1). In this case, the unit vector corresponding to the intermediate point is the one located at (x1, y1) in the direction map. The direction map algorithm also obtains the unit vector corresponding to the keypoint.
[0053] The direction map algorithm calculates, as factor RoD, the variation in the direction of the obtained unit vectors. The variation in the direction of the unit vectors represents the maximum difference between two of those unit vectors.
[0054] As a factor D, the direction map algorithm calculates the distance between the key points of a person and the reference point.
[0055] <<Position Map Algorithm>> The position map algorithm calculates the position map of the target image 10, and uses the position map for key point association to divide the key points into key point groups. The position map is a feature map extracted from the target image 10 and has the same size as the target image 10. In the position map, the pixels of the person area indicate the relative position of the person area. The relative position of the person area is the position of the reference point (e.g., the key point of the neck) with respect to the base position of the target image 10. The base position of the target image 10 can be its origin (e.g., the upper left corner).
[0056] More specifically, the position map may include two feature maps called the H position map and the V position map. In the H position map, the pixels within the person area indicate the horizontal position of the reference point of that person area with respect to the base position (e.g., the left end of the target image 10). On the other hand, in the V position map, the pixels of the person area indicate the vertical position of the reference point of that person area with respect to the base position (e.g., the upper end of the target image 10).
[0057] Let the width of the target image 10 be W and the height be H, and let the absolute coordinates of the reference point of the person area in the target image 10 be (x1, y1). In this case, the relative position of the person area is (x1 / W, y1 / W). Therefore, the pixels of this person area indicate x1 / W in the H position map and y1 / H in the V position map.
[0058] The position map algorithm can be implemented using a machine learning model such as a neural network. Hereinafter, this machine learning model is referred to as a "position map model". The position map model is configured to obtain the target image 10 as input and can be pre-trained to output a position map of the target image 10 in response to the input of the input data. The position map algorithm can input the target image 10 into the position map model in order to obtain the position map. Then, the position map algorithm uses the position map to divide the keypoints into keypoint groups.
[0059] For each keypoint detected from the target image 10, the position map algorithm calculates the distance from each person detected from the target image 10. The distance is calculated using the relative position between the keypoint obtained from the position map and the person.
[0060] Specifically, the position map algorithm obtains a pixel value from a pixel of the position map whose coordinates on the position map are the same as the coordinates of the keypoint on the target image 10, and uses the obtained value as the relative position of the keypoint. In the target image 10, let the coordinates of the keypoint K1 be (x1, y1). Also, the pixel at (x1, y1) in the H position map indicates x2, and the pixel at (x1, y1) in the V position map indicates y2. In this case, the position map algorithm obtains the pixel value x2 from the pixel at (x1, y1) in the H position map as the x coordinate of the relative position of the keypoint K1. Also, the position map algorithm obtains the pixel value y2 of the pixel at (x1, y1) in the V position map as the y coordinate of the relative position of the keypoint K1. Thereby, the relative position of the keypoint K1 is specified as (x2, y2).
[0061] Regarding the relative position of a person, the position map algorithm obtains pixel values from the pixels within the person area corresponding to that person and uses the obtained values as the relative position of the person. Assume that the pixels in the person area of person P1 show the value x3 in the H position map and the pixels in the person area of person P1 show the value y3 in the V position map. In this case, the relative position of person P1 is specified as (x3, y3). As described above, when the relative position of the keypoint is indicated by (x2, y2), the position map algorithm calculates the distance between (x2, y2) and (x3, y3) as the distance between the keypoint and person P1.
[0062] After calculating the distances from each keypoint to each person, the position map algorithm identifies the person with the shortest distance from the keypoint. Then, the position map identifies that the keypoint belongs to the identified person.
[0063] <Model Configuration> <<Regarding the Direction Map Model>> The direction map is a feature map that describes the geometric relationship between a reference point (such as the keypoint of the head) and any other pixel within the entire body area of a person. To improve the quality of the direction map, it is preferable that the direction map model fully understands the body situation of the person (i.e., the connection method between different body parts). The midpoint defined between two pairs of keypoints helps in the proper understanding of the connection between different body parts by the direction map model and can thus help improve the quality of the direction map.
[0064] Based on the above insights, the direction map is preferably configured to generate a direction map using the key points and midpoints detected from the target image 10. Therefore, when the midpoint algorithm and the direction map model algorithm are adopted as the pre-defined key point association algorithms, the direction map model can be configured to obtain not only the target image 10 but also the output of the midpoint model as inputs. In this case, the midpoint model and the direction map model can be trained together using the same training data for each other. Note that the midpoint model may have a function as a key point detection model to detect key points from the target image 10.
[0065] Figure 7 shows the training of the midpoint model and the direction map model. In this example, the midpoint model receives the target image 10 as an input and detects key points and midpoints from the target image 10. The direction map model is connected to the midpoint model so that the direction map model can receive the output of the midpoint (i.e., key points and midpoints) as an input.
[0066] The training data for training the model shown in Figure 7 includes a set of input images in which one or more persons are imaged and ground truth data. The ground truth data shows the key points and midpoints detected from the corresponding input images and the direction map generated from the corresponding input images. The model calculates a loss representing the difference between those outputs (i.e., the key points and midpoints detected by the midpoint and the direction map generated by the direction map model) and the ground truth data, and is further trained by updating the trainable parameters of the model based on the calculated loss.
[0067] <<Regarding the position map model>> For any pixel within the body region of a person excluding the reference point, the values in the person's two-direction maps (denoted by vx and vy respectively) are the X and Y components of the unit vector from the pixel to the reference point. Therefore, vx and vy satisfy the condition "vx^2 + vy^2 = 1", which means that for all pixels within the body region of the person excluding the reference point, the sum of the squares of vx and vy is a constant value. On the other hand, the position map is defined as filling the body region of the person using a constant value corresponding to the position of the person within the image. Therefore, the square of the direction map can help improve the quality of the position map by converging the values of all pixels in the body region of the person to a constant value. Note that the square of the direction map includes the square of the H-direction map where each pixel indicates the square of the value indicated by the corresponding pixel, and the square of the V-direction map where each pixel indicates the square of the value indicated by the corresponding pixel.
[0068] Based on the above insights, it is preferable to configure the position map model to generate the position map using the square of the direction map. Therefore, when the direction map model algorithm and the position map algorithm are adopted as the pre-defined keypoint association algorithm, the position map model can be configured to obtain not only the target image 10 but also the square of the output of the direction map model as an input. In this case, the direction map model and the position map can be trained together using the same training data for each other.
[0069] Figure 8 shows the training of the direction map model and the position map model. In this example, there is a component between the direction map model and the position map model that calculates the square of the output of the direction map model. This component is configured to obtain the output of the direction map model, calculate the square of this output, and supply the result of the calculation to the position map model.
[0070] The training data for training the model shown in FIG. 8 includes a set of input images in which one or more persons are imaged and ground truth data. The ground truth data indicates a direction map and a position map generated from the corresponding input image. The model calculates a loss representing the difference between those outputs (i.e., the direction map generated by the direction map model and the position map generated by the position map model) and the ground truth data, and is further trained by updating the trainable parameters of the model based on the calculated loss.
[0071] By combining the configurations shown in FIGS. 7 and 8, the midpoint model, the direction map model, and the position map model can be trained together when those models are adopted as a pre-defined keypoint association algorithm. FIG. 9 shows the training of the midpoint model, the direction map model, and the position map model.
[0072] The training data for training the model shown in FIG. 9 includes a set of input images in which one or more persons are imaged and ground truth data. The ground truth data indicates the keypoints and midpoints detected from the corresponding input image, and the direction map and the position map generated from the corresponding input image. The model calculates a loss representing the difference between those outputs (i.e., the keypoints and midpoints detected by the midpoint, the direction map generated by the direction map model, and the position map generated by the position map model) and the ground truth data, and is further trained by updating the trainable parameters of the model based on the calculated loss.
[0073] <Selection of Keypoint Association Algorithm: S108> The algorithm selection unit 2060 selects a keypoint association algorithm suitable for the target image 10 based on a selection factor (S108). In some embodiments, the algorithm selection unit 2060 may determine the keypoint association algorithm based on whether the selection factor is greater than a pre-defined threshold.
[0074] FIG. 10 is a flowchart showing a first exemplary flow of a process for selecting a keypoint association algorithm. In this example, the pre-defined keypoint algorithms include a midpoint algorithm, a direction map algorithm, and a position map algorithm. Also, in this example, the selection factors include the density of people in the target image 10 and the resolution of the people in the target image 10.
[0075] Specifically, the algorithm selection unit 2060 determines whether the resolution of the people in the target image 10 is less than a threshold ThR (S302). When the resolution is less than the threshold ThR (S302: YES), the algorithm selection unit 2060 selects the midpoint algorithm as the keypoint association algorithm to be applied to the target image 10 (S304). On the other hand, when the resolution is greater than or equal to the threshold ThR (S302: NO), the algorithm selection unit 2060 determines whether the density of the people in the target image 10 is less than a threshold ThD (S306).
[0076] When the density is less than the threshold ThD (S306: YES), the algorithm selection unit 2060 selects the direction map algorithm as the keypoint association algorithm to be applied to the target image 10 (S308). On the other hand, when the density is greater than or equal to the threshold ThD (S306: NO), the algorithm selection unit 2060 selects the position map algorithm as the keypoint association algorithm to be applied to the target image 10 (S310).
[0077] FIG. 11 is a flowchart showing a second exemplary flow of a process for selecting a keypoint association algorithm. In this example, the pre-defined keypoint algorithms include a midpoint algorithm and a direction map algorithm. Also, in this example, the resolution of the people in the target image 10 is used as a selection factor.
[0078] Specifically, the algorithm selection unit 2060 determines whether the resolution of the person in the target image 10 is smaller than the threshold ThR (S402). When the resolution is smaller than the threshold ThR (S402: YES), the algorithm selection unit 2060 selects the midpoint algorithm as the keypoint association algorithm to be applied to the target image 10 (S404). On the other hand, when the resolution is equal to or higher than the threshold ThR (S402: NO), the algorithm selection unit 2060 selects the direction map algorithm as the keypoint association algorithm to be applied to the target image 10 (S406).
[0079] Note that in the example shown in FIG. 11, instead of the direction map algorithm, a position map algorithm may be adopted as one of the pre-defined keypoint association algorithms. In this case, when the resolution is equal to or higher than the threshold ThR, the algorithm selection unit 2060 selects the position map algorithm as the keypoint association algorithm to be applied to the target image 10.
[0080] FIG. 12 is a flowchart showing a third exemplary flow of the process of selecting a keypoint association algorithm. In this example, the pre-defined keypoint algorithms include a direction map algorithm and a position algorithm. Also, in this example, the density of the people in the target image 10 is used as a selection factor.
[0081] Specifically, the algorithm selection unit 2060 determines whether the density of the people in the target image 10 is smaller than the threshold ThD (S502). When the density is smaller than the threshold ThD (S502: YES), the algorithm selection unit 2060 selects the direction map algorithm as the keypoint association algorithm to be applied to the target image 10 (S504). On the other hand, when the density is equal to or higher than the threshold ThD (S502: NO), the algorithm selection unit 2060 selects the position map algorithm as the keypoint association algorithm to be applied to the target image 10 (S506).
[0082] In the example shown in FIG. 12, instead of the direction map algorithm, a midpoint algorithm may be adopted as one of the pre-defined keypoint association algorithms. In this case, when the density is equal to or greater than the threshold ThD, the algorithm selection unit 2060 selects the midpoint algorithm as the keypoint association algorithm to be applied to the target image 10.
[0083] <Output from the pose estimation device 2000> The pose estimation device 2000 may be configured to output information indicating the result of pose estimation (referred to as output information). For example, the output information may include an identifier of the target image 10 (e.g., frame number), for each keypoint group, an identifier of the estimated pose of the keypoint group, and a set of keypoint information of each keypoint within the keypoint group. The identifier of the estimated pose indicates what kind of pose the person corresponding to the keypoint group is taking. The keypoint information indicates the type (e.g., neck, right shoulder, etc.) and position (e.g., coordinates) of the keypoint.
[0084] There are various ways to output the output information. In some embodiments, the output information may be stored in a storage device, displayed on a display device, or transmitted to another computer such as the user's PC or smartphone of the pose estimation device 2000.
[0085] The program can be stored using various types of non-transitory computer readable media and provided to a computer. Non-transitory computer readable media include various types of tangible storage media. Examples of non-transitory computer readable media include magnetic recording media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical recording media (e.g., magneto-optical disks), CD-ROM, CD-R, CD-R / W, semiconductor memories (e.g., mask ROM, PROM (Programmable ROM), EPROM (Erasable PROM), flash ROM, RAM). Also, the program may be provided to the computer by various types of transitory computer readable media. Examples of transitory computer readable media include electrical signals, optical signals, and electromagnetic waves. The transitory computer readable media can supply the program to the computer via wired communication paths such as electric wires and optical fibers, or wireless communication paths.
[0086] The present disclosure has been described above with reference to the embodiments, but the present disclosure is not limited to the above-described embodiments. Various changes understandable by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present invention.
[0087] Some or all of the above embodiments may be described as follows in the appended claims, but are not limited thereto. <Appended Claims> (Appended Claim 1) A posture estimation device, at least one memory configured to store instructions, acquiring a target image in which one or more persons are imaged, detecting key points from the target image, Calculating one or more selection factors based on the key points, where the one or more selection factors include the density, resolution, or both of the people in the target image, Selecting an algorithm for keypoint association from a pre-defined algorithm for keypoint association based on the one or more selection factors, Performing keypoint association on the keypoints using the selected algorithm to divide the keypoints into one or more keypoint groups, each of which includes keypoints of the same person, Estimating the pose of the person corresponding to each keypoint group based on the keypoints included in the keypoint group, At least one processor configured to execute the instructions for performing the above, a pose estimation device. (Appendix 2) The types of the keypoints include the right shoulder and the left shoulder, The calculation of the density of the people in the target image is as follows: For each keypoint of the left shoulder, linking the keypoint of the left shoulder to the keypoint of the right shoulder closest to it, For each keypoint of the right shoulder linked to multiple keypoints of the left shoulder, when a predetermined multiple of the length of the shortest link having the keypoint of the right shoulder exceeds the length of the longest link, deleting the longest link having the keypoint of the right shoulder, Calculating the density of the people in the target image based on the number of keypoints of the right shoulder linked to multiple keypoints of the left shoulder. The pose estimation device according to Appendix 1 includes this. (Appendix 3) The types of the keypoints include the right shoulder and the left shoulder, The calculation of the resolution of the people in the target image is as follows: For each keypoint of the left shoulder, linking the keypoint of the left shoulder to the keypoint of the right shoulder closest to it, For each of the key points on the right shoulder linked to a plurality of key points on the left shoulder, when the length of the longest link exceeds a predetermined multiple of the length of the shortest link having the key point on the right shoulder, deleting the longest link having the key point on the right shoulder; calculating the resolution of the person in the target image based on the number of the remaining links after the deletion and the number of the links having a length less than a threshold value defined based on one of the dimensions of the target image, the posture estimation device according to Supplementary Note 1 or 2. (Supplementary Note 4) The pre-defined algorithm is the midpoint algorithm, the direction map algorithm, the position map algorithm, or an algorithm including two or three of them, the posture estimation device according to any one of Supplementary Notes 1 to 3. (Supplementary Note 5) The selection algorithm for key point association is determining whether the selection factor is less than a threshold value of the selection factor; when the selection factor is less than the threshold value of the selection factor, selecting a first algorithm for key point association; when the selection factor is greater than or equal to the threshold value of the selection factor, selecting a second algorithm for key point association, the posture estimation device according to any one of Supplementary Notes 1 to 4. (Supplementary Note 6) The selection algorithm for key point association is determining whether the resolution of the person in the target image is less than a threshold value of the resolution; when the resolution of the person in the target image is less than the threshold value of the resolution, selecting the midpoint algorithm; when the resolution of the person in the target image is greater than or equal to the threshold value of the resolution, selecting the direction map algorithm or the position map algorithm, the posture estimation device according to Supplementary Note 5. (Supplementary Note 7) The selection algorithm for key point association is Determining whether the density of the person in the target image is less than the density threshold; When the density of the person in the target image is less than the density threshold, selecting the midpoint algorithm or the direction map algorithm; When the density of the person in the target image is greater than or equal to the density threshold, selecting the position map algorithm, the posture estimation device according to appended note 5 or 6. (Appended note 8) A posture estimation method executed by one or more computers, Obtaining a target image in which one or more persons are imaged; Detecting key points from the target image; Calculating one or more selection factors based on the key points, the one or more selection factors including the density, resolution, or both of the persons in the target image; Selecting an algorithm for key point association from pre - defined algorithms for key point association based on the one or more selection factors; Performing key point association on the key points using the selected algorithm to divide the key points into one or more key point groups, each of which includes key points of the same person; Estimating the posture of the person corresponding to each key point group based on the key points included in the key point group for each key point group. A posture estimation method. (Appended note 9) The types of the key points include the right shoulder and the left shoulder, The calculation of the density of the person in the target image is For each key point of the left shoulder, linking the key point of the left shoulder to the key point of the right shoulder closest to it; For each of the key points on the right shoulder linked to a plurality of key points on the left shoulder, when the length of the longest link exceeds a predetermined multiple of the length of the shortest link having the key point on the right shoulder, deleting the longest link having the key point on the right shoulder; calculating the density of the person in the target image based on the number of the key points on the right shoulder linked to the plurality of key points on the left shoulder, the method for estimating a pose according to Supplementary Note 8, comprising: (Supplementary Note 10) The types of the key points include the right shoulder and the left shoulder. The calculation of the resolution of the person in the target image For each of the key points on the left shoulder, linking the key point on the left shoulder to the key point on the right shoulder that is closest to it; For each of the key points on the right shoulder linked to a plurality of key points on the left shoulder, when the length of the longest link exceeds a predetermined multiple of the length of the shortest link having the key point on the right shoulder, deleting the longest link having the key point on the right shoulder; calculating the resolution of the person in the target image based on the number of the remaining links after the deletion and the number of the links having a length less than a threshold value defined based on one of the dimensions of the target image, the method for estimating a pose according to Supplementary Note 8 or 9, comprising: (Supplementary Note 11) The predefined algorithm is the midpoint algorithm, the direction map algorithm, the position map algorithm, or two or three of them, the method for estimating a pose according to any one of Supplementary Notes 8 to 10. (Supplementary Note 12) The selection algorithm for key point association determining whether the selection factor is less than a threshold value of the selection factor; when the selection factor is less than the threshold value of the selection factor, selecting a first algorithm for key point association; When the selection factor is equal to or greater than the threshold value of the selection factor, selecting a second algorithm for keypoint association, the posture estimation method according to any one of Appendices 8 to 11. (Appendix 13) The selection algorithm for keypoint association is determining whether the resolution of the person in the target image is less than the threshold value of the resolution; when the resolution of the person in the target image is less than the threshold value of the resolution, selecting a midpoint algorithm; when the resolution of the person in the target image is equal to or greater than the threshold value of the resolution, selecting a direction map algorithm or a position map algorithm, the posture estimation method according to Appendix 12. (Appendix 14) The selection algorithm for keypoint association is determining whether the density of the person in the target image is less than the threshold value of the density; when the density of the person in the target image is less than the threshold value of the density, selecting a midpoint algorithm or a direction map algorithm; when the density of the person in the target image is equal to or greater than the threshold value of the density, selecting a position map algorithm, the posture estimation method according to Appendix 12 or 13. (Appendix 15) A non-transitory computer-readable storage medium storing a program, wherein the program causes one or more computers to acquire a target image in which one or more persons are imaged; detect keypoints from the target image; calculating one or more selection factors based on the keypoints, the one or more selection factors including the density, resolution, or both of the persons in the target image; Based on the one or more selection factors, selecting an algorithm for keypoint association from a pre-defined algorithm for keypoint association; Performing keypoint association on the keypoints using the selected algorithm to divide the keypoints into one or more keypoint groups, each containing keypoints of the same person; For each keypoint group, estimating the pose of the person corresponding to the keypoint group based on the keypoints included in the keypoint group, a non-transitory computer-readable storage medium for causing the execution. (Appendix 16) The types of the keypoints include the right shoulder and the left shoulder. The calculation of the density of the person in the target image For each keypoint of the left shoulder, linking the keypoint of the left shoulder to the keypoint of the right shoulder closest to it; For each keypoint of the right shoulder linked to a plurality of keypoints of the left shoulder, when the length of the longest link exceeds a predetermined multiple of the length of the shortest link having the keypoint of the right shoulder, deleting the longest link having the keypoint of the right shoulder; Calculating the density of the person in the target image based on the number of keypoints of the right shoulder linked to a plurality of keypoints of the left shoulder, the storage medium according to Appendix 15. (Appendix 17) The types of the keypoints include the right shoulder and the left shoulder. The calculation of the resolution of the person in the target image For each keypoint of the left shoulder, linking the keypoint of the left shoulder to the keypoint of the right shoulder closest to it; For each keypoint of the right shoulder linked to a plurality of keypoints of the left shoulder, when the length of the longest link exceeds a predetermined multiple of the length of the shortest link having the keypoint of the right shoulder, deleting the longest link having the keypoint of the right shoulder; Calculating the resolution of the person in the target image based on the number of the remaining links after the deletion and the number of the links having a length less than a threshold value defined based on one of the dimensions of the target image, the storage medium according to appended note 15 or 16. (Appended note 18) The predefined algorithm is the storage medium according to any one of appended notes 15 to 17, including a midpoint algorithm, a direction map algorithm, a position map algorithm, or two or three of them. (Appended note 19) The selection algorithm for keypoint association is Determining whether the selection factor is less than a threshold value of the selection factor, When the selection factor is less than the threshold value of the selection factor, selecting a first algorithm for keypoint association, When the selection factor is greater than or equal to the threshold value of the selection factor, selecting a second algorithm for keypoint association, the storage medium according to any one of appended notes 15 to 18. (Appended note 20) The selection algorithm for keypoint association is Determining whether the resolution of the person in the target image is less than a threshold value of the resolution, When the resolution of the person in the target image is less than the threshold value of the resolution, selecting a midpoint algorithm, When the resolution of the person in the target image is greater than or equal to the threshold value of the resolution, selecting a direction map algorithm or a position map algorithm, the storage medium according to appended note 19. (Appended note 21) The selection algorithm for keypoint association is Determining whether the density of the person in the target image is less than a threshold value of the density, When the density of the person in the target image is less than the threshold value of the density, selecting a midpoint algorithm or a direction map algorithm; When the density of the person in the target image is greater than or equal to the threshold value of the density, selecting a position map algorithm, a storage medium according to appended note 19 or 20.
Explanation of symbols
[0088] 10 Target image 22 Keypoint of the right shoulder 24 Keypoint of the left shoulder 1000 Computer 1020 Bus 1040 Processor 1060 Memory 1080 Storage device 1100 Input / output interface 1120 Network interface 2000 Pose estimation device 2020 Acquisition unit 2040 Keypoint detection unit 2060 Algorithm selection unit 2080 Keypoint association unit 2100 Estimation unit
Claims
1. Acquisition means for acquiring a target image in which one or more persons are imaged; Keypoint detection means for detecting keypoints from the target image; Algorithm selection means for calculating one or more selection factors based on the keypoints, and selecting, based on the one or more selection factors, an algorithm to be applied to the keypoints from a predefined algorithm for keypoint association; Keypoint association means for performing keypoint association on the keypoints using the selected algorithm in order to divide the keypoints into one or more keypoint groups each including keypoints of the same person; Estimation means for estimating the posture of the person corresponding to each keypoint group based on the keypoints included in the keypoint group, the posture estimation device having: The one or more selection factors include the density, resolution, or both of the persons in the target image.
2. The types of the keypoints include the right shoulder and the left shoulder. The algorithm selection means calculates, as the selection factor, the density of the persons in the target image. The calculation of the density includes: Linking each keypoint of the left shoulder to the keypoint of the nearest right shoulder; When the length of the longest link having the keypoint of the right shoulder exceeds a predetermined multiple of the length of the shortest link having the keypoint of the right shoulder for each keypoint of the right shoulder linked to a plurality of keypoints of the left shoulder, deleting the longest link; Calculating the density of the persons in the target image based on the number of keypoints of the right shoulder linked to the plurality of keypoints of the left shoulder. The posture estimation device according to Claim 1.
3. The types of the keypoints include the right shoulder and the left shoulder. The algorithm selection means calculates, as the selection factor, the resolution of the persons in the target image. The calculation of the resolution includes: Linking each keypoint of the left shoulder to the keypoint of the nearest right shoulder; For each of the key points on the right shoulder linked to the key points on the left shoulder, when the length of the longest link having the key point on the right shoulder exceeds a predetermined multiple of the length of the shortest link having the key point on the right shoulder, deleting the longest link; calculating the resolution of the person in the target image based on the number of the remaining links after the deletion and the number of the links having a length less than a threshold value defined based on one of the dimensions of the target image, the posture estimation device according to claim 1 or 2.
4. The posture estimation device according to any one of claims 1 to 3, wherein the pre-defined algorithm includes a midpoint algorithm, a direction map algorithm, a position map algorithm, or two or three of them.
5. The selection of the algorithm is determining whether the selection factor is less than a threshold value of the selection factor; when the selection factor is less than the threshold value of the selection factor, selecting the first algorithm; when the selection factor is greater than or equal to the threshold value of the selection factor, selecting the second algorithm, the posture estimation device according to any one of claims 1 to 4.
6. The selection of the algorithm is determining whether the resolution of the person in the target image is less than a threshold value of the resolution; when the resolution of the person in the target image is less than the threshold value of the resolution, selecting a midpoint algorithm; when the resolution of the person in the target image is greater than or equal to the threshold value of the resolution, selecting a direction map algorithm or a position map algorithm, the posture estimation device according to claim 5.
7. The selection of the algorithm is determining whether the density of the person in the target image is less than a threshold value of the density; when the density of the person in the target image is less than the threshold value of the density, selecting a midpoint algorithm or a direction map algorithm; when the density of the person in the target image is greater than or equal to the threshold value of the density, selecting a position map algorithm, the posture estimation device according to claim 5 or 6.
8. A posture estimation method executed by one or more computers, The step of obtaining a target image in which one or more persons are imaged; The step of detecting keypoints from the target image; The step of calculating one or more selection factors based on the keypoints; The step of selecting, based on the one or more selection factors, an algorithm to be applied to the keypoints from a pre-defined algorithm for keypoint association; The step of performing keypoint association on the keypoints using the selected algorithm to divide the keypoints into one or more keypoint groups each including keypoints of the same person; The step of estimating, for each keypoint group, the pose of the person corresponding to the keypoint group based on the keypoints included in the keypoint group, wherein the one or more selection factors include the density, resolution, or both of the persons in the target image, and is a pose estimation method.
9. The step of obtaining a target image in which one or more persons are imaged; The step of detecting keypoints from the target image; The step of calculating one or more selection factors based on the keypoints; The step of selecting, based on the one or more selection factors, an algorithm to be applied to the keypoints from a pre-defined algorithm for keypoint association; The step of performing keypoint association on the keypoints using the selected algorithm to divide the keypoints into one or more keypoint groups each including keypoints of the same person; The step of estimating, for each keypoint group, the pose of the person corresponding to the keypoint group based on the keypoints included in the keypoint group, which is executed by one or more computers; wherein the one or more selection factors include the density, resolution, or both of the persons in the target image, and is a program.
Citation Information
Patent Citations
Method and device for identifying video event
CN110909655A
Openpose-based multi-person posture detection method and system
CN111310625A
Method and device for detecting keypoints
CN111368594A
Special effect processing method for live broadcasting, device, and server
JP2021157835A
Imaging system and method for object detection and localization
US20190019030A1