Learning device, estimation device, learning method, estimation method, and program

JPWO2024128124A5Active Publication Date: 2025-08-13NEC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024564333
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-08-13
Estimated Expiration
2043-12-07

AI Technical Summary

Technical Problem

Conventional methods face challenges in collecting and varying training data for estimating 3D skeletal information from human body images, leading to decreased estimation accuracy due to the need for specialized equipment and limited environmental conditions.

Method used

A learning device and method that acquire and utilize teacher images with correct labels for center positions and key points, along with relative depth directions, to learn an estimation model that estimates the relative position of key points in depth, enabling the collection of sufficient and varied learning data without specialized equipment.

Benefits of technology

This approach allows for the easy construction and high-accuracy estimation of 3D skeletal information from images, improving the estimation model's performance by using images alone to generate learning data and estimate positional information in real space.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The present invention provides a learning device comprising: an acquiring unit for acquiring learning data in which teacher images including human bodies, a correct answer label indicating a central position, in the image, of each human body, a correct answer label indicating a position, in the image, of each keypoint on each human body, and a correct answer label indicating a relative position, in a depth direction, of each keypoint on each human body are associated with one another; and a training unit which, on the basis of the learning data, trains an estimation model for estimating information indicating the likelihood of a central position, in the image, of each human body, the relative position, in the image, of each keypoint associated with each human body, and the relative position, in the depth direction, of each keypoint associated with each human body.
Need to check novelty before this filing date? Find Prior Art

Description

Learning device, estimation device, learning method, estimation method, and recording medium

[0001] The present invention relates to a learning device, an estimation device, a learning method, an estimation method, and a program.

[0002] A technology related to the present invention is disclosed in Non-Patent Document 1. The technology in Non-Patent Document 1 is used to estimate position information (three-dimensional skeletal information) in real space of key points of a human body (joint points / skeletal points of the human body) from an image using a trained estimation model.

[0003] In the conventional technology of Non-Patent Document 1, a single image is input to a trained estimation model configured with a convolutional neural network, and the position coordinates of key points of the human body (joint points / skeleton points of the human body) in real space (three-dimensional) are estimated. The training data used is data in which an image is paired with the position coordinates of the key points of the human body in real space.

[0004] Bugra Tekin et al., Structured Prediction of 3D Human Pose with Deep Neural Networks, [Retrieved September 20, 2022], Internet, <URL: https: / / arxiv.org / abs / 1605.05180>

[0005] The problem with Non-Patent Document 1 is that when training an estimation model, it is difficult to collect training data, such as position coordinates in real space (three-dimensional skeletal information) of key points on the human body, and therefore it is not easy to train the estimation model.

[0006] The reason for this is that position coordinates in real space (three dimensions) of key points on the human body cannot be easily generated / collected manually as long as there is an image of the human body, as can position coordinates on an image (two dimensions).Instead, training data cannot be collected without using large-scale equipment such as a motion capture system.

[0007] Another problem is that it is not possible to collect training data with a wide variety of images paired with position coordinates in real space (three-dimensional).

[0008] The reason for this is that equipment such as motion capture systems are installed in limited environments such as indoor laboratories due to the installation conditions of the equipment, and the images captured are limited in terms of variety in terms of background, number of people, depth, etc.

[0009] Another problem is that if the amount or variety of training data is insufficient, the estimation accuracy of the estimation model, i.e., the accuracy of the process of estimating the position information in real space of key points on the human body from images, will be low.

[0010] An example of an object of the present invention is to provide a learning device, an estimation device, a learning method, an estimation method, and a program that solves any of the above-mentioned problems.

[0011] According to one aspect of the present invention, there is provided a learning device having: an acquisition means for acquiring learning data linking teacher images including human bodies with correct labels indicating the center position of each human body on the image, correct labels indicating the position of each key point for each human body on the image, and correct labels indicating the relative position in the depth direction of each key point for each human body; and a learning means for learning an estimation model that estimates, based on the learning data, information indicating the likelihood of the center position of each human body on the image, the relative position on the image of each key point linked to each human body, and the relative position in the depth direction of each key point linked to each human body.

[0012] According to one aspect of the present invention, there is provided an estimation device having: estimation means for estimating, using an estimation model trained by the learning device, position coordinates on an image of each key point associated with each human body; a relative position in the depth direction of each key point associated with each human body with reference to the center position of the human body in real space; position coordinates indicating the center position of each human body on the image; and an order of each human body in the depth direction.

[0013] According to one aspect of the present invention, a learning method is provided in which one or more computers acquire learning data linking teacher images including human bodies with correct labels indicating the center position of each human body on the image, correct labels indicating the position of each key point for each human body on the image, and correct labels indicating the relative position in the depth direction of each key point for each human body, and learn an estimation model based on the learning data to estimate information indicating the likelihood of the center position of each human body on the image, the relative position on the image of each key point linked to each human body, and the relative position in the depth direction of each key point linked to each human body.

[0014] According to one aspect of the present invention, there is provided a program that causes a computer to function as: an acquisition means that acquires training data that links a teacher image including human bodies with a correct label indicating the center position of each human body on the image, a correct label indicating the position of each key point for each human body on the image, and a correct label indicating the relative position in the depth direction of each key point for each human body; and a learning means that, based on the training data, learns an estimation model that estimates information indicating the likelihood of the center position of each human body on the image, the relative position on the image of each key point linked to each human body, and the relative position in the depth direction of each key point linked to each human body.

[0015] According to one aspect of the present invention, there is provided an estimation method in which one or more computers use an estimation model learned by the learning device to estimate position coordinates on an image of each key point associated with each human body, relative positions in the depth direction of each key point associated with each human body based on the center position of the human body in real space, position coordinates indicating the center position of each human body on the image, and an order of each human body in the depth direction.

[0016] According to one aspect of the present invention, there is provided a program that causes a computer to function as an estimation means that uses an estimation model learned by the learning device to estimate the position coordinates on an image of each key point associated with each human body, the relative position in the depth direction of each key point associated with each human body based on the center position of the human body in real space, the position coordinates indicating the center position of each human body on the image, and the order of each human body in the depth direction.

[0017] According to one aspect of the present invention, the problem of providing a learning device, a learning method, and a program that can easily learn / construct an estimation model trained with a sufficient amount and variety of training data in an apparatus that uses a trained estimation model to estimate real-space position information of key points on a human body from an image is solved.

[0018] Furthermore, according to one aspect of the present invention, the object of the present invention is to provide an estimation device, an estimation method, and a program for estimating position information in real space of key points of a human body from an image with high accuracy.

[0019] The above-mentioned objects, as well as other objects, features, and advantages, will become more apparent from the following description of the preferred embodiments and the accompanying drawings.

[0020] FIG. 1 is a diagram for explaining the technology of the present embodiment. FIG. 2 is a diagram for explaining the technology of the present embodiment. FIG. 3 is a diagram for explaining the technology of the present embodiment. FIG. 4 is a diagram for explaining the technology of the present embodiment. FIG. 5 is an example of a functional block diagram of a learning device of the present embodiment. FIG. 6 is an example of a functional block diagram of a learning device of the present embodiment. FIG. 7 is a flowchart showing an example of a processing flow of the learning device of the present embodiment. FIG. 8 is an example of a functional block diagram of an estimation device of the present embodiment. FIG. 9 is an example of a functional block diagram of an estimation device of the present embodiment. FIG. 10 is a diagram for explaining the processing of the estimation device of the present embodiment. FIG. 11 is a diagram for explaining the technology of the present embodiment. FIG. 12 is a flowchart showing an example of a processing flow of the estimation device of the present embodiment. FIG. 13 is an example of a functional block diagram of a 3D skeleton estimation device of the present embodiment. FIG. 14 is a diagram for explaining the technology of the present embodiment. FIG. 15 is a diagram for explaining the technology of the present embodiment. FIG. 16 is a diagram for explaining the technology of the present embodiment. A diagram showing an example of a hardware configuration of the device of the present embodiment.

[0021] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings, like components are designated by like reference numerals, and the description thereof will be omitted as appropriate.

[0022] First Embodiment FIG. 6 is a functional block diagram showing an overview of a learning device 10 according to a first embodiment. The learning device 10 includes an acquisition unit 11 and a learning unit 12. The acquisition unit 11 acquires learning data that associates teacher images including human bodies with correct labels indicating the center positions of each human body on the image, correct labels indicating the positions of each key point on each human body on the image, and correct labels indicating the relative positions of each key point on each human body in the depth direction. Based on the learning data, the learning unit 12 learns an estimation model that estimates information indicating the likelihood of the center position of each human body on the image, the relative positions on the image of each key point associated with each human body, and the relative positions in the depth direction of each key point associated with each human body.

[0023] According to the learning device 10 having such a configuration, in a device that uses a trained estimation model to estimate position information in real space at key points on the human body from an image, an estimation model trained with a sufficient amount and variety of training data can be easily learned / constructed.

[0024] Second Embodiment Fig. 9 is a functional block diagram showing an overview of an estimation device 20 according to a second embodiment. The estimation device 20 includes an estimation unit 21. Using an estimation model trained by the learning device 10 described in the first embodiment, the estimation unit 21 estimates the position coordinates on the image of each key point associated with each human body, the relative position in the depth direction of each key point associated with each human body with reference to the center position of the human body in real space, the position coordinates indicating the center position of each human body on the image, and the order of each human body in the depth direction.

[0025] The estimation device 20 having such a configuration can estimate position information in real space at key points of a human body from an image with high accuracy.

[0026] "Third embodiment" <Overview> The estimation device 20 of this embodiment estimates similar information, "position information of a human body in real space" and "position information of key points associated with a human body in real space," instead of estimating "position information of a human body in real space" from an image.

[0027] "Position information of the human body in real space" refers to the position coordinates of the human body on the image and the order of the human body in the depth direction in real space. Also, "Position information of the keypoint associated with the human body in real space" refers to the position coordinates of the keypoint associated with the human body on the image and the relative position of the keypoint associated with the human body in the depth direction with the center of the human body as the reference. The real space position of the keypoint in the depth direction can be identified based on this relative position in the depth direction and the corresponding order of the human body in the depth direction in real space.

[0028] The learning device 10 of this embodiment learns a neural network (estimation model) that outputs information necessary to obtain the above information (information related to the above information). The above information and the output information of the estimation model related to the above information can be easily collected / generated manually as long as there is a human body image. Therefore, the learning data for the estimation model can be easily collected / generated. This makes it possible to easily learn / construct an estimation model in a device that estimates position information in real space of key points on a human body from an image.

[0029] <Features of the Technology of the Present Embodiment> The technology of the present embodiment will be described. As shown in Fig. 1, when an image is input to a neural network, a plurality of data as shown in the figure is output. In other words, the neural network of the present embodiment is composed of a plurality of layers that output a plurality of data as shown in the figure.

[0030] FIG. 2 shows examples of the "likelihood of human body position," "correction amount of human body position," and "depth information of human body" among the multiple data shown in FIG. 1. FIG. 15 shows examples of the "relative position of key point a" and "relative depth information of key point a" among the multiple data shown in FIG. 1. FIG. 16 shows examples of the "relative position of key point b" and "relative depth information of key point b" among the multiple data shown in FIG. 1. FIG. 3 shows an image that is the source of the data in FIGS. 2, 15, and 16, to which explanations illustrating the concepts of each of the data in FIGS. 2, 15, and 16 have been added.

[0031] The "likelihood of human body position" data shown in FIG. 2 is data indicating the likelihood of the center position of the human body on the image (the position of the human body), as shown in FIG. 3. As shown in the figure, this data indicates the likelihood that the center position of the human body on the image is located in each of multiple grids obtained by dividing the image. The likelihood may be indicated by a normal distribution or the like, centered on the grid indicating the center position of the human body on the image. Note that the method of dividing the image into grids is a design matter, and the number and size of the grids shown in the figure are merely examples.

[0032] According to the data shown in Fig. 2, the "second grid from the left, third grid from the bottom," "fourth grid from the right, fourth grid from the top," and "second grid from the right, third grid from the top" are identified as grids in which the center positions of the human bodies on the image are located. When an image including multiple human bodies is input as shown in Fig. 3, the grid in which the center positions of each of the multiple human bodies on the image are located is identified.

[0033] The "correction amount of human body position" data shown in FIG. 2 is data indicating the amount of movement in the x direction and the amount of movement in the y direction from the center of the grid identified as the center position of the human body on the image as shown in FIG. 3 to the center position of the human body on the image. The "correction amount of human body position" data is stored at the position of the grid identified as the center position of the human body on the image. As shown in FIG. 3, the center position of the human body on the image is located at a certain position within one grid. By using the likelihood of the human body position and the correction amount of the human body position, the center position of the human body on the image (the position of the human body) can be identified.

[0034] The "depth information of the human body" data shown in Fig. 2 is data that indicates the order in the depth direction in real space for the human body in the grid identified as the center position of the human body on the image as shown in Fig. 3. The "depth information of the human body" data is stored at the position of the grid identified as the center position of the human body on the image.

[0035] The depth direction is the direction indicating the front / back as seen from the camera. Alternatively, it may be the direction of the optical axis of the camera. The order is assigned to the human bodies shown in the image, and for example, the person closest to the camera in the image is numbered 0, and the order increases by one as the person moves further away from the camera. Using FIG. 3 as an example, person 1 is located closest to the camera in the real space, and they are lined up in order from the front toward the back of the camera: person 1, person 3, person 2. Therefore, the depth information of the human bodies is such that person 1 is numbered 0 (=i 1 ), Person 3 is 1 (=i 3 ), Person 2 is 2 (= i 2 )

[0036] The order may be normalized to a value between 0 and 1 by dividing so that the order of the farthest part in the image is 1. Furthermore, although the order uses a numerical value that increases by 1 from the front, a numerical value that reflects the distance between people in real space (or the apparent distance) may also be used. These orders allow correct answers to be generated visually from human body images, making it easy to collect learning data.

[0037] The "relative position of keypoint" data shown in Figures 15 and 16 is data indicating the relative position of each keypoint on the image with respect to the human body of the grid identified as the center position of the human body on the image as shown in Figure 3. Specifically, it is data indicating the amount of movement in the x direction and the y direction from the center of the grid identified as the center position of the human body on the image to the position of each keypoint corresponding to the human body on the image. The "relative position of keypoint" data is stored at the position of the grid identified as the center position of the human body on the image. As shown in Figure 3, the center position of the human body on the image is located at a certain position within one grid. The position of each keypoint on the human body on the image can be identified by using the likelihood of the human body position and the relative position of each keypoint.

[0038] The "relative depth information of key points" data shown in Figures 15 and 16 is data indicating the relative position in the depth direction, with respect to the center position of the human body in real space, for each key point corresponding to the human body in the grid identified as the center position of the human body in the image as shown in Figure 3. The "relative depth information of key points" data is stored at the grid position identified as the center position of the human body in the image. The depth direction refers to the direction indicating the front / back as seen from the camera. Alternatively, it may be the direction of the optical axis of the camera. The relative position in the depth direction, for example, is set to a negative value if the key point is located in front of the center position of the human body in real space, a positive value if the key point is located in the back, and 0 if the key point is located at the center position of the human body in real space. Using Figure 4 as an example, key point a of person 1 is located in front of the camera, close to the camera, with respect to the center position of the human body in real space. Key point b of person 1 is located in the back, far from the camera, with respect to the center position of the human body in real space. Therefore, the relative depth information of the keypoints is that the keypoint a of person 1 is -1 (=d a1 ), Person 1's key point b is 1 (=d b1) In Figure 4, the relative depth information of the keypoints is expressed as three values, -1, 0, and 1, but it is also possible to use a numerical value that directly reflects the position (or the apparent position) of the human body from the center position in real space. These relative positions in the depth direction can be visually generated from the human body image, making it easy to collect learning data.

[0039] Although FIG. 3 shows the positions of two key points for each person, the number of key points can be three or more.

[0040] In the technology of this embodiment, after outputting the above-mentioned multiple data from an input image, the parameters of the estimation model are calculated (learned) by minimizing the value of a predetermined loss function based on the multiple data and a given correct label.

[0041] During estimation, the grid where the center position of each human body on the image is located is identified based on the "likelihood of human body position" data shown in Fig. 2, and the correction amount corresponding to the identified grid position is obtained from the "correction amount of human body position" data shown in Fig. 2. Based on the identified grid position (center position of the grid) and the obtained correction amount, the center position of each human body on the image is identified.

[0042] Next, depth information corresponding to the identified grid positions is obtained from the "depth information of human body" data shown in Fig. 2. The depth direction order of each human body in real space is identified using the obtained depth information. Next, the relative position of each key point corresponding to the identified grid position is obtained from the "relative position of each key point" data shown in Fig. 2. Based on the identified grid position (center position of the grid) and the obtained relative position of each key point, the position of each key point on the image for each human body is identified.

[0043] Next, the relative depth information of each keypoint corresponding to the identified grid position is obtained from the data of "relative depth information of each keypoint" shown in Figures 15 and 16. From the obtained relative depth information of each keypoint, the relative position of each keypoint on each human body in the depth direction with respect to the center position of the human body in real space is identified.

[0044] As described above, during estimation, the grid in which the center position of each human body on the image is located is identified, and based on the position of the identified grid, real-space position information of the human body's key points is identified, which is indicated by real-space position information of the human body (the center position of each human body on the image and the depth-wise order of each human body in real space) and real-space position information of the key points associated with the human body (the position of each key point on each human body on the image and the relative position in the depth direction based on the center position of each key point on each human body in real space).

[0045] Furthermore, by having the above-mentioned features, the technology of this embodiment can easily learn / construct an estimation model in an apparatus that uses a trained estimation model to estimate position information in real space at key points on a human body from an image.

[0046] <Functional Configuration> Next, the functional configuration of the learning device 10 of this embodiment will be described. An example of a functional block diagram of the learning device 10 is described in Figure 5. As shown in the figure, the learning device 10 has an acquisition unit 11, a learning unit 12, and a storage unit 13. Note that, as shown in the functional block diagram of Figure 6, the learning device 10 does not have to have the storage unit 13. In this case, an external device configured to be able to communicate with the learning device 10 has the storage unit 13.

[0047] The acquisition unit 11 acquires training data in which teacher images and correct labels are linked. The teacher images include a person. The teacher images may include only one person or multiple people. The correct labels indicate at least the position of each keypoint on the human body on the image, the relative position of each keypoint on the human body in the depth direction based on the center position of the human body in real space, the center position of the human body on the image, and the order of the human body in the depth direction. The center position of the human body on the image may be calculated from the position of each keypoint on the human body on the image. For example, it may be the center of a rectangle that encompasses the positions of each keypoint on the human body, or the center of gravity using the positions of each keypoint on the human body. The correct labels may also be new correct labels obtained by processing the above-mentioned correct labels. For example, it may be correct labels of the multiple data shown in FIG. 1 that are processed from the above-mentioned correct labels.

[0048] For example, the operator creating the correct labels may perform tasks such as specifying positions in the image for the "position of each key point on the human body on the image" and the "center position of the human body on the image," which are the correct labels. Also, the operator creating the correct labels may perform tasks such as specifying relative positions and orders for the "relative positions in the depth direction of each key point on the human body based on the center position of the human body in real space" and the "order of the human body in the depth direction," which are the correct labels, in accordance with how the person appears in the image.

[0049] Here, the key points may be at least a part of a joint, a predetermined part (eyes, nose, mouth, navel, etc.), or an extremity of the body (tip of the head, fingertips, toes, etc.). Alternatively, the key points may be other parts. There are various ways to define the number and positions of key points, and there are no particular limitations.

[0050] For example, a large amount of learning data is stored in the storage unit 13. The acquisition unit 11 can acquire the learning data from the storage unit 13.

[0051] The learning unit 12 learns an estimation model based on the training data. The memory unit 13 stores the estimation model. The estimation model is configured with the neural network described with reference to FIG. 1. The estimation model outputs a plurality of data shown in FIG. 1. The plurality of data shown in FIG. 1 indicates information necessary for determining position information of a human body in real space and position information of key points associated with the human body in real space, and indicates the likelihood of the human body position, the amount of correction of the human body position, depth information of the human body, the relative position of each key point, and relative depth information of each key point. Details regarding the plurality of data are described above.

[0052] Then, various estimation processes can be performed using the multiple pieces of data output by the estimation model. For example, an estimation device (e.g., the estimation device 20 described in the following embodiment) obtains multiple pieces of data from the estimation model as described using FIGS. 1 to 3, 15, and 16. Using the multiple pieces of data obtained by the estimation model, the estimation device can estimate position information of human bodies in real space (for each human body, the center position of the human body on the image and the order of each human body in the depth direction in real space) and position information of key points associated with the human bodies in real space (for each human body, the position of each key point on the image and the relative position of each key point in the depth direction with respect to the center position of the human body in real space).

[0053] For example, the estimation device identifies the center position of each human body on the image based on the likelihood of the human body position and the correction amount of the human body position shown in Fig. 2. The estimation device also identifies the order of each human body in the depth direction in real space based on the likelihood of the human body position and depth information of the human body shown in Fig. 2. The estimation device also identifies the position of each key point on the image for each human body based on the likelihood of the human body position and the relative position of each key point shown in Fig. 2, Fig. 15, and Fig. 16. The estimation device also identifies the relative position of each key point in the depth direction with respect to the center position of the human body in real space based on the likelihood of the human body position and relative depth information of each key point shown in Fig. 2, Fig. 15, and Fig. 16.

[0054] When training the estimation model, the learning unit 12 can train (adjust) parameters of the estimation model so as to minimize errors between the plurality of data output from the training estimation model and the plurality of data in the training data (correct labels). Additionally, the learning unit 12 can train all lattices for the "likelihood of human body position" data. Furthermore, the learning unit 12 can train only the lattice where the center position of the human body on the image in the training data is located for the "correction amount of human body position," "depth information of human body," "relative position of each key point," and "relative depth information of each key point" data.

[0055] Here, a specific example of the learning method used by the learning unit 12 will be described.

[0056] For the "likelihood of human body position" data, the learning unit 12 can learn (adjust) the parameters of the estimation model so as to minimize the error between the map showing the likelihood of human body position output from the estimation model being trained and the map showing the likelihood of human body position of the training data (correct answer label) for all grid positions.

[0057] Furthermore, for the data on "amount of correction of human body position," the learning unit 12 can learn (adjust) the parameters of the estimation model so as to minimize the error between the amount of correction of human body position output from the estimation model being learned and the amount of correction of human body position in the learning data (correct label) only for the grid position where the center position of the human body on the image in the learning data is located.

[0058] Furthermore, for the data of "depth information of the human body," the learning unit 12 can learn (adjust) the parameters of the estimation model so as to minimize the error between the depth information of the human body output from the estimation model being trained and the depth information of the human body in the training data (correct label) only for the grid positions where the center position of the human body on the image in the training data is located.

[0059] Furthermore, for the data on the "relative position of each key point," the learning unit 12 can learn (adjust) the parameters of the estimation model so as to minimize the error between the relative position of each key point output from the estimation model being trained and the relative position of each key point in the training data (correct label) only for the grid positions where the center position on the image of the human body in the training data is located.

[0060] Furthermore, for the data of "relative depth information of each key point," the learning unit 12 can learn (adjust) the parameters of the estimation model so as to minimize the errors between the relative depth information of each key point output from the estimation model being trained and the relative depth information of each key point in the training data (correct label) only for the grid positions where the center position of the human body on the image in the training data is located.

[0061] An example of the processing flow of the learning device 10 will be described with reference to FIG.

[0062] In S10, the learning device 10 acquires learning data in which teacher images and correct labels are associated with each other. This process is realized by the acquisition unit 11. Details of the process executed by the acquisition unit 11 are as described above.

[0063] In S11, the learning device 10 learns an estimation model using the learning data acquired in S10. This process is realized by the learning unit 12. Details of the process executed by the learning unit 12 are as described above.

[0064] The learning device 10 repeats the loop of S10 and S11 until a termination condition is satisfied. The termination condition is defined using, for example, the value of a loss function.

[0065] <Hardware Configuration> Next, an example of the hardware configuration of the learning device 10 will be described. Each functional unit of the learning device 10 is realized by any combination of hardware and software, centered around a central processing unit (CPU) of any computer, memory, programs loaded into memory, a storage unit such as a hard disk that stores the programs (this can store programs that are pre-loaded when the device is shipped, as well as programs downloaded from recording media such as CDs (Compact Discs) or servers on the Internet), and a network connection interface. Those skilled in the art will understand that there are many variations in the implementation methods and devices.

[0066] FIG. 17 is a block diagram illustrating an example of the hardware configuration of a learning device 10. As shown in FIG. 17, the learning device 10 has a processor 1A, a memory 2A, an input / output interface 3A, a peripheral circuit 4A, and a bus 5A. The peripheral circuit 4A includes various modules. The learning device 10 does not necessarily have to have the peripheral circuit 4A. Note that the learning device 10 may be composed of multiple devices that are physically and / or logically separated. In this case, each of the multiple devices can have the above hardware configuration.

[0067] The bus 5A is a data transmission path for the processor 1A, memory 2A, peripheral circuit 4A, and input / output interface 3A to transmit and receive data to and from each other. The processor 1A is, for example, a processing unit such as a CPU or a graphics processing unit (GPU). The memory 2A is, for example, a random access memory (RAM) or a read-only memory (ROM). The input / output interface 3A includes an interface for acquiring information from an input device, an external device, an external server, an external sensor, a camera, etc., and an interface for outputting information to an output device, an external device, an external server, etc. Examples of input devices include a keyboard, a mouse, a microphone, physical buttons, a touch panel, etc. Examples of output devices include a display, a speaker, a printer, a mailer, etc. The processor 1A can issue commands to each module and perform calculations based on the results of those calculations.

[0068] <Effects> The estimation model learned by the learning device 10 of this embodiment has the feature of outputting multiple data, namely, "likelihood of human body position," "correction amount of human body position," "depth information of human body," "relative position of each key point," and "relative depth information of each key point."

[0069] Then, by using the multiple data output from the estimation model, it is possible to obtain information similar to the positional information of the key points of the human body in real space, such as "positional information of the human body in real space (position coordinates on the image of the human body and the order of the human body in the depth direction in real space)" and "positional information of the key points linked to the human body in real space (position coordinates on the image of the key points linked to the human body and the relative position in the depth direction with the center of the human body as the reference point)."

[0070] Furthermore, the multiple data sets output from the estimation model can be easily collected / generated manually as long as there are human body images. Therefore, a sufficient amount and variety of training data can be easily collected / generated. According to the learning device 10, a device that uses a trained estimation model to estimate real-world position information of key points on a human body from an image can easily learn / construct an estimation model trained with a sufficient amount and variety of training data.

[0071] Furthermore, the learning device 10 of this embodiment can use a trained estimation model to estimate the real-space position information of a human body and the real-space position information of key points linked to the human body from processed images, and can easily collect / generate learning data for the estimation model using only images without any special equipment. Furthermore, even if multiple people appear in an image, it can estimate each key point (each joint point), which is three-dimensional skeletal information, while linking it to each person.

[0072] Fourth Embodiment An estimation device 20 of this embodiment uses an estimation model trained by the learning device 10 of the third embodiment to estimate position information of a human body in real space (position coordinates of the human body on an image and the order of the human body in the depth direction in real space) and position information of key points associated with the human body in real space (position coordinates of the key points associated with the human body on an image and relative positions of the key points associated with the human body in the depth direction with the center of the human body as the reference). This will be described in detail below.

[0073] Fig. 8 shows an example of a functional block diagram of the estimation device 20. As shown in the figure, the estimation device 20 has an estimation unit 21 and a storage unit 22. Note that, as shown in the functional block diagram of Fig. 9, the estimation device 20 does not have to have the storage unit 22. In this case, an external device configured to be able to communicate with the estimation device 20 has the storage unit 22.

[0074] The estimation unit 21 acquires any image as the processed image. For example, the estimation unit 21 may acquire an image captured by a camera or an image from a stored video as the processed image.

[0075] Then, using the estimation model learned by the learning device 10, the estimation unit 21 estimates and outputs position information of the human body in real space (position coordinates of the human body on the image and the order of the human body in the depth direction in real space) and position information of the key points linked to the human body in real space (position coordinates of the key points linked to the human body on the image and the relative positions of the key points linked to the human body in the depth direction with the center of the human body as the reference).

[0076] As described in the third embodiment, when an image is input, the estimation model outputs the data described with reference to FIGS. 1 to 3 , 15 , and 16 . The estimation unit 21 performs further estimation processing using the data output by the estimation model to estimate position information of the human body in real space (position coordinates of the human body on the image and the order of the human body in the depth direction in real space) and position information of key points associated with the human body in real space (position coordinates of the key points associated with the human body on the image and the relative positions of the key points associated with the human body in the depth direction with respect to the center of the human body), and outputs the estimation results. The trained estimation model is stored in the storage unit 22. The estimation results can be output using any means, such as a display, a projection device, a printer, or email. The estimation unit 21 may also output the data output by the estimation model as the estimation result as is.

[0077] An example of the process performed by the estimation unit 21 will be described below with reference to FIGS. 10 and 11. FIG.

[0078] (Step 1): The processed image is processed by the estimation model to obtain a plurality of data as shown in FIGS. 1 to 3, 15 and 16.

[0079] (Step 2): Based on the "likelihood of human body position" data, the grid (P2 in FIG. 10) in which the center position (P1 in FIG. 10) of each person (each human body) on the image is located (is included) is identified. Specifically, the grid whose likelihood is equal to or greater than a threshold is identified. Furthermore, the center position of the grid (P3 in FIG. 10) is identified from the identified grid.

[0080] (Step 3): From the data on "correction amount of human body position," obtain the correction amount (P4 in FIG. 10) corresponding to the grid position identified in (Step 2).

[0081] (Step 4): Based on the center position of the grid identified in (Step 2) and the correction amount obtained in (Step 3), the coordinates of the center position of each person in the processed image (P1 in FIG. 10) are identified. This identifies the position coordinates of each human body on the image.

[0082] (Step 5): From the "depth information of the human body" data, obtain the depth information corresponding to the grid positions identified in (Step 2), i.e., the order in the depth direction in real space. This identifies the order in the depth direction of each human body in real space.

[0083] (Step 6): From the data on the "relative position of each keypoint," obtain the relative position (P6 in FIG. 11) corresponding to the grid position identified in (Step 2).

[0084] (Step 7): Based on the center position of the grid identified in (Step 2) and the relative position acquired in (Step 6), the position coordinates of each key point on the image (P7 in FIG. 11) are identified for each person included in the processed image. This identifies the position coordinates on the image of each key point associated with each human body.

[0085] (Step 8): From the data of "relative depth information of each keypoint," obtain relative depth information corresponding to the grid position identified in (Step 2), i.e., the relative position in the depth direction with respect to the center position of the human body in real space (P8 in FIG. 11). This identifies the relative position in the depth direction with respect to the center of the human body of each keypoint associated with each human body.

[0086] (Step 9): Output the position coordinates on the image of each human body identified in (Step 4), the order in the depth direction in real space of each human body identified in (Step 5), the position coordinates on the image of each key point associated with each human body identified in (Step 7), and the relative position in the depth direction with respect to the center of the human body of each key point associated with each human body identified in (Step 8).

[0087] As a result, the estimation unit 21 can use the estimation model learned by the learning device 10 to estimate and output, from the image, position information of the human body in real space (position coordinates of the human body on the image and the order of the human body in the depth direction in real space) and position information of the key points linked to the human body in real space (position coordinates of the key points linked to the human body on the image and the relative positions of the key points linked to the human body in the depth direction with the center of the human body as the reference).

[0088] The estimation unit 21 can superimpose and display the estimated information on the image used for estimation, as shown in Fig. 12. The estimation unit 21 can superimpose and display the order of the estimated human bodies in the depth direction in real space (P13 in Fig. 12) corresponding to the human bodies on a position based on the position coordinates on the image of each estimated human body (P11 in Fig. 12) or the position coordinates on the image of each key point associated with each estimated human body (P12 in Fig. 12).

[0089] Furthermore, the estimation unit 21 can superimpose and display an object indicating the key point on the position coordinates on the image of each key point associated with each estimated human body (P12 in FIG. 12 ). The estimation unit 21 can then change the color (or shape or size) of the object to correspond to the value of the estimated relative position in the depth direction corresponding to the key point. As described above, superimposing information about depth makes it easier to understand visually and intuitively information about depth that is difficult to express on an image.

[0090] In addition, the above display method can also be used as support when manually generating correct labels (training data) necessary for training an estimation model. When manually inputting correct labels on an image, it is difficult to represent the correct labels in the depth direction on the image, leading to errors in labeling. Therefore, by using the above display method, it is possible to reduce errors in labeling. When manually inputting the correct level in the depth direction, by sequentially displaying the input status on the image using the above display method, the state of the correct labels in the depth direction can be visually understood even on the image, reducing errors in label input.

[0091] Next, an example of the processing flow of the estimation device 20 will be described with reference to the flowchart of FIG.

[0092] In S20, the estimating device 20 acquires a processed image. For example, an operator inputs a processed image to the estimating device 20. Then, the estimating device 20 acquires the input processed image.

[0093] In S21, the estimation device 20 estimates, from the processed image, position information of the human body in real space (position coordinates of the human body on the image and the order of the human body in the depth direction in real space) and position information of the key points associated with the human body in real space (position coordinates of the key points associated with the human body on the image and relative positions of the key points associated with the human body in the depth direction with the center of the human body as the reference), using the estimation model learned by the learning device 10. This processing is realized by the estimation unit 21. Details of the processing executed by the estimation unit 21 are as described above.

[0094] In S22, the estimation device 20 outputs the estimation result of S21. The estimation device 20 can use any means such as a display, a projection device, a printer, or email.

[0095] Next, an example of the hardware configuration of the estimation device 20 will be described. Each functional unit of the estimation device 20 is realized by any combination of hardware and software, centered around a CPU of any computer, memory, a program loaded into the memory, a storage unit such as a hard disk that stores the program (which can store programs pre-loaded at the time of shipping the device, as well as programs downloaded from a recording medium such as a CD or a server on the Internet), and a network connection interface. Those skilled in the art will understand that there are various variations in the implementation method and device. Figure 17 is a block diagram illustrating an example of the hardware configuration of the estimation device 20.

[0096] According to the estimation device 20 of the present embodiment described above, it is possible to estimate, from a processed image, the position information of a human body in real space and the position information of key points linked to the human body in real space using an estimation model trained by the learning device 10 of the third embodiment. With this estimation device 20, it is possible to easily collect and generate learning data for an estimation model using only images, without the need for special equipment, and it is possible to easily train and construct an estimation model. Furthermore, even if multiple people appear in an image, it is possible to estimate each key point (each joint point), which is three-dimensional skeletal information, while linking it to each person.

[0097] Fifth Embodiment Next, a fifth embodiment will be described in detail with reference to the drawings.

[0098] Referring to FIG. 14, in the fifth embodiment, a computer-readable storage medium 102 storing a three-dimensional skeleton estimation program 101 is connected to a computer 100 .

[0099] The computer-readable storage medium 102 is composed of a magnetic disk, semiconductor memory, etc., and the 3D skeleton estimation program 101 stored therein is read by the computer 100 when the computer 100 is started up, etc., and controls the operation of the computer 100, causing the computer 100 to function as each of the functional units 11, 12, and 13 within the learning device 10 in the first and third embodiments described above, and to perform the processing shown in Figure 7.

[0100] In this embodiment, the learning device 10 according to the first and third embodiments is realized by a computer and a program, but the estimation device 20 according to the second and fourth embodiments can also be realized by a computer and a program in a similar manner.

[0101] "Industrial Applicability" The first to fifth embodiments described above can be adapted for the following applications: A three-dimensional skeletal estimation device that can estimate, from an image, positional information of a human body in real space and positional information of key points linked to the human body in real space; A three-dimensional skeletal estimation device that can estimate three-dimensional skeletal information by linking each key point (each joint point) to each person, even if multiple people appear in the image; A three-dimensional skeletal estimation device that can easily learn / construct an estimation model in a device that estimates positional information of key points on a human body in real space from an image using a trained estimation model; and A program for implementing these three-dimensional skeletal estimation devices on a computer.

[0102] Furthermore, the above-described first to fifth embodiments can be adapted to applications such as devices and functions that perform image recognition that require estimation of real-space position information of a human body and real-space position information of key points linked to the human body from camera / stored video.

[0103] Furthermore, the first to fifth embodiments described above can be adapted to applications such as devices and functions for analyzing behavior in the marketing and monitoring fields.

[0104] Furthermore, the above-described first to fifth embodiments can be applied to applications such as an input interface that receives as input position information of a human body in real space estimated from a camera or stored video, and position information of key points associated with the human body in real space.

[0105] In addition, the above-described first to fifth embodiments can be applied to applications such as video / image search devices and functions that use estimated real-space position information of a human body and real-space position information of key points associated with the human body as trigger keys.

[0106] Although the embodiments of the present invention have been described above with reference to the drawings, these are merely examples of the present invention, and various other configurations may be adopted. The configurations of the above-described embodiments may be combined with each other, or some of the configurations may be replaced with other configurations. Furthermore, various modifications may be made to the configurations of the above-described embodiments without departing from the spirit of the invention. Furthermore, the configurations and processes disclosed in the above-described embodiments and modified examples may be combined with each other.

[0107] In addition, although the flowcharts used in the above description show multiple steps (processes) in a sequential order, the order of steps executed in each embodiment is not limited to the order shown. In each embodiment, the order of the steps shown in the drawings can be changed as long as it does not cause any problems in terms of the content. Furthermore, the above-described embodiments can be combined as long as the content is not contradictory.

[0108] In this specification, "acquisition" includes at least one of the following: "the device retrieves data stored in another device or storage medium (active acquisition)" based on user input or program instructions, such as receiving data by making a request or inquiry to another device, or accessing and reading out another device or storage medium; "inputting data output from another device into the device (passive acquisition)" based on user input or program instructions, such as receiving data that is distributed (or transmitted, push notification, etc.), and selecting and acquiring data from received data or information; and "editing data (converting data to text, rearranging data, extracting some data, changing the file format, etc.) to generate new data and acquire the new data."

[0109] Some or all of the above embodiments may be described as, but are not limited to, the following supplementary notes: 1. A learning device comprising: an acquisition means for acquiring learning data linking teacher images including human bodies with correct labels indicating the center position of each human body on the image, correct labels indicating the position of each key point on each human body on the image, and correct labels indicating the relative position of each key point on each human body in the depth direction; and a learning means for training an estimation model that estimates, based on the learning data, information indicating the likelihood of the center position on the image for each human body, the relative position on the image for each key point linked to each human body, and the relative position in the depth direction for each key point linked to each human body. 2. The learning device described in 1, characterized in that correct labels indicating position information of each human body in the depth direction are added and linked to the learning data, and the estimation model estimates by adding the relative position on the image indicating the center position of each human body on the image for each human body, and position information of each human body in the depth direction. 3. The learning device according to claim 2, wherein the depth direction position information is the depth direction order of a person in the image in real space. 4. The learning device according to any one of claims 1 to 3, wherein the relative positions on the image are relative positions on the image based on the center of a grid identified as the location of the center position of the human body on the image in the information indicating the likelihood, and the relative positions of each keypoint in the depth direction are relative positions in the real space in the depth direction based on the center position of the human body in real space. 5. The learning device according to claim 4, wherein the relative positions of each keypoint in the depth direction are set to a negative value if the keypoint is located in the foreground in the image, a positive value if the keypoint is located behind the center position of the human body in real space, and 0 if the keypoint is located at the center position of the human body in real space.6. The learning means, based on the estimation model during learning, estimates information indicating the likelihood of the central position of each human body on the image, the relative position of each key point associated with each human body on the image, the relative position in the depth direction of each key point associated with each human body based on the central position of the human body in real space, the relative position on the image indicating the central position of each human body on the image, and the order of each human body in the depth direction, adjusts parameters of the estimation model so as to minimize, for all grid positions, an error between the estimated result of the information indicating the likelihood of the central position of each human body on the image and information obtained from the correct label indicating the likelihood of the central position of each human body on the image, and adjusts parameters of the estimation model so as to minimize, for only grid positions where the central position of the human body on the image is located in the learning data, 6. The learning device according to claim 4 or 5, wherein parameters of the estimation model are adjusted so as to minimize an error between an estimation result of a relative position in the depth direction of each key point associated with each human body, with reference to a center position of the human body in real space, and a relative position in the depth direction of each key point associated with each human body obtained from the correct label, with reference to a center position of the human body in real space, only for grid positions where the center positions of the human body on the image are located in the training data; adjust parameters of the estimation model so as to minimize an error between an estimation result of a relative position on the image indicating the center position of each human body on the image of the human body, and a relative position on the image indicating the center position of each human body on the image of the human body obtained from the correct label, only for grid positions where the center positions of the human body on the image are located in the training data; and adjust parameters of the estimation model so as to minimize an error between an estimation result of an order in the depth direction of each human body and the order in the depth direction of each human body obtained from the correct label, only for grid positions where the center positions of the human body on the image are located in the training data. 7. The learning device according to any one of 1 to 6, wherein the depth direction is the direction of the optical axis of the camera.8. An estimation device having an estimation means for estimating, using an estimation model trained by the learning device according to any one of 1 to 7, the position coordinates on an image of each key point associated with each human body, the relative position in the depth direction of each key point associated with each human body with reference to the center position of the human body in real space, the position coordinates indicating the center position of each human body on the image, and the order of each human body in the depth direction, the estimation means: identifies a grid in which the center position of each human body on the image is located, based on information indicating the likelihood of the center position of each human body on the image obtained from the estimation model; acquires a relative position on the image corresponding to the identified grid position from the relative position on the image of each key point associated with each human body obtained from the estimation model, and estimates position coordinates on the image of each key point associated with each human body based on the identified grid center position and the relative position on the acquired image; acquires a relative position in the depth direction corresponding to the identified grid position from the relative position in the depth direction of each key point associated with each human body obtained from the estimation model, with the center position of the human body in real space as the reference, and estimates the relative position in the depth direction of each key point associated with each human body based on the center position of the human body in real space; acquires a relative position on the image corresponding to the identified grid position from the relative position on the image indicating the center position of each human body on the image obtained from the estimation model, and estimates position coordinates indicating the center position of each human body on the image of each human body based on the identified grid center position and the relative position on the acquired image; 10. The estimation device according to claim 8, wherein the estimation means acquires a depth direction order corresponding to the specified grid position from the depth direction order of each human body obtained from the estimation model, and estimates the depth direction order of each human body. 11. The estimation device according to claim 8 or 9, wherein the estimation means superimposes and displays the depth direction order of each estimated human body corresponding to the human body at a position based on position coordinates indicating a center position on the image of each estimated human body for each estimated human body, or position coordinates on the image of each key point associated with each estimated human body, in the image used for estimation.11. The estimation device according to any one of 8 to 10, wherein the estimation means displays an object indicating a key point superimposed on the position coordinates on the image of each key point associated with each estimated human body in the image used for estimation, and sets the color, shape, or size of the object to a value corresponding to the relative position in the depth direction based on the center position in real space of the human body of each key point associated with each estimated human body corresponding to that key point. 12. The estimation device according to any one of 8 to 11, wherein the depth direction is the direction of the optical axis of a camera. 13. 13. A learning method in which one or more computers acquire learning data linking teacher images including human bodies with correct labels indicating the center position of each human body on the image, correct labels indicating the position of each key point for each human body on the image, and correct labels indicating the relative position of each key point for each human body in the depth direction, and learn an estimation model that estimates information indicating the likelihood of the center position of each human body on the image, the relative position on the image of each key point linked to each human body, and the relative position in the depth direction of each key point linked to each human body, based on the learning data. 14. A program that causes a computer to function as: an acquisition means that acquires training data linking a teacher image including a human body with a correct label indicating the center position of each human body on the image, a correct label indicating the position of each key point for each human body on the image, and a correct label indicating the relative position of each key point for each human body in the depth direction, and a learning means that learns an estimation model that estimates, based on the training data, information indicating the likelihood of the center position of each human body on the image, the relative position of each key point associated with each human body on the image, and the relative position of each key point associated with each human body in the depth direction. 15. An estimation method in which one or more computers use an estimation model trained by the learning device described in any one of 1 to 7 to estimate position coordinates on the image of each key point associated with each human body, the relative position in the depth direction of each key point associated with each human body based on the center position of the human body in real space, position coordinates indicating the center position of each human body on the image, and the order of each human body in the depth direction.16. A program causing a computer to function as estimation means for estimating, using an estimation model trained by the learning device described in any one of 1 to 7, the position coordinates on an image of each key point associated with each human body, the relative position in the depth direction of each key point associated with each human body based on the center position of the human body in real space, the position coordinates indicating the center position of each human body on the image, and the order of each human body in the depth direction,

[0110] This application claims priority based on Japanese Patent Application No. 2022-200103, filed December 15, 2022, the disclosure of which is incorporated herein in its entirety.

[0111] REFERENCE SIGNS LIST 10 Learning device 11 Acquisition unit 12 Learning unit 13 Storage unit 20 Estimation device 21 Estimation unit 22 Storage unit 100 Computer 101 Three-dimensional skeleton estimation program 102 Computer-readable storage medium 1A Processor 2A Memory 3A Input / output I / F 4A Peripheral circuit 5A Bus

Claims

1. an acquisition means for acquiring learning data that associates teacher images including human bodies with correct labels indicating the center positions of each human body on the image, correct labels indicating the positions of each key point on each human body on the image, and correct labels indicating the relative positions of each key point on each human body in the depth direction; a learning means for learning an estimation model that estimates, based on the learning data, information indicating the likelihood of a center position of each human body on an image, a relative position on the image of each key point associated with each human body, and a relative position in the depth direction of each key point associated with each human body; A learning device having the above configuration.

2. In the learning data, a correct answer label indicating position information of each human body in the depth direction is added and linked; the estimation model is estimated by adding relative position information on the image indicating the center position of each human body on the image and position information in the depth direction of each human body; The depth direction position information is an order of the person in the image in the depth direction in real space.

2. The learning device according to claim 1 .

3. the relative position on the image is a relative position on the image based on a center of a grid identified as a location of a center position of the human body on the image in the information indicating the likelihood, The relative position of each key point in the depth direction is a relative position in the depth direction in real space based on the center position of the human body in real space.

3. The learning device according to claim 1 or 2.

4. An estimation means for estimating the position coordinates on the image of each key point associated with each human body, the relative position in the depth direction of each key point associated with each human body using the estimation model learned by a learning device having: an acquisition means for acquiring learning data linking teacher images including human bodies with correct labels indicating the center position of each human body on the image, correct labels indicating the position of each key point on each human body on the image, and correct labels indicating the relative position in the depth direction of each key point on each human body; and a learning means for learning an estimation model that estimates, based on the learning data, information indicating the likelihood of the center position on the image of each human body, the relative position on the image of each key point associated with each human body, and the relative position in the depth direction of each key point associated with each human body, using the estimation model learned by the learning device; An estimation device having:

5. the estimation means superimposes and displays the order of the estimated human bodies in the depth direction corresponding to the human bodies on the image used for estimation at a position based on position coordinates indicating the center position of the estimated human body on the image, or position coordinates on the image of each key point associated with the estimated human body; 5. The estimation device according to claim 4.

6. the estimation means displays an object indicating a key point superimposed on the position coordinates on the image of each key point associated with each estimated human body in the image used for estimation, and sets the color, shape or size of the object to a value corresponding to a relative position in the depth direction based on a center position in real space of each key point associated with each estimated human body corresponding to the key point; 6. The estimation device according to claim 4 or 5.

7. One or more computers Acquire learning data that associates teacher images containing human bodies with correct labels indicating the center positions of each human body on the image, correct labels indicating the positions of each key point on each human body on the image, and correct labels indicating the relative positions in the depth direction of each key point on each human body, Based on the learning data, an estimation model is trained to estimate information indicating the likelihood of a center position of each human body on an image, a relative position on the image of each key point associated with each human body, and a relative position in the depth direction of each key point associated with each human body. How to learn.

8. Computer, an acquisition means for acquiring learning data that associates teacher images including human bodies with correct labels indicating the center positions of each human body on the image, correct labels indicating the positions of each key point on each human body on the image, and correct labels indicating the relative positions in the depth direction of each key point on each human body; a learning means for learning an estimation model that estimates, based on the learning data, information indicating the likelihood of a center position of each human body on an image, a relative position on the image of each key point associated with each human body, and a relative position in the depth direction of each key point associated with each human body; A program that functions as a

9. One or more computers a learning device including: an acquisition means for acquiring learning data linking teacher images including human bodies with correct labels indicating the center position of each human body on the image, correct labels indicating the position of each key point on each human body on the image, and correct labels indicating the relative position of each key point on each human body in the depth direction; and a learning means for learning an estimation model that estimates, based on the learning data, information indicating the likelihood of the center position on the image of each human body, the relative position on the image of each key point linked to each human body, and the relative position in the depth direction of each key point linked to each human body, using the estimation model trained by the learning device to estimate position coordinates on the image of each key point linked to each human body, the relative position in the depth direction of each key point linked to each human body with respect to the center position in real space of the human body, position coordinates indicating the center position of each human body on the image, and the order of each human body in the depth direction; Estimation method.

10. Computer, an estimation means for estimating position coordinates on the image of each key point associated with each human body, the relative position in the depth direction of each key point associated with each human body with respect to the center position of the human body in real space of each key point associated with each human body, position coordinates indicating the center position of each human body on the image, and the order of each human body in the depth direction using the estimation model trained by a learning device having: an acquisition means for acquiring training data linking teacher images including human bodies with correct labels indicating the center position of each human body on the image, correct labels indicating the position of each key point on each human body on the image, and correct labels indicating the relative position of each key point on each human body in the depth direction; and a learning means for training an estimation model that estimates, based on the training data, information indicating the likelihood of the center position of each human body on the image, the relative position on the image of each key point associated with each human body, and the relative position in the depth direction of each key point associated with each human body; A program that functions as a