Learning device, estimation device, learning method, estimation method, and program

The learning device and method address the challenge of collecting training data for real-space human body key point estimation by using ground truth labels, enabling accurate estimation of human body positions and key point positions in real space.

JP7852744B2Active Publication Date: 2026-04-28NEC CORP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NEC CORP
Filing Date
2023-12-07
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing methods face challenges in collecting sufficient and varied training data for estimating real-space position information of human body key points, as motion capture systems are limited to specific environments, leading to low estimation accuracy.

Method used

A learning device and method that acquires training data linking images with ground truth labels for human body positions and key point positions, enabling the construction of an estimation model to estimate central and relative positions in real space.

Benefits of technology

Facilitates easy collection and generation of training data, allowing for the accurate estimation of real-space position information of human body key points using a pre-trained estimation model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007852744000001
    Figure 0007852744000001
  • Figure 0007852744000002
    Figure 0007852744000002
  • Figure 0007852744000003
    Figure 0007852744000003
Patent Text Reader

Abstract

The present invention provides a learning device comprising: an acquiring unit for acquiring learning data in which teacher images including human bodies, a correct answer label indicating a central position, in the image, of each human body, a correct answer label indicating a position, in the image, of each keypoint on each human body, and a correct answer label indicating a relative position, in a depth direction, of each keypoint on each human body are associated with one another; and a training unit which, on the basis of the learning data, trains an estimation model for estimating information indicating the likelihood of a central position, in the image, of each human body, the relative position, in the image, of each keypoint associated with each human body, and the relative position, in the depth direction, of each keypoint associated with each human body.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a learning device, an estimation device, a learning method, an estimation method, and a program.

Background Art

[0002] A technique related to the present invention is disclosed in Non-Patent Document 1. The technique of Non-Patent Document 1 is used to estimate position information (3D skeleton information) in real space at key points of the human body (joint points / skeleton points of the human body) from an image using a learned estimation model.

[0003] In the prior art of Non-Patent Document 1, by inputting a single image into a learned estimation model composed of a convolutional neural network, the position coordinates in real space (3D) at key points of the human body (joint points / skeleton points of the human body) are estimated. For the learning data, data in which an image and the position coordinates in real space at the key points of the human body are paired is used.

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] The problem with Non-Patent Document 1 is that in the learning of the estimation model, learning data such as the position coordinates in real space (3D skeleton information) at key points of the human body is difficult to collect, and the estimation model cannot be easily learned.

[0006] The reason for this is that the position coordinates of key points in the human body in real space (3D), unlike position coordinates in images (2D), cannot be easily generated / collected manually with just a human body image. Training data cannot be collected without using large-scale equipment such as a motion capture system.

[0007] A further problem is the inability to collect training data with a wide variety of images that correspond to position coordinates in real space (3D).

[0008] The reason is that equipment like motion capture systems are installed in limited environments such as indoor laboratories due to their installation requirements, and the images captured have limited variations in background, number of people, depth, etc.

[0009] A further problem is that if the amount and variety of training data are insufficient, the estimation accuracy of the estimation model, that is, the accuracy of the process of estimating the real-space position information of key points of the human body from images, will be low.

[0010] One example of the object of the present invention is to provide a learning device, an estimation device, a learning method, an estimation method, and a program that solve any of the above-mentioned problems. [Means for solving the problem]

[0011] According to one aspect of the present invention, An acquisition means for acquiring training data that links a training image containing a human body with a ground truth label indicating the central position of each human body in the image, a ground truth label indicating the position of each key point in each human body in the image, and a ground truth label indicating the relative position of each key point in the depth direction in each human body. A learning means for learning an estimation model that, based on the aforementioned training data, estimates information indicating the likelihood of the central position on the image for each human body, the relative position on the image of each key point associated with each human body, and the relative position in the depth direction of each key point associated with each human body. A learning device having the following is provided.

[0012] According to one aspect of the present invention, Estimation means for estimating the position coordinates on the image of each key point associated with each human body, the relative position in the depth direction of each key point associated with each human body with respect to the center position of the human body in real space, the position coordinates indicating the center position of the human body on the image for each human body, and the order in the depth direction for each human body, using the estimation model learned by the learning device. An estimation device having the following is provided.

[0013] According to one aspect of the present invention, One or more computers, We obtain training data that links training images containing human bodies with ground truth labels indicating the central position of each human body in the image, ground truth labels indicating the position of each key point in each human body in the image, and ground truth labels indicating the relative position of each key point in the depth direction in each human body. Based on the aforementioned training data, an estimation model is trained to estimate information indicating the likelihood of the central position on the image for each human body, the relative position on the image of each key point associated with each human body, and the relative position in the depth direction of each key point associated with each human body. Learning methods are provided.

[0014] According to one aspect of the present invention, Computers, Acquisition means for acquiring training data that links a training image containing a human body with a ground truth label indicating the central position of each human body in the image, a ground truth label indicating the position of each key point in each human body in the image, and a ground truth label indicating the relative position in the depth direction of each key point in each human body. A learning means for learning an estimation model that estimates information indicating the likelihood of the central position on an image for each human body, the relative position on the image of each key point associated with each human body, and the relative position in the depth direction of each key point associated with each human body, based on the aforementioned training data. A program is provided to enable it to function as such.

[0015] According to one aspect of the present invention, one or more computers use the estimation model learned by the learning device to estimate the position coordinates on the image of each keypoint associated with each human body, the relative position in the depth direction based on the center position in the real space of the human body of each keypoint associated with each human body, the position coordinates indicating the center position on the image of the human body in each human body, and the order in the depth direction in each human body. An estimation method is provided.

[0016] According to one aspect of the present invention, a computer is caused to function as an estimation means for estimating the position coordinates on the image of each keypoint associated with each human body, the relative position in the depth direction based on the center position in the real space of the human body of each keypoint associated with each human body, the position coordinates indicating the center position on the image of the human body in each human body, and the order in the depth direction in each human body, using the estimation model learned by the learning device. A program is provided.

Advantages of the Invention

[0017] According to one aspect of the present invention, in an apparatus for estimating position information in the real space of keypoints of a human body from an image using a learned estimation model, the problem of easily learning / constructing an estimation model learned with a sufficient amount and variation of learning data, a learning apparatus, a learning method, and a program are solved.

[0018] Also, according to one aspect of the present invention, the problem of providing an estimation apparatus, an estimation method, and a program for accurately estimating position information in the real space of keypoints of a human body from an image is solved.

Brief Description of the Drawings

[0019] The above-described object, and other objects, features, and advantages will become more apparent from the following Suitable described embodiments and the accompanying drawings below.

[0020] [Figure 1] This is a diagram illustrating the technology of this embodiment. [Figure 2] This is a diagram illustrating the technology of this embodiment. [Figure 3] This is a diagram illustrating the technology of this embodiment. [Figure 4] This is a diagram illustrating the technology of this embodiment. [Figure 5] This is an example of a functional block diagram of the learning device of this embodiment. [Figure 6] This is an example of a functional block diagram of the learning device of this embodiment. [Figure 7] This flowchart shows an example of the processing flow of the learning device of this embodiment. [Figure 8] This is an example of a functional block diagram of the estimation device of this embodiment. [Figure 9] This is an example of a functional block diagram of the estimation device of this embodiment. [Figure 10] This is a diagram illustrating the processing of the estimation device of this embodiment. [Figure 11] This is a diagram illustrating the processing of the estimation device of this embodiment. [Figure 12] This is a diagram illustrating the technology of this embodiment. [Figure 13] This flowchart shows an example of the processing flow of the estimation device of this embodiment. [Figure 14] This is an example of a functional block diagram of the 3D skeleton estimation device of this embodiment. [Figure 15] This is a diagram illustrating the technology of this embodiment. [Figure 16] This is a diagram illustrating the technology of this embodiment. [Figure 17] This figure shows an example of the hardware configuration of the device according to this embodiment. [Modes for carrying out the invention]

[0021] Embodiments of the present invention will be described below with reference to the drawings. In all drawings, similar components are denoted by the same reference numerals, and their descriptions are omitted as appropriate.

[0022] "First Embodiment" Figure 6 is a functional block diagram showing an overview of the learning device 10 according to the first embodiment. The learning device 10 has an acquisition unit 11 and a learning unit 12. The acquisition unit 11 acquires training data that links a training image containing a human body with a ground truth label indicating the center position of each human body on the image, a ground truth label indicating the position of each key point in each human body on the image, and a ground truth label indicating the relative position in the depth direction of each key point in each human body. Based on the training data, the learning unit 12 learns an estimation model that estimates information indicating the likelihood of the center position of each human body on the image, the relative position in the image of each key point associated with each human body, and the relative position in the depth direction of each key point associated with each human body.

[0023] With a learning device 10 having such a configuration, a device that estimates the real-space position information of key points of the human body from an image using a pre-trained estimation model can easily train / construct an estimation model that has been trained with a sufficient amount and variety of training data.

[0024] "Second Embodiment" Figure 9 is a functional block diagram showing an overview of the estimation device 20 according to the second embodiment. The estimation device 20 has an estimation unit 21. The estimation unit 21 uses the estimation model learned by the learning device 10 described in the first embodiment to estimate the position coordinates on the image of each key point associated with each human body, the relative position in the depth direction of each key point associated with each human body with respect to the center position of the human body in real space, the position coordinates indicating the center position of the human body on the image for each human body, and the order in the depth direction for each human body.

[0025] According to the estimation device 20 with such a configuration, it is possible to estimate the real-space position information of key points of the human body from an image with high accuracy.

[0026] "Third Embodiment" <Overview> Instead of estimating "real-space position information of key points on the human body" from an image, the estimation device 20 of this embodiment estimates similar information: "real-space position information of the human body" and "real-space position information of key points associated with the human body."

[0027] "The positional information of the human body in real space" refers to the position coordinates of the human body on the image and the order of the human body in the depth direction in real space. Furthermore, "the positional information of keypoints linked to the human body in real space" refers to the position coordinates of the keypoints linked to the human body on the image and their relative positions in the depth direction relative to the center of the human body. This relative position in the depth direction, along with the corresponding order of the human body in the depth direction, allows for the identification of the keypoint's position in real space in the depth direction.

[0028] The learning device 10 of this embodiment learns a neural network (estimation model) that outputs information necessary to obtain the above information (related information to the above information). The above information, and the output information of the estimation model related to the above information, can be easily collected / generated manually as long as human body images are available. Therefore, the training data for the estimation model can be easily collected / generated. This makes it possible to easily learn / construct an estimation model in a device that estimates the real-space position information of key points of the human body from images.

[0029] <Technical Features of This Embodiment> The technology of this embodiment will now be described. As shown in Figure 1, when an image is input to a neural network, multiple data points as shown in the figure are output. In other words, the neural network of this embodiment consists of multiple layers that output multiple data points as shown in the figure.

[0030] Figure 2 shows an example of "likelihood of human body position," "correction amount of human body position," and "depth information of the human body" from the multiple data shown in Figure 1. Figure 15 shows an example of "relative position of keypoint a" and "relative depth information of keypoint a" from the multiple data shown in Figure 1. Figure 16 shows an example of "relative position of keypoint b" and "relative depth information of keypoint b" from the multiple data shown in Figure 1. Figure 3 shows the images from which the data in Figures 2, 15, and 16 were derived, with explanations of the concepts behind each of the data in Figures 2, 15, and 16 added.

[0031] The "likelihood of human body position" data shown in Figure 2 represents the likelihood of the central position of the human body on the image (the position of the human body), as shown in Figure 3. As illustrated, this data shows the likelihood that the central position of the human body on the image is located in each of the multiple grids obtained by dividing the image. The likelihood may be expressed using a normal distribution or the like, centered on the grid that shows the central position of the human body on the image. Note that the method of dividing the image into a grid is a design matter, and the number and size of the grids shown are merely examples.

[0032] According to the data shown in Figure 2, the "second grid from the left and third grid from the bottom," the "fourth grid from the right and fourth grid from the top," and the "second grid from the right and third grid from the top" are identified as grids where the central position of the human body on the image is located. As shown in Figure 3, when an image containing multiple human bodies is input, the grids where the central position of each of the multiple human bodies on the image is located are identified.

[0033] The "human body position correction amount" data shown in Figure 2 represents the x-direction movement and y-direction movement from the center of the grid where the central position of the human body on the image is identified to the central position of the human body on the image, as shown in Figure 3. The "human body position correction amount" data is stored at the grid position where the central position of the human body on the image is identified. As shown in Figure 3, the central position of the human body on the image exists at a certain position within a single grid. By using the likelihood of the human body position and the human body position correction amount, the central position of the human body on the image (the position of the human body) can be identified.

[0034] The "depth information of the human body" data shown in Figure 2 represents the order in the depth direction in real space for the human body in the grid where the central position of the human body in the image is identified, as shown in Figure 3. The "depth information of the human body" data is stored at the grid position where the central position of the human body in the image is identified.

[0035] Furthermore, the depth direction refers to the direction of front / back as seen from the camera. Alternatively, it can be the direction of the camera's optical axis. The order is assigned to the human bodies that appear in the image. For example, the person closest to the camera in the image is assigned the number 0, and the number increases by one as you move further away from the camera. To explain using Figure 3 as an example, Person 1 is located closest to the camera in real space, and the people are lined up in order from front to back as Person 1, Person 3, and Person 2. Therefore, the depth information for the human bodies is 0 (=i1) for Person 1, 1 (=i3) for Person 3, and 2 (=i2) for Person 2.

[0036] Furthermore, the order can be normalized to a value between 0 and 1 by dividing so that the furthest object in the image is number 1. In addition, although the order currently uses numbers that increase by 1 from the foreground, it is also possible to use numbers that reflect the distance between people in real space (the apparent distance is also acceptable). These orders allow for the generation of the correct answer by visual inspection from human body images, making it easy to collect training data.

[0037] The "relative position of keypoints" data shown in Figures 15 and 16 represents the relative position of each keypoint on the image relative to the human body, as shown in Figure 3, within the grid where the central position of the human body on the image is identified. Specifically, it represents the amount of movement in the x-direction and the y-direction from the center of the grid where the central position of the human body on the image is identified to the position of each keypoint on the image corresponding to that human body. The "relative position of keypoints" data is stored at the grid position where the central position of the human body on the image is identified. As shown in Figure 3, the central position of the human body on the image exists at a certain position within a single grid. By using the likelihood of the human body position and the relative position of each keypoint, the position of each keypoint on the image of the human body can be determined.

[0038] The "relative depth information of keypoints" data shown in Figures 15 and 16 represents the relative position in the depth direction relative to the center position of the human body in real space, for each keypoint corresponding to the human body in the grid where the center position of the human body in the image is identified, as shown in Figure 3. The "relative depth information of keypoints" data is stored at the grid position where the center position of the human body in the image is identified. The depth direction refers to the direction of front / back as seen from the camera. Alternatively, it may be the direction of the camera's optical axis. The relative position in the depth direction is, for example, a negative value if the keypoint is in front of the human body's center position in real space, a positive value if the keypoint is in the background, and 0 if the keypoint is at the center position of the human body in real space. To explain using Figure 4 as an example, keypoint a of person 1 is located in front, closer to the camera, relative to the human body's center position in real space. Keypoint b of person 1 is located far from the camera, further away, relative to the human body's center position in real space. Therefore, the relative depth information of the key point is that key point a of person 1 is -1 (=d a1 ), the key point b of person 1 is 1 (=d b1) This is the result. In Figure 4, the relative depth information of the keypoint is given in the form of three values: -1, 0, and 1. However, numerical values ​​that directly reflect the position from the center of the human body in real space (the apparent position may also be used) may be used. These relative positions in the depth direction can be visually generated from human body images, making it easy to collect training data.

[0039] Note that while Figure 3 shows the locations of two key points for each person, the number of key points can be three or more.

[0040] In the technology of this embodiment, after outputting multiple data points as described above from an input image, the parameters of the estimation model are calculated (learned) by minimizing the value of a predetermined loss function based on these multiple data points and a pre-given correct label.

[0041] Furthermore, during estimation, based on the "likelihood of human body position" data shown in Figure 2, the grid where the center position of each human body on the image is located is identified, and the correction amount corresponding to the identified grid position is obtained from the "correction amount of human body position" data shown in Figure 2. Based on the identified grid position (center position of the grid) and the obtained correction amount, the center position of each human body on the image is determined.

[0042] Next, depth information corresponding to the identified grid positions is obtained from the "Human Body Depth Information" data shown in Figure 2. The obtained depth information is used to determine the order in the depth direction in real space for each human body. Next, the relative positions of each keypoint corresponding to the identified grid positions are obtained from the "Relative Positions of Each Keypoint" data shown in Figure 2. Based on the identified grid positions (center positions of the grid) and the obtained relative positions of each keypoint, the image positions of each keypoint in each human body are determined.

[0043] Next, relative depth information for each keypoint corresponding to the identified grid position is obtained from the "relative depth information of each keypoint" data shown in Figures 15 and 16. Based on the obtained relative depth information of each keypoint, the center position of each keypoint in the human body in real space is used as a reference. and Identify the relative position in the depth direction.

[0044] As described above, during estimation, a grid is identified where the central position of each human body on the image is located. Based on the position of this identified grid, the spatial position information of the human body (the central position of the human body on the image for each human body and the order in the depth direction in real space for each human body) and the spatial position information of key points associated with the human body (the position of each key point on the image for each human body and the central position of each key point on the human body in real space are used as a reference). and This identifies the real-space position information of key points on the human body, indicated by the relative position in the depth direction.

[0045] Furthermore, the technology of this embodiment, by having the above features, allows for easy training and construction of an estimation model in a device that uses a pre-trained estimation model to estimate the real-space position information of key points of the human body from an image.

[0046] <Functional Configuration> Next, the functional configuration of the learning device 10 of this embodiment will be described. Figure 5 shows an example of a functional block diagram of the learning device 10. As shown in the figure, the learning device 10 has an acquisition unit 11, a learning unit 12, and a storage unit 13. However, as shown in the functional block diagram of Figure 6, the learning device 10 does not necessarily have a storage unit 13. In this case, an external device configured to communicate with the learning device 10 has a storage unit 13.

[0047] The acquisition unit 11 acquires training data that links training images with correct labels. The training images include people. The training images may include only one person or multiple people. The correct labels indicate at least the position of each key point on the human body in the image, the relative position in the depth direction of each key point on the human body with respect to the center position of the human body in real space, the center position of the human body in the image, and the order of the human body in the depth direction. The center position of the human body in the image may be calculated from the position of each key point on the human body in the image. For example, it may be the center of a rectangle that encompasses the positions of each key point on the human body, or the centroid using the positions of each key point on the human body. The correct labels may also be new correct labels obtained by processing the above correct labels. For example, the correct labels may be the multiple data shown in Figure 1, which are obtained by processing the above correct labels.

[0048] For example, an operator creating a ground truth label would only need to specify the positions within the image for the ground truth labels, namely "the image position of each key point on the human body" and "the image center position of the human body." Furthermore, the operator creating the ground truth label would only need to specify the relative positions and order of the ground truth labels, namely "the relative depth position of each key point on the human body relative to the real-space center position of the human body" and "the depth order of the human body," in accordance with how the person appears within the image.

[0049] Here, key points may be at least a part of a joint, a specific body part (eyes, nose, mouth, navel, etc.), or an extremity (head, hands, feet, etc.). Key points may also be other parts. The number and location of key points can be defined in various ways. each Yes, there are no particular restrictions.

[0050] For example, a large amount of training data is stored in the memory unit 13. The acquisition unit 11 can then acquire the training data from the memory unit 13.

[0051] The learning unit 12 learns an estimation model based on the training data. The memory unit 13 stores the estimation model. The estimation model consists of a neural network as explained using Figure 1. The estimation model outputs multiple data shown in Figure 1. The multiple data shown in Figure 1 represent the information necessary to determine the real-space position information of the human body and the real-space position information of key points associated with the human body, and show the likelihood of the human body position, the amount of correction for the human body position, the depth information of the human body, the relative position of each key point, and the relative depth information of each key point. Details regarding the multiple data are explained above.

[0052] Then, various estimation processes can be performed using the multiple data output by the estimation model. For example, the estimation device (for example, the estimation device 20 described in the following embodiment) obtains multiple data from the estimation model as described with reference to Figures 1 to 3, 15 and 16. The estimation device uses the multiple data obtained by the estimation model to obtain real-space position information of the human body (the central position of the human body on the image and the order in the depth direction in real space for each human body) and real-space position information of key points associated with the human body (the position of each key point on the image and the central position of each key point in real space for each human body). and The relative position in the depth direction can be estimated.

[0053] For example, the estimation device identifies the center position of each human body in the image based on the likelihood of the human body position and the correction amount of the human body position shown in Figure 2. The estimation device also identifies the order of the human body in the depth direction in real space based on the likelihood of the human body position and the depth information of the human body shown in Figure 2. Furthermore, the estimation device identifies the position of each keypoint in the image for each human body based on the likelihood of the human body position and the relative position of each keypoint shown in Figures 2, 15, and 16. Finally, the estimation device uses the center position of each keypoint in real space of the human body as a reference, based on the likelihood of the human body position and the relative depth information of each keypoint shown in Figures 2, 15, and 16. and Identify the relative position in the depth direction.

[0054] When training an estimation model, the learning unit 12 can learn (adjust) the parameters of the estimation model to minimize the error between the multiple data output from the estimation model during training and the multiple data in the training data (ground truth labels). In addition, the learning unit 12 can train on all grids for the "likelihood of human body position" data. Furthermore, for the "correction amount of human body position," "depth information of the human body," "relative position of each key point," and "relative depth information of each key point" data, the learning unit 12 can train only on the grids where the center position of the human body on the image is located in the training data.

[0055] Here, we will explain a specific example of the learning method used by the learning unit 12.

[0056] The learning unit 12 can learn (adjust) the parameters of the estimation model to minimize the error between the map showing the likelihood of human body position output from the estimation model under training and the map showing the likelihood of human body position of the training data (ground truth labels) for all grid positions.

[0057] Furthermore, with respect to the "human body position correction amount" data, the learning unit 12 can learn (adjust) the parameters of the estimation model to minimize the error between the human body position correction amount output from the estimation model during training and the human body position correction amount in the training data (ground truth labels), using only the grid position where the center position of the human body on the image is located in the training data.

[0058] Furthermore, with respect to the "depth information of the human body" data, the learning unit 12 can learn (adjust) the parameters of the estimation model to minimize the error between the depth information of the human body output from the estimation model during training and the depth information of the human body in the training data (ground truth labels), using only the grid position where the center position of the human body on the image is located in the training data.

[0059] Furthermore, with respect to the data of "relative position of each keypoint," the learning unit 12 can learn (adjust) the parameters of the estimation model to minimize the error between the relative position of each keypoint output from the estimation model during learning and the relative position of each keypoint in the learning data (ground truth labels), using only the grid position where the center position on the image of the human body is located in the learning data.

[0060] Furthermore, with respect to the "relative depth information of each keypoint" data, the learning unit 12 can learn (adjust) the parameters of the estimation model to minimize the error between the relative depth information of each keypoint output from the estimation model during learning and the relative depth information of each keypoint in the learning data (ground truth labels), using only the grid position where the central position on the image of the human body is located in the learning data.

[0061] An example of the processing flow of the learning device 10 will be explained using Figure 7.

[0062] In S10, the learning device 10 acquires training data that links training images with correct labels. This process is carried out by the acquisition unit 11. The details of the process performed by the acquisition unit 11 are as described above.

[0063] In S11, the learning device 10 trains an estimation model using the training data acquired in S10. This process is carried out by the learning unit 12. The details of the process executed by the learning unit 12 are as described above.

[0064] The learning device 10 repeats the loops S10 and S11 until the termination condition is met. The termination condition is defined, for example, using the value of the loss function.

[0065] <Hardware Configuration> Next, an example of the hardware configuration of the learning device 10 will be described. Each functional unit of the learning device 10 is realized by any combination of hardware and software, centered around a CPU (Central Processing Unit) of any computer, memory, programs loaded into memory, a storage unit such as a hard disk that stores those programs (which can store programs pre-installed at the time of shipment, as well as programs downloaded from recording media such as CDs (Compact Discs) or from servers on the Internet), and a network connection interface. It will be understood by those skilled in the art that there are various modifications to the implementation method and the device.

[0066] Figure 17 is a block diagram illustrating the hardware configuration of the learning device 10. As shown in Figure 17, the learning device 10 has a processor 1A, memory 2A, input / output interface 3A, peripheral circuitry 4A, and bus 5A. Peripheral circuitry 4A includes various modules. The learning device 10 does not necessarily have peripheral circuitry 4A. The learning device 10 may also be composed of multiple physically and / or logically separated devices. In this case, each of the multiple devices may have the above hardware configuration.

[0067] Bus 5A is a data transmission path for the processor 1A, memory 2A, peripheral circuits 4A, and input / output interface 3A to send and receive data to and from each other. Processor 1A is a processing unit such as a CPU or GPU (Graphics Processing Unit). Memory 2A is a memory such as RAM (Random Access Memory) or ROM (Read Only Memory). Input / output interface 3A includes interfaces for acquiring information from input devices, external devices, external servers, external sensors, cameras, etc., and interfaces for outputting information to output devices, external devices, external servers, etc. Input devices include, for example, keyboards, mice, microphones, physical buttons, touch panels, etc. Output devices include, for example, displays, speakers, printers, mailers, etc. Processor 1A can issue commands to each module and perform calculations based on the results of those calculations.

[0068] <Effects and Effects> The estimation model learned by the learning device 10 of this embodiment has the characteristic of outputting multiple data, including "likelihood of human body position," "amount of correction for human body position," "depth information of the human body," "relative position of each key point," and "relative depth information of each key point."

[0069] Furthermore, by using the multiple data output from this estimation model, it is possible to obtain information similar to the real-space position information of key points on the human body, namely, "real-space position information of the human body (the position coordinates on the image of the human body and the order of the human body in the depth direction in real space)" and "real-space position information of key points associated with the human body (the position coordinates on the image of the key points associated with the human body and the relative position in the depth direction of the key points associated with the human body with respect to the center of the human body)."

[0070] Furthermore, the multiple data points output from the estimation model can be easily collected / generated manually as long as human body images are available. Therefore, a sufficient amount and variety of training data can be easily collected / generated. With such a learning device 10, in a device that estimates the real-space position information of key points of the human body from images using a trained estimation model, an estimation model trained with a sufficient amount and variety of training data can be easily trained / constructed.

[0071] Furthermore, according to the learning device 10 of this embodiment, it is possible to estimate the real-world position information of the human body and the real-world position information of key points associated with the human body from a processed image using a trained estimation model. This allows for easy collection and generation of training data for the estimation model using only images, without the need for special equipment. In addition, even if multiple people are depicted in the image, it is possible to estimate each key point (each joint point), which is 3D skeletal information, while associating it with each person.

[0072] "Fourth Embodiment" The estimation device 20 of this embodiment uses the estimation model learned by the learning device 10 of the third embodiment to estimate the real-space position information of the human body (the position coordinates of the human body on the image and the order of the human body in the depth direction in real space) and the real-space position information of key points associated with the human body (the position coordinates of the key points associated with the human body on the image and the relative position in the depth direction of the key points associated with the human body with respect to the center of the human body). A detailed explanation follows below.

[0073] Figure 8 shows an example of a functional block diagram of the estimation device 20. As shown in the figure, the estimation device 20 has an estimation unit 21 and a storage unit 22. However, as shown in the functional block diagram of Figure 9, the estimation device 20 does not necessarily have a storage unit 22. In this case, an external device configured to communicate with the estimation device 20 provides the storage unit 22.

[0074] The estimation unit 21 acquires an arbitrary image as a processing image. For example, the estimation unit 21 may acquire an image captured by a camera or an image from a stored video as a processing image.

[0075] The estimation unit 21 then uses the estimation model learned by the learning device 10 to estimate and output the real-space position information of the human body (the position coordinates of the human body on the image and the order of the human body in the depth direction) and the real-space position information of key points associated with the human body (the position coordinates of the key points associated with the human body on the image and the relative position of the key points associated with the human body in the depth direction relative to the center of the human body).

[0076] As described in the third embodiment, when an image is input, the estimation model outputs the data described with reference to Figures 1 to 3, 15 and 16. The estimation unit 21 uses the data output by the estimation model to perform further estimation processing to estimate the real-space position information of the human body (the position coordinates of the human body on the image and the order of the human body in the depth direction in real space) and the real-space position information of key points associated with the human body (the position coordinates of the key points associated with the human body on the image and the relative position in the depth direction of the key points associated with the human body with respect to the center of the human body), and outputs the estimation results. The trained estimation model is stored in the storage unit 22. The output of the estimation results can be realized using any means such as a display, projection device, printer, or email. Alternatively, the estimation unit 21 may output the data output by the estimation model as the estimation result.

[0077] Below, an example of the processing performed by the estimation unit 21 will be explained using Figures 10 and 11.

[0078] (Step 1): Process the processed images with an estimation model to obtain multiple data points as shown in Figures 1 to 3, 15 and 16.

[0079] (Step 2): Based on the "likelihood of human body position" data, identify the grid (P2 in Figure 10) in which the central position (P1 in Figure 10) of each person (each body) on the image is located (contained). Specifically, identify the grids in which the likelihood is above a threshold. Furthermore, identify the central position of the grid (P3 in Figure 10) from the identified grids.

[0080] (Step 3): From the "human body position correction amount" data, obtain the correction amount (P4 in Figure 10) corresponding to the grid position identified in (Step 2).

[0081] (Step 4): Based on the grid center position identified in (Step 2) and the correction amount obtained in (Step 3), the coordinates of the center position of each person in the processed image (P1 in Figure 10) are identified. This identifies the position coordinates of each person's body in the image.

[0082] (Step 5): From the "human body depth information" data, obtain the depth information corresponding to the grid positions identified in (Step 2), that is, the order in the depth direction in real space. This identifies the order of each human body in the depth direction in real space.

[0083] (Step 6): From the data of "relative position of each keypoint," obtain the relative position (P6 in Figure 11) corresponding to the grid position identified in (Step 2).

[0084] (Step 7): Based on the grid center position identified in (Step 2) and the relative position obtained in (Step 6), the position coordinates of each key point on the image (P7 in Figure 11) are identified for each person included in the processed image. This identifies the position coordinates on the image of each key point associated with each person.

[0085] (Step 8): From the "relative depth information of each keypoint" data, use the relative depth information corresponding to the grid position identified in (Step 2), that is, the center position of the human body in real space as the reference. and The relative position in the depth direction (P8 in Figure 11) is obtained. This identifies the relative position in the depth direction of each key point associated with each human body, relative to the center of the human body.

[0086] (Step 9): Outputs the image position coordinates of each human body identified in (Step 4), the depth order of each human body in real space identified in (Step 5), the image position coordinates of each key point associated with each human body identified in (Step 7), and the relative depth position of each key point associated with each human body identified in (Step 8) with respect to the human body center.

[0087] As a result, the estimation unit 21 can use the estimation model learned by the learning device 10 to estimate and output the real-space position information of the human body (the position coordinates of the human body on the image and the order of the human body in the depth direction) and the real-space position information of key points associated with the human body (the position coordinates of the key points associated with the human body on the image and the relative position in the depth direction of the key points associated with the human body with respect to the center of the human body).

[0088] Furthermore, as shown in Figure 12, the estimation unit 21 can superimpose and display the estimated information onto the image used for estimation. The estimation unit 21 can superimpose and display the order of the estimated human bodies in the depth direction in real space (P13 in Figure 12) corresponding to each human body, based on the position coordinates on the image of each estimated human body (P11 in Figure 12), or the position coordinates on the image of each key point associated with each estimated human body (P12 in Figure 12).

[0089] Furthermore, the estimation unit 21 can overlay an object representing a key point onto the image position coordinates (P12 in Figure 12) of each key point associated with each estimated human body. The estimation unit 21 can then make the color (or shape or size) of that object correspond to the estimated relative position value in the depth direction that corresponds to that key point. As described above, by overlaying depth-related information, depth-related information that is difficult to represent on an image becomes visually and intuitively easier to understand.

[0090] In addition, the above display method can also be used as support when manually generating the correct labels (training data) necessary for training an estimation model. When manually inputting correct labels on an image, the correct labels in the depth direction are difficult to represent on the image, leading to errors in correct labeling. Therefore, by using the above display method, errors in correct labeling can be reduced. When manually inputting the correct level in the depth direction, displaying the input status sequentially on the image using the above display method makes the state of the correct labels in the depth direction visually clear on the image, reducing errors in label input.

[0091] Next, an example of the processing flow of the estimation device 20 will be explained using the flowchart in Figure 13.

[0092] In S20, the estimation device 20 acquires the processed image. For example, an operator inputs a processed image to the estimation device 20. The estimation device 20 then acquires the input processed image.

[0093] In S21, the estimation device 20 uses the estimation model learned by the learning device 10 to estimate the real-space position information of the human body (the position coordinates of the human body on the image and the order of the human body in the depth direction in real space) and the real-space position information of key points associated with the human body (the position coordinates of the key points associated with the human body on the image and the relative position in the depth direction of the key points associated with the human body with respect to the center of the human body) from the processed image. This process is carried out by the estimation unit 21. Details of the process executed by the estimation unit 21 are as described above.

[0094] In S22, the estimation device 20 outputs the estimation result from S21. The estimation device 20 can utilize any means such as a display, projection device, printer, or email.

[0095] Next, an example of the hardware configuration of the estimation device 20 will be described. Each functional unit of the estimation device 20 is realized by any combination of hardware and software, centered around the CPU of any computer, memory, a program loaded into memory, a storage unit such as a hard disk that stores that program (which can store programs that are pre-installed at the time of shipment, as well as programs downloaded from recording media such as CDs or from servers on the Internet), and a network connection interface. It will be understood by those skilled in the art that there are various modifications to the implementation method and the device. Figure 17 is a block diagram illustrating the hardware configuration of the estimation device 20.

[0096] According to the estimation device 20 of this embodiment described above, it is possible to estimate the real-space position information of the human body and the real-space position information of key points associated with the human body from a processed image using the estimation model learned by the learning device 10 of the third embodiment. With such an estimation device 20, it is possible to easily collect and generate training data for the estimation model using only an image, without the need for any special equipment, and to easily learn and build the estimation model. Furthermore, with such an estimation device 20, even if multiple people are shown in the image, it is possible to estimate each key point (each joint point), which is 3D skeletal information, while associating it with each person.

[0097] "Fifth Embodiment" Next, a fifth embodiment will be described in detail with reference to the drawings.

[0098] Referring to Figure 14, in the fifth embodiment, a computer-readable storage medium 102 that stores the 3D skeleton estimation program 101 is connected to the computer 100.

[0099] The computer-readable storage medium 102 is composed of a magnetic disk, semiconductor memory, etc., and the 3D skeleton estimation program 101 stored therein is read by the computer 100 when the computer 100 is started up, and by controlling the operation of the computer 100, the computer 100 is made to function as the respective functional units 11, 12, and 13 in the learning device 10 of the first and third embodiments described above, and to perform the processing shown in Figure 7.

[0100] In this embodiment, the learning device 10 according to the first and third embodiments is implemented using a computer and a program, but the estimation device 20 according to the second and fourth embodiments can also be implemented using a computer and a program in a similar manner.

[0101] "Industrial applicability" The first to fifth embodiments described above are A 3D skeletal estimation device that can estimate the real-world positional information of the human body and the real-world positional information of key points associated with the human body from an image. • A 3D skeleton estimation device that can estimate the 3D skeletal information of each key point (each joint point) by linking it to each person, even if multiple people are shown in the image. A device that uses a pre-trained estimation model to estimate the real-space position information of key points in the human body from an image, and a 3D skeleton estimation device that can easily train / construct the estimation model. • Programs to implement those 3D skeleton estimation devices on a computer. It can be used for the following purposes.

[0102] Furthermore, the first to fifth embodiments described above can be adapted to applications such as devices and functions that perform image recognition requiring the estimation of the real-world positional information of a human body and the real-world positional information of key points associated with the human body, based on camera and stored images.

[0103] Furthermore, the first to fifth embodiments described above can be adapted for applications such as devices and functions for behavioral analysis in the marketing and surveillance fields.

[0104] Furthermore, the first to fifth embodiments described above can be applied to applications such as input interfaces that take as input information the real-world position of a human body estimated from camera and stored images, and the real-world position information of key points associated with the human body.

[0105] In addition, the first to fifth embodiments described above can be applied to applications such as video / image retrieval devices and functions that use estimated real-space position information of a human body and real-space position information of key points associated with the human body as trigger keys.

[0106] The embodiments of the present invention have been described above with reference to the drawings, but these are illustrative examples of the present invention, and various other configurations can be adopted. The configurations of the embodiments described above may be combined with each other, or some configurations may be replaced with other configurations. Furthermore, the configurations of the embodiments described above may be modified in various ways without departing from the spirit of the invention. In addition, the configurations and processes disclosed in each of the embodiments and modifications described above may be combined with each other.

[0107] Furthermore, while the flowcharts used in the above description show multiple steps (processes) in sequence, the execution order of the steps performed in each embodiment is not limited to the order in which they are described. In each embodiment, the order of the illustrated steps can be changed to the extent that it does not impede the content. Also, the above embodiments can be combined to the extent that their contents do not conflict.

[0108] In this specification, "acquisition" includes at least one of the following: "active acquisition" based on user input or program instructions, such as "the device retrieving data stored in another device or storage medium," for example, receiving data by requesting or inquiring with another device, or accessing and reading data from another device or storage medium; "passive acquisition" based on user input or program instructions, such as "inputting data output from another device into the device," for example, receiving data that is distributed (or transmitted, push notification, etc.), or selecting and acquiring data from the received data or information; and "generating new data by editing data (converting to text, rearranging data, extracting some data, changing file format, etc.) and acquiring said new data."

[0109] Some or all of the above embodiments may also be described as follows, but are not limited to the following. 1. An acquisition means for acquiring training data that links a training image containing a human body with a ground truth label indicating the central position of each human body in the image, a ground truth label indicating the position of each key point in each human body in the image, and a ground truth label indicating the relative position in the depth direction of each key point in each human body. A learning means for learning an estimation model that, based on the aforementioned training data, estimates information indicating the likelihood of the central position on the image for each human body, the relative position on the image of each key point associated with each human body, and the relative position in the depth direction of each key point associated with each human body. A learning device having 2. In the aforementioned training data, correct labels indicating the depth-direction position information of each human body are added and linked. The estimation model adds the relative position on the image, which indicates the central position of each human body on the image, and the depth position information for each human body to the estimation. The learning device according to claim 1, characterized in that it is a learning device. 3. The positional information in the depth direction refers to the order of the people in the image in the depth direction in real space. The learning device according to claim 2, characterized in that it is a learning device. 4. The relative position on the image is the relative position on the image with respect to the center of the grid identified in the likelihood information as the location of the central position of the human body on the image. The relative positions of each of the aforementioned keypoints in the depth direction are relative positions in the depth direction in real space, with reference to the center position of the human body in real space. A learning device according to any one of 1 to 3, characterized by the features described herein. 5. The relative position of each keypoint in the depth direction shall be determined based on the center position of the human body in real space. A negative value shall be used if the keypoint is in front of the image, a positive value if the keypoint is behind the image, and 0 if the keypoint is at the center position of the human body in real space. The learning device according to claim 4, characterized in that it is a learning device. 6. The learning means is, Based on the aforementioned estimation model under training, the following information is estimated: the likelihood of the central position on the image for each human body, the relative position on the image for each key point associated with each human body, the relative position in the depth direction of each key point associated with each human body relative to the central position of the human body in real space, the relative position on the image indicating the central position of the human body for each human body, and the order in the depth direction for each human body. The parameters of the estimation model are adjusted to minimize the error between the estimated information indicating the likelihood of the central position on the image for each human body and the information indicating the likelihood of the central position on the image for each human body obtained from the ground truth labels, for all grid positions. The parameters of the estimation model are adjusted so that the error between the estimated relative position on the image of each key point associated with each human body and the relative position on the image of each key point associated with each human body obtained from the ground truth labels is minimized only with respect to the grid position where the center position on the image of the human body is located in the training data. The parameters of the estimation model are adjusted to minimize the error between the estimated relative position in the depth direction of each key point associated with each human body, relative to the center position of the human body in real space, obtained from the ground truth labels, and the grid position where the center position of the human body on the image is located in the training data. The parameters of the estimation model are adjusted so that the error between the estimated relative position on the image, which indicates the central position of each human body on the image, and the relative position on the image, which indicates the central position of each human body obtained from the ground truth label, is minimized only for the grid position where the central position of the human body on the image is located in the training data. The parameters of the estimation model are adjusted so that the error between the estimated depth order of each human body and the depth order of each human body obtained from the ground truth labels is minimized only for the grid positions where the central position of the human body on the image is located in the training data. A learning device according to claim 4 or 5, characterized in that it is a learning device. 7. The depth direction is the direction of the camera's optical axis. A learning device according to any one of 1 to 6, characterized by the features described herein. 8. Estimation means for estimating the image position coordinates of each key point associated with each human body, the relative position in the depth direction of each key point associated with each human body with respect to the center position of the human body in real space, the position coordinates indicating the center position of the human body in the image for each human body, and the order in the depth direction for each human body, using an estimation model learned by any of the learning devices described in 1 to 7. An estimation device having the following features. 9. The estimation means is, Based on the information indicating the likelihood of the central position on the image for each human body obtained from the estimation model, the grid in which the central position on the image of each human body is located is identified. From the relative positions on the image of each key point associated with each human body obtained from the estimation model, the relative positions on the image corresponding to the positions of the identified grid are obtained, and based on the center position of the identified grid and the obtained relative positions on the image, the position coordinates on the image of each key point associated with each human body are estimated. From the relative depth position of each key point associated with each human body obtained from the estimation model, relative depth position corresponding to the identified grid position is obtained, and the relative depth position of each key point associated with each human body, relative depth position of the human body, relative depth position is estimated. From the relative position on the image indicating the center position of each human body obtained from the estimation model, the relative position on the image corresponding to the identified grid position is obtained, and based on the identified grid center position and the obtained relative position on the image, the position coordinates indicating the center position of each human body on the image are estimated. From the depth-direction order of each human body obtained from the estimation model, the depth-direction order corresponding to the identified grid positions is obtained, and the depth-direction order of each human body is estimated. The estimation device according to claim 8, characterized in that it is a spectroscopy device. 10. The estimation means superimposes the depth order of each estimated human body onto the image used for estimation, based on the position coordinates indicating the central position of each estimated human body on the image, or the position coordinates on the image of each key point associated with each estimated human body. The estimation device according to 8 or 9, characterized in that it is a device that provides an estimate. 11. The estimation means displays an object representing a key point superimposed on the image coordinates of each key point associated with each estimated human body in the image used for estimation, and the color, shape, or size of the object corresponds to the value of the relative position in the depth direction of the key point, which corresponds to the center position of each key point associated with each estimated human body in real space. An estimation device according to any one of 8 to 10, characterized in that it is a device. 12. The aforementioned depth direction is the direction of the camera's optical axis. An estimation device according to any one of 8 to 11, characterized in that it is a device. 13. One or more computers, We obtain training data that links training images containing human bodies with ground truth labels indicating the central position of each human body in the image, ground truth labels indicating the position of each key point in each human body in the image, and ground truth labels indicating the relative position of each key point in the depth direction in each human body. Based on the aforementioned training data, an estimation model is trained to estimate information indicating the likelihood of the central position on the image for each human body, the relative position on the image of each key point associated with each human body, and the relative position in the depth direction of each key point associated with each human body. Learning methods. 14. Computers, Acquisition means for acquiring training data that links a training image containing a human body with a ground truth label indicating the central position of each human body in the image, a ground truth label indicating the position of each key point in each human body in the image, and a ground truth label indicating the relative position in the depth direction of each key point in each human body. A learning means for learning an estimation model that estimates information indicating the likelihood of the central position on an image for each human body, the relative position on the image of each key point associated with each human body, and the relative position in the depth direction of each key point associated with each human body, based on the aforementioned training data. A program that makes it function as such. 15. One or more computers, Using an estimation model trained by one of the learning devices described in 1 to 7, the image position coordinates of each key point associated with each human body, the relative position in the depth direction of each key point associated with each human body relative to the center position of the human body in real space, the position coordinates indicating the center position of each human body in the image, and the order in the depth direction for each human body are estimated. Estimation method. 16. Computers, Estimation means for estimating the image position coordinates of each key point associated with each human body, the relative position in the depth direction of each key point associated with each human body with respect to the center position of the human body in real space, the position coordinates indicating the center position of the human body in the image for each human body, and the order in the depth direction for each human body, using an estimation model learned by any of the learning devices described in 1 to 7. A program that makes it function as such.

[0110] This application claims priority based on Japanese Patent Application No. 2022-200103, filed on 15 December 2022, and incorporates all of its disclosures herein. [Explanation of symbols]

[0111] 10 Learning device 11 Acquisition Department 12. Learning Department 13 Storage section 20 Estimation device 21 Estimation part 22 Memory section 100 Computers 101 Program for 3D Skeleton Estimation 102 Computer-readable storage media 1A Processor 2A Memory 3A input / output I / F 4A Peripheral Circuits 5A Bus

Claims

1. An acquisition means for acquiring training data that links a training image containing a human body with a ground truth label indicating the central position of each human body in the image, a ground truth label indicating the position of each key point in each human body in the image, a ground truth label indicating the relative position in the depth direction of each key point in each human body, and a ground truth label indicating the order in the depth direction of each human body in real space. A learning means for learning an estimation model that, based on the aforementioned training data, estimates information indicating the likelihood of the central position on the image for each human body, the relative position on the image of each key point associated with each human body, the relative position in the depth direction of each key point associated with each human body, and the order in the depth direction in real space for each human body. A learning device having the following features.

2. The estimation model estimates by adding the relative position on the image that indicates the central position of each human body on the image, The learning device according to feature 1.

3. The relative position on the image is the relative position on the image with respect to the center of the grid, which is identified in the likelihood information as the location of the central position of the human body on the image. The relative positions of each of the aforementioned keypoints in the depth direction are relative positions in the depth direction in real space, with reference to the center position of the human body in real space. The learning device according to feature 1 or 2.

4. An estimation means for acquiring training data that links a training image containing a human body with a ground truth label indicating the central position of each human body in the image, a ground truth label indicating the position of each key point in each human body in the image, and a ground truth label indicating the relative position in the depth direction of each key point in each human body; and a learning means for learning an estimation model that estimates information indicating the likelihood of the central position of each human body in the image, the relative position of each key point associated with each human body in the image, and the relative position in the depth direction of each key point associated with each human body, using the estimation model learned by the learning device, estimates the position coordinates of each key point associated with each human body in the image, the relative position in the depth direction of each key point associated with each human body with respect to the central position of the human body in real space, the position coordinates indicating the central position of each human body in the image, and the order in the depth direction of each human body. An estimation device having the following features.

5. The estimation means superimposes and displays the depth order of each estimated human body on the image used for estimation, based on the position coordinates indicating the center position of each estimated human body on the image, or the position coordinates on the image of each key point associated with each estimated human body. The estimation device according to feature 4.

6. The estimation means displays an object representing a key point superimposed on the image coordinates of each key point associated with each estimated human body, and the color, shape, or size of the object corresponds to the value of the relative position in the depth direction of the key point, which is the center position of each key point associated with each estimated human body in real space. The estimation device according to claim 4 or 5.

7. One or more computers, We obtain training data that links training images containing human bodies with ground truth labels indicating the central position of each human body in the image, ground truth labels indicating the position of each key point in each human body in the image, ground truth labels indicating the relative position of each key point in the depth direction in each human body, and ground truth labels indicating the order in the depth direction in real space for each human body. Based on the aforementioned training data, the system learns an estimation model that estimates information indicating the likelihood of the central position on the image for each human body, the relative position on the image of each key point associated with each human body, the relative position in the depth direction of each key point associated with each human body, and the order in the depth direction in real space for each human body. Learning methods.

8. Computers, Acquisition means for acquiring training data that links a training image containing a human body with a ground truth label indicating the central position of each human body in the image, a ground truth label indicating the position of each key point in each human body in the image, a ground truth label indicating the relative position in the depth direction of each key point in each human body, and a ground truth label indicating the order in the depth direction of each human body in real space. A learning means for learning an estimation model that learns information indicating the likelihood of the central position on an image for each human body, the relative position on the image of each key point associated with each human body, the relative position in the depth direction of each key point associated with each human body, and the order in the depth direction in real space for each human body, based on the aforementioned training data. A program that makes it function as such.

9. One or more computers, A learning device comprising: acquisition means for acquiring training data that links a training image containing a human body with a ground truth label indicating the central position of each human body in the image, a ground truth label indicating the position of each key point in each human body in the image, and a ground truth label indicating the relative position in the depth direction of each key point in each human body; and learning means for learning an estimation model that estimates information indicating the likelihood of the central position of each human body in the image, the relative position of each key point associated with each human body in the image, and the relative position in the depth direction of each key point associated with each human body, is used to estimate the position coordinates of each key point associated with each human body in the image, the relative position in the depth direction of each key point associated with each human body with respect to the central position of the human body in real space, the position coordinates indicating the central position of each human body in the image, and the order in the depth direction of each human body. Estimation method.

10. Computers, An estimation means for acquiring training data that links a training image containing a human body with a ground truth label indicating the central position of each human body in the image, a ground truth label indicating the position of each key point in each human body in the image, and a ground truth label indicating the relative position in the depth direction of each key point in each human body; and a learning means for learning an estimation model that estimates information indicating the likelihood of the central position of each human body in the image, the relative position of each key point associated with each human body in the image, and the relative position in the depth direction of each key point associated with each human body, using the estimation model learned by the learning device, estimates the position coordinates of each key point associated with each human body in the image, the relative position in the depth direction of each key point associated with each human body with respect to the central position of the human body in real space, the position coordinates indicating the central position of each human body in the image, and the order in the depth direction of each human body. A program that makes it function as such.

Citation Information

Patent Citations

  • Multi-person three-dimensional attitude estimation method and device and electronic equipment

    CN114550282A

  • Method and system for monocular depth estimation of a person

    JP2022536790A