Information Processing Apparatus, Information Processing Method, and Program
By constructing a training model that can learn and estimate joint coordinates outside the perspective, the problem of posture estimation in the case where traditional techniques are difficult to deal with the object part is outside the perspective, and accurate posture estimation in this case is achieved.
Patent Information
- Application Number
- CN201980100231.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-09-20
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2039-09-20
AI Technical Summary
Traditional techniques are difficult to deal with pose estimation when part of an object is located outside the perspective.
By learning the relationship between a first image of an object having multiple joints and coordinate information indicating the location of the multiple joints, a training model is constructed to estimate coordinate information of at least one joint located outside the perspective of the newly acquired second image of the object.
Even if part of the object is outside the perspective, it can accurately estimate the object's posture, solving the problem that traditional techniques are difficult to deal with this situation.
Smart Images

Figure CN114391156B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing device, an information processing method, and a program. Background Art
[0002] Various techniques for estimating an object's pose have been proposed, such as past estimation methods based on an object image, estimation methods based on the output of sensors attached to the object, and estimation methods based on a knowledge model. PTL 1 describes a motion model learning device using a total pose matrix and a partial pose matrix.
[0003] [Citation List]
[0004] [Patent Document]
[0005] [PTL 1] JP 2012-83955A Summary of the Invention
[0006] [Technical Problem]
[0007] However, the conventional techniques are premised on the whole object being within the view of the image pickup device. Therefore, for example, it is difficult to handle a case where a component of the object has a part outside the view in the generated image.
[0008] An object of the present invention is to provide an information processing device, an information processing method, and a program that can perform pose estimation based on an image even when a part of the object is outside the view.
[0009] [Solution to Problem]
[0010] According to one aspect of the present invention, there is provided an information processing device including: a relationship learning part that learns a relationship between a first image of an object having a plurality of joints and coordinate information indicating positions of the plurality of joints to construct a training model, the coordinate information indicating positions of the plurality of joints being defined in a range extended to be larger than a view of the first image, the training model estimating coordinate information of at least one joint located outside a view of a newly acquired second image of the object.
[0011] According to another aspect of the present invention, there is provided an information processing device including: a coordinate estimation part that estimates coordinate information of at least one joint located outside a view of a newly acquired second image of the object based on a training model constructed by learning a relationship between a first image of an object having a plurality of joints and coordinate information indicating positions of the plurality of joints, the coordinate information indicating positions of the plurality of joints being defined in a range extended to be larger than a view of the first image.
[0012] According to another aspect of the present invention, there is provided an information processing method, including: a step of constructing a training model by learning the relationship between a first image of an object having a plurality of joints and coordinate information indicating the positions of the plurality of joints, the coordinate information indicating the positions of the plurality of joints being defined within a range extended to be larger than the viewing angle of the first image, the training model being used to estimate the coordinate information of at least one joint located outside the viewing angle of a newly acquired second image of the object; and a step of estimating the coordinate information of at least one joint located outside the viewing angle of the second image based on the training model.
[0013] According to another aspect of the present invention, there is provided a program that causes a computer to implement: a function of constructing a training model by learning the relationship between a first image of an object having a plurality of joints and coordinate information indicating the positions of the plurality of joints, the coordinate information indicating the positions of the plurality of joints being defined within a range extended to be larger than the viewing angle of the first image, the training model being used to estimate the coordinate information of at least one joint located outside the viewing angle of a newly acquired second image of the object.
[0014] According to another aspect of the present invention, there is provided a program that causes a computer to implement: a step of estimating the coordinate information of at least one joint located outside the viewing angle of a newly acquired second image of the object based on a training model constructed by learning the relationship between a first image of an object having a plurality of joints and coordinate information indicating the positions of the plurality of joints, the coordinate information indicating the positions of the plurality of joints being defined within a range extended to be larger than the viewing angle of the first image. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a block diagram showing a schematic configuration of a system of an information processing apparatus including an embodiment of the present invention.
[0016] Figure 2 is a diagram showing Figure 1 an example of input data in an example.
[0017] Figure 3 depicts an example for explaining Figure 1 an example of estimating joint coordinates in an example.
[0018] Figure 4 depicts an example for explaining Figure 3 another example of estimating joint coordinates in an example.
[0019] Figure 5 depicts an example for explaining Figure 4 yet another example of estimating joint coordinates in an example.
[0020] Figure 6It is a flowchart showing an example of a process according to an embodiment of the present invention.
[0021] Figure 7 It is another flowchart showing an example of a process according to an embodiment of the present invention. Detailed Description of the Invention
[0022] Hereinafter, some embodiments of the present invention will be described in detail with reference to the accompanying drawings. It should be noted that in this specification and the drawings, components having substantially the same functional configuration are denoted by the same reference numerals, and thus repeated descriptions will be omitted.
[0023] Figure 1 It is a block diagram showing a schematic configuration of a system of an information processing apparatus including an embodiment of the present invention. In the illustrated example, the system 10 includes information processing apparatuses 100 and 200. Both the information processing apparatuses 100 and 200 are connected to a wired or wireless network, and a training model 300 constructed by the information processing apparatus 100 and stored in a memory on the network is read by the information processing apparatus 200 via the network, for example.
[0024] For example, the information processing apparatuses 100 and 200 are implemented by computers having a communication interface, a processor, and a memory. In the information processing apparatuses 100 and 200, the functions of the respective parts described below are implemented in software by the processor operating according to a program stored in the memory or received via the communication interface.
[0025] The information processing apparatus 100 includes an input section 110, a relationship learning section 120, and an output section 130. By using the training model 300 constructed by the information processing apparatus 100, the information processing apparatus 200, which will be described later, performs estimation processing based on an image of an object, so that the coordinates of the joints of the object can be estimated in a region including a part outside the perspective of the image pickup device.
[0026] The input section 110 receives the input of input data 111 used for learning performed by the relationship learning section 120. In the present embodiment, the input data 111 includes an image of an object having a plurality of joints and joint coordinate information of the object in the image.
[0027] Figure 2 It is a diagram showing Figure 1 an example of the input data 111. In the present embodiment, the input data 111 includes an image generated to have a composition in which a part of the object obj is outside the perspective, as shown in Figure 2 images A1 to A4. In images A1 to A4, the part of the contour line of the object obj included in the perspective is represented by a solid line, and the part located outside the perspective is represented by a dotted line. In addition, the solid lines representing the pose of the object obj represent the respective joints of the object obj and their mutual relationships.
[0028] In addition, the input data 111 may include an image generated to have a composition in which the entire object is within the viewing angle. The image may be a two-dimensional image generated by an RGB (Red-Green-Blue) sensor or the like, or may be a three-dimensional image generated by an RGB-D sensor or the like.
[0029] Furthermore, the input data 111 includes coordinate information indicating the positions of a plurality of joints of the object, as Figure 2 shown by the joint coordinate information Bc in. In the present embodiment, since the joint coordinate information Bc is defined within a range extended to be larger than the viewing angle of the image, the input data 111 includes images A1 to A4, and the joint coordinate information Bc indicating the positions of all the joints of the object obj in the images A1 to A4. Each of the images A1 to A4 is generated to have a composition in which a part of the object obj is outside the viewing angle.
[0030] For example, in the case of image A1, the joints J1 of the two wrists are outside the viewing angle of the image, but the input data 111 includes image A1 and the joint coordinate information Bc, and the joint coordinate information Bc includes the coordinate information of each joint located within the viewing angle of image A1 and the double-wrist joints J1.
[0031] Here, in the present embodiment, as Figure 2 shown, the joint coordinate information Bc is input independently of the images A1 to A4 based on three-dimensional pose data. Such three-dimensional pose data is obtained, for example, from images captured by a plurality of cameras different from the camera that captures the images A1 to A4, from the actions of an IMU (Inertial Measurement Unit) sensor attached to the object obj, etc. Incidentally, the detailed description of the acquisition of such three-dimensional pose data will be omitted because various known techniques can be used.
[0032] Referring again to Figure 1 , the relationship learning section 120 of the information processing device 100 learns the relationship between the image and the joint coordinate information input via the input section 110 to construct the training model 300. In the present embodiment, the relationship learning section 120 performs supervised learning to construct the training model 300 by using the image and the joint coordinate information input via the input section 110 as input data and using the three-dimensional pose data as the correct answer data. Incidentally, as for the specific method of machine learning, since various known techniques can be used, the detailed description thereof will be omitted. The relationship learning section 120 outputs the parameters of the constructed training model 300 via the output section 130.
[0033] The information processing device 200 includes an input section 210, a coordinate estimation section 220, a three-dimensional pose estimation section 230, and an output section 240. The information processing device 200 performs an estimation process based on an image of an object by using the training model 300 constructed by the information processing device 100, so as to estimate the coordinates of joints in a region of a part of the object outside the viewing angle of the image pickup device.
[0034] The input section 210 receives an input of an input image 211 for the estimation by the coordinate estimation section 220. The input image 211 is, for example, an image newly acquired by the image pickup device 212. The input image 211 is an image of an object obj having a plurality of joints as described above with reference to Figure 2 Note that the object in the image of the input data 111 and the object in the input image 211 have the same joint structure, but they do not necessarily have to be the same object. Specifically, for example, when the object in the image of the input data 111 is a person, the object in the input image 211 is also a person, but the object does not have to be the same person.
[0035] In addition, the input image 211 is not limited to an image acquired by the image pickup device 212. For example, an image stored in a storage device connected to the information processing device 200 by wire or wirelessly can be input as the input image 211 via the input section 210. In addition, an image acquired from a network can be input as the input image 211 via the input section 210. In addition, the input image 211 can be a still image or a moving image.
[0036] The coordinate estimation section 220 estimates the coordinates of a plurality of joints of the object from the input image 211 input via the input section 210 based on the training model 300. As described above, since the training model 300 is constructed based on the coordinate information of joints defined in a range larger than the viewing angle of the image, even in a region outside the viewing angle of the input image 211, it is possible to infer the positions of the respective joints and the link structure between these joints. As a result, the coordinate estimation section 220 can estimate that "there is no joint within the viewing angle of the image input to the input section 210, but there is a joint at the coordinates (X, Y, Z) outside the viewing angle." In addition, the coordinate estimation section 220 can also estimate the positional relationship of the plurality of joints based on the estimated coordinates of the plurality of joints.
[0037] Figure 3 Describes for illustration in Figure 1A diagram of an example for estimating joint coordinates in an example. In this embodiment, the training model 300 includes a first training model M1 that estimates the coordinates of joints located within the viewing angle from an image, and a second training model M2 that estimates the coordinates of at least one joint located outside the viewing angle from information about the joint coordinates within the viewing angle. The coordinate estimation section 220 performs a two-step estimation process using the first training model M1 and the second training model M2.
[0038] Here, in Figure 3 the example shown in (a) of , the input image 211 includes image A5, where the joints J2 of both ankles are located outside the viewing angle. Figure 3 The first training model M1 shown in (b) of is a training model based on CNN (Convolutional Neural Network). The coordinate estimation section 220 uses the first training model M1 to estimate the coordinates of the joints within the viewing angle of image A5, that is, the coordinates of the joints excluding the joints J2 of both ankles. As a result, intermediate data DT1 for identifying the coordinates of the joints within the viewing angle of the image can be obtained.
[0039] In addition, as shown in (c) of Figure 3 , the coordinate estimation section 220 uses the intermediate data DT1 to perform an estimation process using the second training model M2. Figure 3 The second training model M2 shown in (d) of is a training model based on RNN (Recurrent Neural Network), and can estimate the joint coordinates that are outside the viewing angle and not included in the intermediate data DT1, that is, the coordinates of the joints J2 of both ankles in this example.
[0040] In addition, in the illustrated example, the intermediate data DT1 is data representing the coordinates of the joints as two-dimensional coordinates, but the second training model M2 can estimate the joint coordinates as three-dimensional coordinates by inputting the time-series intermediate data DT1, as shown in Figure 3 in (e) of . As shown in (e) of Figure 3 , the final data DT2 obtained by estimation using the second training model M2 includes the estimation results of the coordinates of all the joints of the object, including the joints J2 of both ankles that are outside the viewing angle of image A5 shown in Figure 3 (a) of .
[0041] Referring again to Figure 1, the three-dimensional pose estimation section 230 of the information processing device 200 estimates the full-body pose of the object based on the joint coordinates estimated by the coordinate estimation section 220. The three-dimensional pose estimation section 230 outputs, via the output section 240, data 241 representing the estimated full-body pose of the object. For example, the data 241 representing the full-body pose of the object can be displayed on the display as an extended image of the input image 211. For example, the data 241 representing the full-body pose of the object can be output as the movement in the image of the user avatar imitating the pose of the object, or the image of the character in the game, the moving image, etc. Alternatively, the data 241 representing the full-body pose of the object can be output as the movement of the robot imitating the pose of the object, or instead of the output of the display.
[0042] According to the configuration of the present embodiment as described above, based on the training model 300 constructed by learning the relationship between the image of the object with multiple joints and the coordinate information of the joints defined within a range that extends beyond the viewing angle of the image, the coordinates of multiple joints including at least one joint located outside the viewing angle of the image in the input section 210 are estimated. Therefore, even when a part of the object is located outside the viewing angle of the image, the pose of the object can be estimated based on the image.
[0043] Figure 4 Depicts a diagram for explaining another example of estimating joint coordinates in the example of Figure 1 . In the example of Figure 4 , the training model 300 includes a first training model M1 similar to the example of Figure 3 , and a set of training models (third training model M3, fourth training model M4, and fifth training model M5) constructed for each joint.
[0044] In the examples shown in (a) and (b) of Figure 4 , similar to (a) and (b) of Figure 3 , the coordinate estimation section 220 uses the first training model M1 to estimate the coordinates of the joints located within the viewing angle of the image A5, that is, the joints other than the joints J2 of the two ankles. As a result, intermediate data DT1 for identifying the coordinates of the joints located within the viewing angle of the image can be obtained.
[0045] Next, as shown in (c) of Figure 4 , the coordinate estimation section 220 uses the intermediate data DT1 to perform the estimation process using the third training model M3 to the fifth training model M5.
[0046] Figure 4 The third training model M3 to the fifth training model M5 shown in (d) of Figure 3 are training models based on RNN (recurrent neural network) and similar to the second training model M2 shown in Figure 4The third training model M3 to the fifth training model M5 shown in (d) are training models constructed in a limited manner for estimating the coordinates of a single joint (or a group of joints). For example, the third training model M3 is a training model constructed in a limited manner for estimating the joint coordinates of two ankles. In this case, the coordinate estimation part 220 can estimate the coordinates of the joints J2 of the two ankles from the intermediate data DT1 by using the third training model M3. In the case where there are other joints located outside the field of view, the estimation using the fourth training model M4 or the fifth training model M5 is performed in parallel, and the estimation results are combined.
[0047] It should be noted that in Figure 4 The same is true in the example, the intermediate data DT1 is Figure 3 The example of similarly represents the data of joint coordinates in two-dimensional coordinates. Figure 4 As shown in (e), the third training model M3 to the fifth training model M5 can estimate the coordinates of the joints as three-dimensional coordinates by inputting the intermediate data DT1 of the time series. Figure 4 As shown in (e), the final data DT3 estimated using the third training model M3 to the fifth training model M5 includes the estimated results of the coordinates of all joints of the object, including the coordinates of Figure 4 The joints J2 of the two ankles are outside the field of view of the image A5 shown in (a).
[0048] according to Figure 4 Another example of joint coordinate estimation is shown, where the coordinates of the joints are estimated by using different training models, depending on which joint is outside the viewing angle of the image. It can be expected that by constructing each training model (in the case of the above example, the third training model M3, the fourth training model M4, and the fifth training model M5) that estimates the coordinates of a single joint (or a group of joints) in a limited manner, the size of each model can be reduced and the processing load can also be reduced.
[0049] Furthermore, for limited requests such as “estimate face position only”, results can be obtained with minimal processing load.
[0050] Figure 5 Depicted to illustrate the Figure 1 A diagram of another example of estimating joint coordinates in the example of . Figure 5 In the example, the training model 300 is a Figure 3 The first training model M1 and the second training model M2 shown are training models that jointly perform the function of the two-step estimation process, and include a sixth training model M6 that estimates the coordinates of the joints including the objects located outside the view from the time series input image 211. The coordinate estimation part 220 performs the estimation process by using the sixth training model M6.
[0051] In Figure 5 In the example shown in (a) of, the input image 211 includes the image A5 in which the joint J2 of both ankles is out of the viewing angle, as shown in Figure 3 (a) of. In the illustrated example, the input image 211 is basically a two-dimensional image, but the sixth training model M6 can estimate the coordinates of the joints as three-dimensional coordinates as shown in Figure 5 (c) of by inputting the input image 211 of the time series.
[0052] Figure 5 The sixth training model M6 shown in (b) of is a training model obtained by adding a time control element such as the second training model M2 shown in Figure 3 (b) of to the first training model M1 shown in Figure 3 (d) of. The coordinate estimation section 220 estimates the coordinates of the joints located inside and outside the viewing angle of the image A5 by using the sixth training model M6, that is, the coordinates of all joints including the joint J2 of both ankles.
[0053] As a result, as shown in Figure 5 (c) of, the final data DT4 estimated by using the 6th training model M6 includes the estimation results of the coordinates of all joints of the object, including the joint J2 of both ankles out of the viewing angle of the image A5 shown in Figure 5 (a) of.
[0054] Note that in the above embodiments of the present invention, the construction of the training model 300 performed by the information processing device 100 and the estimation of the full body posture of the object performed by the information processing device 200 can be independently executed. For example, the training model 300 can be pre-constructed by the information processing device 100, and any information processing device 200 can estimate the full body posture of the object based on the training model 300. In addition, for example, the information processing device 100 and the information processing device 200 can be installed as a single computer that can be connected to the training model 300.
[0055] In addition, in the embodiments of the present invention, the functions described as being implemented in the information processing device 100 and the information processing device 200 can be implemented in a server. For example, the image generated by the image pickup device is sent from the information processing device to the server, and the server can estimate the full body posture of the object.
[0056] In addition, the training model 300 of the embodiment of the present invention may be a model for estimating all joint positions of an object, or may be a model for estimating only some joint positions. In addition, the coordinate estimation part 220 of the present embodiment may estimate the positions of all joints of the object, or may estimate only the positions of some joints. Further, the three-dimensional pose estimation part 230 of the present embodiment may estimate the three-dimensional pose of the whole body of the object, or may estimate only a part of the three-dimensional pose such as the upper body.
[0057] In addition, in the embodiment of the present invention, a human is taken as an example, but the present invention is not limited to this example. For example, any object having a plurality of joints, such as an animal or a robot, may be a candidate object. The information processing device 200 of the present embodiment can be mounted on a robot, for example, for controlling the actions of the robot. In addition, the information processing device 200 according to the present embodiment can be used for monitoring suspicious persons, for example, by being installed on a surveillance camera device.
[0058] Figure 6 and Figure 7 is a flowchart showing a processing example according to an embodiment of the present invention.
[0059] Figure 6 Shows the process of constructing the training model 300 of the information processing device 100. First, the input part 110 of the information processing device 100 receives the input data 111 to be used for learning by the relational learning part 120, that is, the input data including the image and the coordinate information of the object joints (step S101). Here, since the joint coordinate information is defined in a range extended beyond the perspective of the image, even for an image in which some joints of the object are outside the perspective, the input data 111 includes the image and the coordinate information of all joints, including the joints located outside. Next, the relational learning part 120 constructs the training model 300 by learning the relationship between the image and the coordinate information in the input data 111 (step S102). In the information processing device 100, the output part 130 outputs the training model 300 constructed in the memory on the network, for example, in particular, the parameters of the training model 300 (step S103).
[0060] On the other hand, Figure 7Shows the process in which the information processing device 200 estimates joint coordinates from an image by using the training model 300. When the input section 210 of the information processing device 200 receives the input of a new input image 211 (step S201), the coordinate estimation section 220 estimates the coordinates of the joints from the image by using the training model 300 (step S202). Further, the three-dimensional pose estimation section 230 estimates the full-body pose of the object based on the estimation result of the joint coordinates by the coordinate estimation section 220 (step S203). In the information processing device 200, the output section 240 outputs data representing the estimated full-body pose of the object (step S204).
[0061] Although some embodiments of the present invention have been described in detail above with reference to the accompanying drawings, the present invention is not limited to these examples. Obviously, those having ordinary knowledge in the technical field to which the present invention pertains can make various modifications or changes within the scope of the technical idea described in the claims, and thus it is naturally understood that these modifications or changes also belong to the technical scope of the present invention.
[0062] [List of reference signs]
[0063] 10: System
[0064] 100, 200: Information processing device
[0065] 110, 210: Input section
[0066] 111: Input data
[0067] 120: Relationship learning section
[0068] 130, 240: Output section
[0069] 211: Input image
[0070] 212: Image pickup device
[0071] 220: Coordinate estimation section
[0072] 230: Three-dimensional pose estimation section
[0073] 300: Training model
Claims
1. An information processing device, comprising: a relationship learning section that learns the relationship between a first image of an object having a plurality of joints and coordinate information indicating the positions of the plurality of joints to construct a plurality of training models, the coordinate information indicating the positions of the plurality of joints being defined within a range extended to a perspective larger than that of the first image, and each training model estimates the coordinate information of at least one joint located outside the perspective of a newly acquired second image of the object, wherein the coordinates of the joints are estimated by using different training models depending on which joint is located outside the perspective of the image.
2. The information processing device according to claim 1, wherein, the relationship learning section constructs the training model that estimates the coordinate information of the plurality of joints including the at least one joint.
3. The information processing device according to claim 1 or 2, wherein, the relationship learning section learns the relationship between a plurality of first images acquired in time series and the coordinate information of the plurality of first images to construct the training model, and the training model estimates the three-dimensional coordinate information of the plurality of joints in the second image.
4. The information processing device according to claim 1, wherein, the plurality of joints includes a first joint and a second joint, and the training models include a first training model and a second training model. The first training model estimates the coordinate information of the first joint when the first joint is located outside the perspective of the second image, and the second training model estimates the coordinate information of the second joint when the second joint is located outside the perspective of the second image.
5. The information processing device according to claim 1, wherein, the training models include a third training model and a fourth training model. The third training model estimates the coordinate information of the joints located within the perspective of the second image, and the fourth training model estimates the coordinate information of at least one joint located outside the perspective of the second image.
6. An information processing device, comprising: a coordinate estimation section that estimates the coordinate information of at least one joint located outside the perspective of a newly acquired second image of the object based on a plurality of training models constructed by learning the relationship between a first image of an object having a plurality of joints and coordinate information indicating the positions of the plurality of joints, the coordinate information indicating the positions of the plurality of joints being defined within a range extended to a perspective larger than that of the first image, wherein the coordinates of the joints are estimated by using different training models depending on which joint is located outside the perspective of the image.
7. The information processing device according to claim 6, wherein, the coordinate estimation section estimates the coordinate information of the plurality of joints including the at least one joint.
8. The information processing device according to claim 6 or 7, wherein, the second image includes a plurality of images acquired in time series, and the coordinate estimation section estimates the three-dimensional coordinate information of the plurality of joints.
9. The information processing device according to claim 6, wherein, the plurality of joints includes a first joint and a second joint, The plurality of training models includes a first training model and a second training model. When the first joint is outside the view angle of the second image, the first training model estimates the coordinate information of the first joint; and when the second joint is outside the view angle of the second image, the second training model estimates the coordinate information of the second joint, and the coordinate estimation unit estimates the coordinate information of the first joint based on the first training model when the first joint is outside the view angle of the second image; and estimates the coordinate information of the second joint based on the second training model when the second joint is outside the view angle of the second image.
10. The information processing apparatus according to claim 6, wherein, the plurality of training models includes a third training model and a fourth training model. The third training model estimates the coordinate information of joints within the view angle of the second image, and the fourth training model estimates the coordinate information of at least one joint outside the view angle of the second image, and the coordinate estimation unit estimates the coordinate information of joints within the view angle of the second image based on the third training model, and estimates the coordinate information of at least one joint outside the view angle of the second image based on the fourth training model.
11. An information processing method, comprising: a step of constructing a plurality of training models by learning the relationship between a first image of an object having a plurality of joints and coordinate information indicating the positions of the plurality of joints, the coordinate information indicating the positions of the plurality of joints being defined in a range extended to be larger than the view angle of the first image, each training model being used to estimate the coordinate information of at least one joint outside the view angle of a newly acquired second image of the object; and a step of estimating the coordinate information of at least one joint outside the view angle of the second image based on the training models, wherein the coordinates of the joints are estimated by using different training models depending on which joint is outside the view angle of the image.
12. A non-transitory storage medium including a program that causes a computer to implement: a function of constructing a plurality of training models by learning the relationship between a first image of an object having a plurality of joints and coordinate information indicating the positions of the plurality of joints, the coordinate information indicating the positions of the plurality of joints being defined in a range extended to be larger than the view angle of the first image, each training model being used to estimate the coordinate information of at least one joint outside the view angle of a newly acquired second image of the object, wherein the coordinates of the joints are estimated by using different training models depending on which joint is outside the view angle of the image.
13. A non-transitory storage medium including a program that causes a computer to implement: A step of estimating coordinate information of at least one joint located outside a viewing angle of a newly acquired second image of an object, based on a plurality of trained models constructed by learning a relationship between a first image of the object having a plurality of joints and coordinate information indicating positions of the plurality of joints, the coordinate information indicating the positions of the plurality of joints being defined in a range extended to be larger than a viewing angle of the first image, wherein coordinates of the joints are estimated by using different trained models, depending on which joint is located outside the viewing angle of the image.
Citation Information
Patent Citations
Motion model learning device, three-dimensional attitude estimation device, motion model learning method, three-dimensional attitude estimation method and program
JP2012083955A
Learning data generation apparatus, learning apparatus, estimation apparatus, learning data generation method, and computer program
JP2018129007A
Learning data generation device, estimation device, estimation method, and computer program
JP2019016164A