Learning program, generation program, learning method and generation method
The learning program addresses the inefficiencies of manual joint definition adjustments by using a regression model to generate 3D skeletal data that matches existing datasets from any 3D body CG model, enhancing recognition accuracy and efficiency.
Patent Information
- Application Number
- JP2024502743
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-02-28
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-02-28
AI Technical Summary
Existing methods for generating 3D skeletal data require manual adjustment of joint definitions, which is time-consuming and prone to errors, especially when dealing with large datasets or varying body shapes.
A learning program that uses a regression model to generate 3D skeletal data matching the joint definition of an existing dataset from any 3D body CG model, eliminating the need for manual adjustments.
This approach allows for efficient generation of 3D skeletal data that aligns with existing datasets, improving recognition accuracy and reducing the complexity of data collection and processing.
Smart Images

Figure 0007673869000002 
Figure 0007673869000003 
Figure 0007673869000004
Abstract
Description
[Technical field]
[0001] The present invention relates to a learning program and the like. [Background technology]
[0002] Regarding the detection of three-dimensional human movements, 3D sensing technology has been established that uses multiple 3D laser sensors to detect a person's 3D skeletal coordinates with an accuracy of ±1 cm. This 3D sensing technology is expected to be applied to gymnastics judging support systems, and to other sports and fields. The method using 3D laser sensors is referred to as the laser method.
[0003] In the laser method, a laser is emitted about 2 million times per second, and the depth and information of each irradiation point, including the target person, is calculated based on the laser's travel time (Time of Flight: ToF). The laser method can obtain highly accurate depth data, but has the disadvantage that the hardware is complex and expensive because the configuration and processing of the laser scan and ToF measurement are complicated.
[0004] Instead of the laser method, 3D skeletal recognition may be performed using an image method. The image method uses a CMOS (Complementary Metal Oxide Semiconductor) imager to acquire RGB (Red Green Blue) data for each pixel, and an inexpensive RGB camera can be used. With the recent improvement of deep learning technology, the accuracy of 3D skeletal recognition is also improving.
[0005] In image-based 3D skeleton recognition based on deep learning, it is necessary to prepare a large amount of training data to train the skeleton recognition model, for example, the training data includes body images, 3D joint positions (or 2D joint positions), and camera parameters.
[0006] As a method for collecting the above training data, for example, there is a Motion Capture system that measures body movements by attaching special markers to the surface. There are also methods that semi-automatically annotate joints from images taken from multiple viewpoints without using special equipment.
[0007] In general, if data with conditions not included in the training data is input to a trained skeleton recognition model, the skeleton may not be recognized correctly. For this reason, a method is adopted in which new training data is collected as necessary and combined with an existing training dataset to retrain the skeleton recognition model. The training data conditions include body posture, camera angle, appearance, etc. Appearance refers to information on the appearance of the foreground and background. In the following description, an existing training dataset is referred to as an "existing dataset."
[0008] Here, when collecting new training data using a motion capture system or a method for semi-automatically annotating joints from images taken from multiple viewpoints, images must be taken on-site. In addition, the conditions under which data can be collected are limited, and the cost of data collection is high, making it difficult to efficiently collect training data under arbitrary conditions.
[0009] In response to this, a method of synthesizing virtual training data using a 3D body CG (Computer Graphics) model can be considered to efficiently collect training data under any conditions and reinforce the existing dataset. In this case, it is required to adjust the number of joints and the joint positions of the 3D body CG model so that it matches the joint definition of the existing dataset. This is because retraining using training data generated using a 3D body CG model with a joint definition different from that of the existing dataset will result in a decrease in the recognition accuracy of the trained skeleton recognition model.
[0010] FIG. 12 is a diagram showing an example of joint definitions of an existing data set and a 3D body CG model. For example, joint definition 5a of the existing data set is defined with 21 joints. On the other hand, joint definition 5b of the 3D body CG model is defined with 24 joints. In the conventional method, a user visually checks the difference between joint definition 5a and joint definition 5b, and manually adjusts the number of joints and the joint positions of joint definition 5b so that joint definition 5a and joint definition 5b match, thereby creating joint definition 5c. The number of joints in joint definition 5c is the same as the number of joints in joint definition 5a (21 joints).
[0011] Here, if the existing data set contains position information of markers attached to the surface of a person, such position information can be used to automatically calculate a 3D body CG model that defines joints corresponding to the existing data set. In the following explanation, the position information of the markers attached to the surface is referred to as marker position information.
[0012] FIG. 13 is a diagram for explaining an example of a conventional technique. For example, when training data is generated using a motion capture system, marker position information 6c is obtained in addition to image data 6a and 3D skeletal data 6b in the existing data set. By using the marker position information 6c to execute 3D body CG model estimation using the marker position information, a 3D body CG model 7 can be obtained. The joint definition of the 3D body CG model 7 is similar to the joint definition of the existing data set.
[0013] By using the conventional technique of Fig. 13, a pair of 3D skeletal data 6b and a 3D body CG model 7 can be obtained. By using this pair, a regression model that generates 3D skeletal data (3D joint information) from the 3D body CG model can be generated. In addition, by using this regression model, 3D skeletal data can be generated from various 3D body CG models, so that 3D joint positions to be used in training data under any condition can be easily added. [Prior art documents] [Patent documents]
[0014] [Patent Document 1] JP 2014-044653 A [Non-patent literature]
[0015] [Non-Patent Document 1] M. Naureen et al., “AMASS: Archive of Motion Capture as Surface Shapes,” ICCV2019 Summary of the Invention [Problem to be solved by the invention]
[0016] In the above-mentioned method of manually adjusting the joint definition of a 3D body CG model, the adjustment results of the internal body joint positions, whose positions are difficult to determine from the body shape, are subjective, and the accuracy of the joint definition depends on the skill of the operator. In addition, when preparing a large number of 3D body CG models, it is not realistic to manually adjust the joint definition of each 3D body CG model.
[0017] If an existing dataset contains marker position information, it is possible to automatically generate a 3D body CG model that corresponds to the joint definitions of the existing dataset by performing 3D body CG model estimation using the marker position information, but this cannot be used if the marker position information is not included.
[0018] Furthermore, if retraining is performed using training data generated using a 3D body G model with joint definitions that differ from the joint definitions of the existing dataset, the recognition accuracy of the trained skeletal recognition model will decrease.
[0019] For this reason, there is a need to generate 3D skeletal data that matches the joint definitions of existing datasets from any 3D body CG model, even if the existing datasets do not contain marker position information.
[0020] In one aspect, the present invention aims to provide a learning program, a generation program, a learning method, and a generation method that can generate 3D skeletal data that conforms to the joint definitions of an existing dataset from any 3D body CG model. [Means for solving the problem]
[0021] In the first proposal, a computer executes the following process. The computer obtains a model of an object composed of a three-dimensional surface. The computer generates image data in which the model of the object is rendered. The computer specifies three-dimensional skeletal data of the rendered image data by inputting the rendered image data to a first learning device that has been trained using the image data of the object in the training data as an explanatory variable and the three-dimensional skeletal data of the training data as an objective variable. The computer executes training of a second learning device using the specified three-dimensional skeletal data as an objective variable and the model of the object as an explanatory variable. Effect of the Invention
[0022] 3D skeletal data that matches the joint definitions of existing datasets can be generated from any 3D body CG model. [Brief description of the drawings]
[0023] [Figure 1] FIG. 1 is a diagram for explaining the process of the information processing device according to the present embodiment. [Diagram 2] FIG. 2 is a diagram for explaining the effect of the information processing device according to the present embodiment. [Diagram 3] FIG. 3 is a functional block diagram illustrating a configuration of an information processing device according to the present embodiment. [Figure 4] FIG. 4 is a diagram illustrating an example of a data structure of an existing data set. [Diagram 5] FIG. 5 is a diagram showing an example of a 3D body CG model. [Figure 6] FIG. 6 is a diagram illustrating an example of a data structure of the regression data set. [Figure 7] FIG. 7 is a diagram for explaining the first removal process. [Figure 8] FIG. 8 is a diagram for explaining the second removal process. [Figure 9] FIG. 9 is a flowchart (1) showing a processing procedure of the information processing device according to the present embodiment. [Figure 10] FIG. 10 is a flowchart (2) showing the processing procedure of the information processing device according to the present embodiment. [Figure 11] FIG. 11 is a diagram illustrating an example of a hardware configuration of a computer that realizes the same functions as the information processing apparatus of the embodiment. [Figure 12] FIG. 12 is a diagram showing an example of joint definitions of an existing data set and a 3D body CG model. [Figure 13] FIG. 13 is a diagram for explaining an example of the conventional technology. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0024] Hereinafter, the learning program, the generating program, the learning method, and the generating method disclosed in the present application will be described in detail with reference to the accompanying drawings. Note that the present invention is not limited to these embodiments. EXAMPLES
[0025] An example of processing by the information processing device according to this embodiment will be described. Fig. 1 is a diagram for explaining processing by the information processing device according to this embodiment. The information processing device trains a 3D skeleton recognizer M1 in advance using an existing data set 50. The existing data set 50 includes a plurality of training data.
[0026] For example, the training data includes a set of image data, 3D skeletal data, and camera parameters. The image data is image data of a person captured by a camera. The 3D skeletal data is information consisting of the three-dimensional joint positions (coordinates of the joints) of a person. The camera parameters are parameters of the camera that captured the image (image data).
[0027] The information processing device trains a 3D skeleton recognizer M1 using image data included in the training data as input (explanatory variables) and 3D skeleton data as output (correct label, objective variable). The 3D skeleton recognizer M1 is a neural network or the like. In the following description, the trained 3D skeleton recognizer M1 is simply referred to as the 3D skeleton recognizer M1. When image data is input to the 3D skeleton recognizer M1, 3D skeleton data is output.
[0028] The information processing device acquires a 3D body CG model 10. The 3D body CG model 10 is a model of an object (person) composed of three-dimensional surfaces. For example, the multiple surfaces that make up the model of the object (person) are multiple meshes. The information processing device generates synthetic image data 11 by rendering the 3D body CG model 10. The information processing device infers 3D skeletal data 12 by inputting the synthetic image data 11 to a 3D skeleton recognizer M1. The information processing device repeatedly executes this process to obtain multiple pairs of the 3D body CG model 10 and the 3D skeletal data 12.
[0029] After the above processing, the information processing device trains a regression model M2 using the 3D body CG model 10 as an input (explanatory variable) and the 3D skeletal data 12 as an output (answer label, objective variable). The regression model M2 is a neural network or the like. By inputting an arbitrary 3D body CG model to the trained regression model M2, it is possible to generate 3D skeletal data corresponding to the arbitrary 3D body CG model and conforming to the joint definition of the existing data set 50.
[0030] For example, the information processing device adds 3D skeletal data inferred by inputting an arbitrary 3D body CG model into a trained regression model M2 to an existing data set 50. When adding the inferred 3D skeletal data, the information processing device may generate paired image data based on the arbitrary 3D body CG model. The information processing device retrains the 3D skeletal recognizer M1 using the existing data set 50 to which the inferred 3D skeletal data has been added.
[0031] As described above, the information processing device according to this embodiment generates synthetic image data 11 by rendering the 3D body CG model 10, and infers the 3D skeletal data 12 by inputting the synthetic image data 11 to the 3D skeleton recognizer M1. The information processing device repeatedly executes this process to obtain a plurality of pairs of the 3D body CG model 10 and the 3D skeletal data 12. Here, rendering is a process of projecting the 3D body CG model 10 onto an image using image processing. For example, the 3D body CG model 10 is converted into 2D image information by rendering.
[0032] The information processing device also trains a regression model M2 using the 3D body CG model 10 as input and the 3D skeletal data 12 as output. By inputting an arbitrary 3D body CG model to the trained regression model M2, it is possible to generate 3D skeletal data that corresponds to the arbitrary 3D body CG model and that matches the joint definition of the existing data set 50.
[0033] As a result, even if the existing data set does not have marker position information, 3D skeletal data that matches the joint definition of the existing data set 50 can be generated from any 3D body CG model. For example, a 3D body CG model with conditions not included in the original training data can be input to the regression model M2 to easily generate 3D skeletal data. By adding such 3D skeletal data to the existing data set 50 and retraining the 3D skeleton recognizer M1, it is possible to improve the recognition accuracy when image data with conditions not included in the original training data is input.
[0034] Fig. 2 is a diagram for explaining the effect of the information processing device according to the present embodiment. In Fig. 2, 3D skeletal data 8a is data based on the joint definition of an existing data set. 3D skeletal data 8b is 3D skeletal data of a 3D body CG model that does not use the regression model M2. 3D skeletal data 8c is 3D skeletal data of a 3D body CG model that is estimated using the regression model M2.
[0035] Comparing the 3D skeletal data 8a and the 3D skeletal data 8b, a difference in joints of about 5 to 7 cm occurs in areas A1 and A2. Therefore, if the 3D skeletal data 8b is added to the existing data set 50 and the 3D skeletal recognizer M1 is retrained, the recognition accuracy of the 3D skeletal recognizer M1 will decrease.
[0036] On the other hand, when the 3D skeletal data 8a and the 3D skeletal data 8c are compared, the difference in the joints in the areas A3 and A4 is about 1 cm. Therefore, the 3D skeletal data 8c is 3D skeletal data that matches the joint definitions of the existing data set 50, and by adding the 3D skeletal data 8c to the existing data set 50 and retraining the 3D skeleton recognizer M1, the recognition accuracy of the 3D skeleton recognizer M1 can be improved.
[0037] Next, a configuration example of an information processing device that executes the process described in Fig. 1 will be described. Fig. 3 is a functional block diagram showing the configuration of an information processing device according to this embodiment. As shown in Fig. 3, this information processing device 100 has a communication unit 110, an input unit 120, a display unit 130, a storage unit 140, and a control unit 150.
[0038] The communication unit 110 executes data communication with an external device etc. via a network. The communication unit 110 receives an existing data set 50 etc. from an external device.
[0039] The input unit 120 is an input device that accepts operations from a user, and is realized by, for example, a keyboard, a mouse, and the like.
[0040] The display unit 130 is a display device for outputting the processing results of the control unit 150, and is realized by, for example, a liquid crystal monitor, a printer, or the like.
[0041] The storage unit 140 is a storage device that stores various types of information, and is realized by, for example, a semiconductor memory element such as a random access memory (RAM) or a flash memory, or a storage device such as a hard disk or an optical disk.
[0042] The memory unit 140 has an existing data set 50, a 3D skeleton recognizer M1, a regression model M2, a 3D body CG model table 141, and a regression data set 142.
[0043] The existing dataset 50 has a plurality of training data used when training the 3D skeleton recognizer M1. FIG. 4 is a diagram showing an example of the data structure of the existing dataset. As shown in FIG. 4, the existing dataset 50 associates an item number, image data, 3D skeleton data, and camera parameters. The item number is a number that identifies a record (training data) of the existing dataset 50. The image data is image data of a person captured by a camera. The 3D skeleton data is information composed of the three-dimensional joint positions of the person. The camera parameters are parameters of the camera that captured the image.
[0044] The 3D skeleton recognizer M1 is a learning model that outputs 3D skeleton data when image data is input. The 3D skeleton recognizer M1 is a neural network or the like.
[0045] The regression model M2 is a learning model that outputs 3D skeletal data when a 3D body CG model is input. The regression model M2 is a neural network or the like.
[0046] The 3D body CG model table 141 has a plurality of 3D body CG models. Fig. 5 is a diagram showing an example of a 3D body CG model. As shown in Fig. 5, the 3D body CG model table 141 includes 3D body CG models mo1, mo2, mo3, mo4, mo5, mo6, and mo7 of various conditions. Fig. 5 shows the 3D body CG models mo1 to mo7 as an example, but is not limited to this.
[0047] The regression dataset 142 stores a plurality of pairs of a 3D body CG model and 3D skeletal data. FIG. 6 is a diagram showing an example of the data structure of the regression dataset. As shown in FIG. 6, the regression dataset 142 associates an item number with a 3D body CG model and 3D skeletal data. The item number is a number for identifying a record of the regression dataset 142. The 3D body CG model is, for example, data of the 3D body CG model described in FIG. 5. The 3D skeletal data is 3D skeletal data obtained by inputting the 3D body CG model into the trained regression dataset 142.
[0048] Returning to the description of Fig. 3, the control unit 150 is realized by a processor such as a CPU (Central Processing Unit) or an MPU (Micro Processing Unit) executing various programs stored in a storage device inside the information processing device 100 using a RAM or the like as a working area. The control unit 150 may also be realized by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).
[0049] The control unit 150 has an acquisition unit 151, a first learning execution unit 152, an image generation unit 153, an inference unit 154, a second learning execution unit 155, a skeleton generation unit 156, and an additional processing unit 157.
[0050] The acquisition unit 151 acquires various data from an external device or the like, and stores the acquired data in the storage unit 140. The acquisition unit 151 may acquire various data from the input unit 120. For example, the acquisition unit 151 acquires an existing data set 50, and stores the acquired existing data set 50 in the storage unit 140. The acquisition unit 151 acquires data of a 3D body CG model, and stores the acquired data of the 3D body CG model in the 3D body CG model table 141.
[0051] The first learning execution unit 152 trains (performs machine learning) the 3D skeleton recognizer M1 using training data stored in the existing data set 50. The first learning execution unit 152 trains the 3D skeleton recognizer M1 using image data of the training data as input and 3D skeleton data as output. For example, the first learning execution unit 152 trains parameters of the 3D skeleton recognizer M1 based on a backpropagation learning method or the like.
[0052] Furthermore, when new training data is added to the existing data set 50, the first learning execution unit 152 retrains the 3D skeleton recognizer M1 using the training data stored in the existing data set 50 as well.
[0053] The image generating unit 153 acquires the data of the 3D body CG model from the 3D body CG model table 141, and generates composite image data by executing rendering on the 3D body CG model.
[0054] For example, the image generating unit 153 performs rendering to generate a plurality of image data from N viewpoints, and combines the plurality of image data to generate composite image data. N corresponds to the number of image data (a natural number equal to or greater than 1) input to the 3D skeleton recognizer M1, and is set in advance.
[0055] The image generating unit 153 outputs to the inference unit 154 a set of the generated synthetic image data and the 3D body CG model that is the subject of rendering.
[0056] The inference unit 154 infers 3D skeletal data by inputting the synthetic image data acquired from the image generation unit 153 to the trained 3D skeleton recognizer M1. The inference unit 154 associates the inferred 3D skeletal data with the data of the 3D body CG model corresponding to the synthetic image data, and registers them in the regression dataset 142.
[0057] Incidentally, when the inference unit 154 infers 3D skeletal data, the inference unit 154 may perform a process of determining whether or not each joint position of the 3D skeletal data is appropriate, and excluding joints at inappropriate positions. For example, the inference unit 154 executes a first removal process or a second removal process described below. The inference unit 154 may execute either the first removal process or the second removal process, or may execute both removal processes.
[0058] The first removal process executed by the inference unit 154 will be described. FIG. 7 is a diagram for explaining the first removal process. The inference unit 154 compares the 3D skeletal data 9a conforming to the joint definition of the existing data set 50 prepared in advance with the 3D skeletal data 9b generated by the skeleton generation unit 156. When the distance between a certain joint (for example, ankle) in the 3D skeletal data 9a and a certain joint in the 3D skeletal data 9b is equal to or greater than a threshold, the inference unit 154 removes the certain joint in the 3D skeletal data 9b. The same applies to joints other than the ankle.
[0059] The second removal process executed by the inference unit 154 will be described. Fig. 8 is a diagram for explaining the second removal process. The inference unit 154 compares the surface positions of the 3D body CG model 20 corresponding to the joint definitions of the existing data set 50 with the positions of the joints of the 3D skeletal data, and among the joints of the 3D skeletal data, the joints existing inside the 3D body CG model 20 are left, and the joints existing outside are removed. In the example shown in Fig. 8, the inference unit 154 leaves the joint 21b existing inside the 3D body CG model 20, and removes the joint 21a existing outside.
[0060] The image generation unit 153 and the inference unit 154 repeatedly execute the above process on the data of each 3D body CG model stored in the 3D body CG model table 141, thereby generating multiple pairs of 3D skeletal data and 3D body CG models, and storing them in the regression dataset 142.
[0061] The second learning execution unit 155 trains the regression model M2 based on a set (training data) of the 3D skeletal data and the 3D body CG model stored in the regression data set 142. For example, the second learning execution unit 155 trains the regression model M2 parameters (weights J i Search for the value of
[0062]
number
[0063] In formula (1), "N data " is the number of regression training data. For example, the number of regression training data is the number of pairs of 3D skeleton data and 3D body CG models stored in the regression data set 142. "N joint " is the number of joints in the joint definition of the existing dataset 50.
[0064] "N vert " is the number of 3D vertices of the 3D body CG model, in formula (1), "V i " is the set of vertex coordinates of the 3D body CG model in the i-th training data (a set of 3D skeleton data and a 3D body CG model) (3 × N vert matrix). "J i " is V i Weights (N vert dimensional vector).
[0065] In formula (1), "p i,j “ is the 3D coordinate of the jth joint in the 3D skeletal data in the i-th training data (3D vector). “λ” is a weighting coefficient (scalar) for the regularization term.
[0066] The second learning execution unit 155 optimizes the equation (1) using the non-negative least squares method to obtain J j ≧0, and the 3D joints can be regressed to the inside of the vertex group Vi of the 3D body CG model.
[0067] The second learning execution unit 155 calculates a set of vertices V i,j ⊆V i Assuming that only V affects the determination of joint position, i,j From p i,j By "regressing" the ankle position, it is possible to generate more stable 3D skeletal data. For example, it is possible to prevent the position of unrelated head vertices from being affected when regressing the ankle position.
[0068] Here, we have described the process in which the second learning execution unit 155 trains the regression model M2 based on the objective function of equation (1), but training may also be performed using a backpropagation learning method with a 3D body CG model as input and 3D skeletal data as output.
[0069] Next, a description will be given of the processing of the skeleton generation unit 156 and the additional processing unit 157. After the training of the regression model M2 is completed, the skeleton generation unit 156 and the additional processing unit 157 acquire data of an arbitrary 3D body CG model, and generate training data to be added to the existing data set 50. The data of the arbitrary 3D body CG model may be acquired from the input unit 120, or may be acquired from an external device.
[0070] When the skeleton generation unit 156 acquires an arbitrary 3D body CG model, the skeleton generation unit 156 inputs the acquired 3D body CG model to the regression model M2 to generate 3D skeleton data. The skeleton generation unit 156 outputs the arbitrary 3D body CG model and the generated 3D skeleton data to the additional processing unit 157.
[0071] The additional processing unit 157 adds the acquired 3D skeletal data as training data to the existing data set 50. The additional processing unit 157 may perform rendering on any 3D body CG model to generate image data paired with the 3D skeletal data and add it to the existing data set 50. In addition, the additional processing unit 157 may add camera parameters specified by an external device or the like to the existing data set 50.
[0072] The skeleton generation unit 156 and the addition processing unit 157 execute the above process each time they acquire an arbitrary 3D body CG model, generate new training data, and add the training data to the existing data set 50.
[0073] When new training data is added to the existing data set 50, the first learning execution unit 152 retrains the 3D skeleton recognizer M1 using the training data stored in the existing data set 50 as well.
[0074] Next, an example of a processing procedure of the information processing device 100 according to this embodiment will be described. Fig. 9 is a flowchart (1) showing the processing procedure of the information processing device according to this embodiment. As shown in Fig. 9, the first learning execution unit 152 of the information processing device 100 trains the 3D skeleton recognizer M1 based on the existing data set 50 (step S101).
[0075] The acquisition unit 151 of the information processing device 100 acquires 3D body CG models of various postures and registers them in the 3D body CG model table 141 (step S102). The image generation unit 153 of the information processing device 100 performs rendering on the 3D body CG model to generate synthetic image data (step S103).
[0076] The inference unit 154 of the information processing device 100 inputs the composite image data to the 3D skeleton recognizer M1 and infers 3D skeleton data (step S104). The inference unit 154 removes inappropriate joints included in the 3D skeleton data (step S105).
[0077] The inference unit 154 registers the pair of the 3D body CG model and the 3D skeletal data in the regression data set 142 (step S106). The second learning execution unit 155 of the information processing device 100 trains the regression model M2 based on the regression data set 142 (step S107).
[0078] 10 is a flowchart (2) showing the processing procedure of the information processing device according to this embodiment. The skeleton generating unit 156 of the information processing device 100 acquires an object model (any 3D body CG model) (step S201).
[0079] The skeleton generating unit 156 inputs the object model into the regression model M2, and generates 3D skeleton data (step S202).
[0080] The additional processing unit 157 of the information processing device 100 generates training data based on the relationship between the image data of the object model and the 3D skeletal data (step S203). The additional processing unit 157 adds the generated training data to the existing data set (step S204).
[0081] The first learning execution unit 152 of the information processing device 100 retrains the 3D skeleton recognizer M1 based on the existing data set 50 (step S205).
[0082] Next, the effect of the information processing device 100 according to this embodiment will be described. The information processing device 100 generates synthetic image data by rendering a 3D body CG model, and infers 3D skeletal data by inputting the synthetic image data to a 3D skeleton recognizer M1. The information processing device 100 inputs a 3D body CG model, outputs the inferred 3D skeletal data, and trains a regression model M2. This makes it possible to input an arbitrary 3D body CG model to the trained regression model M2, and generate 3D skeletal data that corresponds to the arbitrary 3D body CG model and is compatible with the joint definition of the existing data set 50.
[0083] The information processing device 100 inputs a 3D body CG model with conditions not included in the original training data into the trained regression model M2 to generate 3D skeletal data, and adds the 3D skeletal data to the existing data set 50. By retraining the 3D skeletal recognizer M1, the information processing device 100 can improve the recognition accuracy when image data with conditions not included in the original training data is input.
[0084] The information processing device 100 inputs the synthetic image data to the 3D skeleton recognizer M1, and when the 3D skeleton data is inferred, performs a process of determining whether or not each joint position of the 3D skeleton data is appropriate, and excluding joints in inappropriate positions. This makes it possible to prevent outlier joints from being registered in the regression dataset 142, and improve the accuracy of the regression model M2 trained using the regression dataset 142.
[0085] When training the regression model M2, the information processing device 100 searches for parameters that reduce the absolute value of the difference between the multiplication value of the vertex coordinates of the object model and the parameters and the positions of each joint of the inferred three-dimensional skeleton data. For example, the information processing device 100 searches for parameters (weights J i This makes it possible to generate a model that can accurately infer 3D skeletal data from a 3D body CG model.
[0086] The information processing device 100 can infer 3D skeletal data by inputting a 3D body CG model into the trained regression model M2, thereby reducing the processing load compared to directly analyzing the 3D body CG model and inferring 3D skeletal data.
[0087] Next, an example of a hardware configuration of a computer that realizes the same functions as the information processing device 100 described in the above embodiment will be described. Fig. 11 is a diagram showing an example of a hardware configuration of a computer that realizes the same functions as the information processing device of the embodiment.
[0088] 11, the computer 200 has a CPU 201 that executes various arithmetic processes, an input device 202 that accepts data input from a user, and a display 203. The computer 200 also has a communication device 204 that transmits and receives data to and from external devices, etc., via a wired or wireless network, and an interface device 205. The computer 200 also has a RAM 206 that temporarily stores various information, and a hard disk device 207. The devices 201 to 207 are connected to a bus 208.
[0089] The hard disk device 207 stores an acquisition program 207a, a first learning execution program 207b, an image generation program 207c, an inference program 207d, a second learning execution program 207e, a skeleton generation program 207f, and an additional processing program 207g. The CPU 201 reads out each of the programs 207a to 207g and expands them in the RAM 206.
[0090] The acquisition program 207a functions as an acquisition process 206a. The first learning execution program 207b functions as a first learning execution process 206b. The image generation program 207c functions as an image generation process 206c. The inference program 207d functions as an inference process 206d. The second learning execution program 207e functions as a second learning execution process 206e. The skeleton generation program 207f functions as a skeleton generation process 206f. The addition processing program 207g functions as an addition processing process 206g.
[0091] The processing of the acquisition process 206a corresponds to the processing of the acquisition unit 151. The processing of the first learning execution process 206b corresponds to the processing of the first learning execution unit 152. The processing of the image generation process 206c corresponds to the processing of the image generation unit 153. The processing of the inference process 206d corresponds to the processing of the inference unit 154. The processing of the second learning execution process 206e corresponds to the processing of the second learning execution unit 155. The processing of the skeleton generation process 206f corresponds to the processing of the skeleton generation unit 156. The processing of the addition processing process 206g corresponds to the processing of the addition processing unit 157.
[0092] It should be noted that each of the programs 207a to 207g does not necessarily have to be stored in the hard disk device 207 from the beginning. For example, each of the programs may be stored in a "portable physical medium" such as a flexible disk (FD), CD-ROM, DVD, magneto-optical disk, or IC card that is inserted into the computer 200. Then, the computer 200 may read and execute each of the programs 207a to 207g. [Explanation of symbols]
[0093] 50 Existing Data Sets 100 Information processing device 110 Communications Department 120 Input section 130 Display section 140 Storage section 150 Control section 151 Acquisition Department 152 First Learning Execution Department 153 Image Generation Unit 154 Reasoning part 155 Second Learning Execution Department 156 Skeleton Generation Unit 157 Additional Processing Unit
Claims
1. Obtaining a model of an object consisting of three-dimensional surfaces; generating image data in which the model of the object is rendered; inputting the rendered image data into a first learning device that has been trained using image data of an object in training data as an explanatory variable and three-dimensional skeletal data of the training data as a target variable, thereby identifying three-dimensional skeletal data of the rendered image data; The identified three-dimensional skeleton data is used as a target variable, and the object model is used as an explanatory variable, and learning is performed by a second learning device. A learning program that causes a computer to execute a process.
2. The learning program according to claim 1, characterized in that the three-dimensional joint positions and the number of joints set in the model of the object composed of the three-dimensional surface are different from the three-dimensional joint positions and the number of joints set in the three-dimensional skeletal data of the training data.
3. The learning program according to claim 1, further comprising causing the computer to execute a process of removing outlier joints of skeletal data based on reference positions of the joints of the object identified based on a model of the object composed of the three-dimensional surface and joint positions of the three-dimensional skeletal data of the rendered image data.
4. The process of executing learning of the second learning device includes: The learning program according to claim 1, characterized in that parameters are searched for such that the absolute value of the difference between the multiplied value of the vertex coordinates of the object model and the parameters and each joint position of the three-dimensional skeletal data of the rendered image data is small.
5. inputting image data in which a model of a first object composed of a three-dimensional surface is rendered into a first learning device trained with first training data, thereby generating three-dimensional skeletal data as a response variable, and acquiring a second learning device trained with the model of the first object as an explanatory variable; generating second training data consisting of a three-dimensional skeleton of the second object model by inputting a second object model to the second learning device; Generate a data set consisting of the first training data and the second training data. A generating program that causes a computer to execute a process.
6. Obtaining a model of an object consisting of three-dimensional surfaces; generating image data in which the model of the object is rendered; inputting the rendered image data into a first learning device that has been trained using image data of an object in training data as an explanatory variable and three-dimensional skeletal data of the training data as a target variable, thereby identifying three-dimensional skeletal data of the rendered image data; The identified three-dimensional skeleton data is used as a target variable, and the object model is used as an explanatory variable, and learning is performed by a second learning device. A learning method characterized in that the processing is executed by a computer.
7. inputting image data in which a model of a first object composed of a three-dimensional surface is rendered into a first learning device trained with first training data, thereby generating three-dimensional skeletal data as a response variable, and acquiring a second learning device trained with the model of the first object as an explanatory variable; generating second training data consisting of a three-dimensional skeleton of the second object model by inputting a second object model to the second learning device; Generate a data set consisting of the first training data and the second training data. A generating method characterized by causing a computer to execute processing.
Citation Information
Patent Citations
Information processor, information processing system, information processing method and program
JP2014044653A
Coordinate detection device and learnt model
JP2019046007A
Model generation device, system, parameter computation device, model generation method, parameter computation method, and program
JP2020190959A
Distance image processing device, distance image processing system, distance image processing method, and distance image processing program
WO2018207365A1
Recognition method, recognition program, recognition device, learning method, learning program, and learning device
WO2020084667A1