3D Pose Detector Optimization Method, Device, and Storage Medium Based on a Unified Mask
Through the 3D attitude detector optimization method based on a unified mask, the skeleton feature map and body type feature map are used to calculate the difference between the unified mask, and the loss function value is calculated in combination with the pixel point weight, and the 3D attitude detector is optimized, which solves the problems of high demand for manual labeling data and low pose estimation accuracy in the prior art, and realizes accurate 3D attitude estimation without manual labeling data.
Patent Information
- Application Number
- CN202311675694.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-07
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-12-07
AI Technical Summary
In the three-dimensional pose estimation, the existing technology has problems such as high demand for manual labeling data, time-consuming and labor-consuming, low pose estimation accuracy, and inability to judge the left and right of the human body.
The 3D attitude detector optimization method based on a unified mask is adopted. By obtaining the skeleton feature map and body type feature map, the difference is calculated from the unified mask, and combining the weights of each pixel point, the pixel point deviation is calculated as the loss function value, and the 3D attitude detector is optimized to achieve accurate 3D attitude estimation without manual labeling of data.
It ensures that the 3D attitude detector has relatively accurate 3D attitude estimation without manual data labeling, solves the problems of high cost of labeling data and low pose estimation accuracy in the prior art, and improves the optimization efficiency of the detector.
Smart Images

Figure CN117876295B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of 3D pose recognition, and in particular to an optimization method, device and storage medium for a 3D pose detector based on a unified mask. Background Technique
[0002] Accurate three-dimensional human pose estimation plays a crucial role in many fields including human-computer interaction. Such as fields of human-computer interaction, robotics, sports performance analysis, human body reconstruction, and augmented / virtual reality. Compared with two-dimensional pose estimation, it requires more labeled data and imposes stronger geometric constraints. Therefore, obtaining three-dimensional data poses more challenges than two-dimensional data.
[0003] Some existing technologies have designed a weakly supervised three-dimensional pose estimation method, which uses unpaired two-dimensional pose information and geometric constraints provided by three-dimensional rigid body changes to provide a supervision signal. First, this method needs to use a detector to obtain predicted two-dimensional key points. The method will select specified key points according to the human skeleton connection relationship, connect them into line segments, and merge them as a two-dimensional skeleton graph. The two-dimensional key points will obtain three-dimensional key points by using a conversion network. The supervision signal of this method comes from the fact that the three-dimensional key points are similar to the unpaired two-dimensional poses after being projected onto the two-dimensional plane. On the other hand, it is required that the three-dimensional key points coincide with the original three-dimensional key points after passing through the process of conversion network - rigid transformation - projection - conversion network - inverse rigid transformation.
[0004] Some other existing technologies have designed an unsupervised three-dimensional pose estimation method, which uses two images with the same background to provide a supervision signal. Specifically, one image provides texture information through a deep network, and the other image uses a detector to obtain two-dimensional poses. Thereafter, the two-dimensional poses will be upsampled under the constraint of the human pose kinematics prior and re-projected back to two dimensions. The ellipse after affine transformation is calculated as the human skeleton graph and the Gaussian heat map containing human key points. Finally, the texture information and the human pose information will be used together to reconstruct the image, and the reconstruction loss is used as the supervision signal. In addition, the SMPL (Skinned Multi-Person Linear Model) prior trained from a large number of motion capture datasets is also considered as a constraint.
[0005] In the first existing technology, weakly supervised information is used, that is, unpaired two-dimensional poses are used as the prior of the human skeleton, and the two-dimensional poses still need to be manually labeled. In the second existing technology, the ellipse after affine transformation is used as the human skeleton, while the actual skeleton form is not an ellipse, which will lead to a loss of pose estimation accuracy. In addition, this method is more cumbersome in practice. In addition, in both of the above two existing technologies, there is a problem that it is impossible to judge the left and right of the human body, and a supervised post-processing step is still required.
[0006] However, obtaining labeled three-dimensional data is still a costly and time-consuming process. Summary of the Invention
[0007] The object of the present invention is to provide an optimization method, device and storage medium for a 3D pose detector based on a unified mask. Based on the obtained skeleton feature map and body shape feature map, subtracting the unified mask, and then cooperating with the weights of each pixel point, it is possible to ensure that the optimized 3D pose detector has relatively accurate 3D pose estimation on the basis of realizing the need for manual annotation of data.
[0008] The object of the present invention can be achieved by the following technical solutions:
[0009] An optimization method for a 3D pose detector based on a unified mask includes:
[0010] Step S1: Obtain the human key point coordinates generated by the 3D pose detector based on the three-dimensional feature map;
[0011] Step S2: Based on the obtained human key point coordinates, register with a pre-configured human skeleton to obtain the axis segments of all bones, where the two ends of the axis segment of the bone respectively correspond to two human key points;
[0012] Step S3: Calculate the distances from all points in the three-dimensional feature map to each bone, and generate the bone feature maps of all bones based on the obtained distances, where the distance from a point to a bone is specifically the distance from the point to the straight line where the axis segment of the bone is located;
[0013] Step S4: Synthesize the bone feature maps of all bones into a skeleton feature map;
[0014] Step S5: Generate a body shape feature map based on the skeleton feature map;
[0015] Step S6: Obtain the mask centroid of the pre-stored unified mask, obtain the distance from each pixel point in the foreground area of the skeleton feature map to the mask centroid as the skeleton weight of the pixel point, and obtain the distance from each pixel point in the foreground area of the body shape feature map to the mask centroid as the body shape weight of the pixel point,
[0016] Obtain the shortest distance from each pixel point in the background area of the skeleton feature map to the foreground area as the skeleton weight of the pixel point, and obtain the shortest distance from each pixel point in the background area of the body shape feature map to the foreground area as the body shape weight of the pixel point;
[0017] Step S7: Based on the skeleton feature map and the skeleton weights of all its pixel points, and the body shape feature map and the body shape weights of all its pixel points, combine the pixel values of each pixel point in the unified mask to calculate the pixel point deviation as the loss function value;
[0018] Step S8: Optimize the 3D pose detector based on the obtained loss function value.
[0019] In the step S3, for the bone feature map of a single bone, its generation process includes:
[0020] Calculating the Euclidean distance from all points in the three-dimensional feature map to the bone;
[0021] Generating the pixel value of the corresponding pixel point based on the Euclidean distance from each point to the bone;
[0022] Obtaining the bone feature map of the bone based on the pixel values of each point.
[0023] Both the skeleton feature map and the bone feature map are two-dimensional images. In the step S3, for the bone feature map of a single bone, its generation process includes:
[0024] Projecting the three-dimensional feature map onto a two-dimensional plane to obtain a two-dimensional feature map, and obtaining the projection of the axis segment of each bone in the two-dimensional plane as a projection line segment;
[0025] Calculating the Euclidean distance from all points in the two-dimensional feature map to the bone, where the distance from a point to the bone is specifically the distance from the point to the straight line where the projection line segment of the bone is located;
[0026] Generating the pixel value of the corresponding pixel point based on the Euclidean distance from each point to the bone;
[0027] Obtaining the bone feature map of the bone based on the pixel values of each pixel point.
[0028] In the step S5, the body shape feature map is obtained by processing through the U-Net network, where the skeleton feature map is used as the input of the U-Net network.
[0029] In the human skeleton, the axis segments of two connected bones share a human key point.
[0030] The step S4 specifically includes:
[0031] Registering the bone feature maps of all bones;
[0032] Summing the pixel values of the same pixel point in all bone feature maps as the pixel value of the corresponding pixel point in the preliminary skeleton feature map;
[0033] Normalizing the pixel values of all pixel points in the preliminary skeleton feature map to obtain the final skeleton feature map.
[0034] The value range of the pixel values of all pixel points in the body shape feature map is 0 - 1.
[0035] The mathematical expression of the pixel point deviation is:
[0036]
[0037] Where: L is the pixel deviation, w_Skel is the skeleton deviation coefficient, which is a constant, w_Physo is the body shape deviation coefficient, which is a constant, and L_Skel i is the difference between the pixel value of pixel i in the skeleton feature map and the pixel value of pixel i in the unified mask, is the skeleton weight of pixel i, and L_Physo i is the difference between the pixel value of pixel i in the body shape feature map and the pixel value of pixel i in the unified mask, is the body shape weight of pixel i, i is the serial number of the pixel, and M is the number of pixels.
[0038] An optimization device for a 3D pose detector based on a unified mask, comprising a memory, a processor, and a program stored in the memory. When the processor executes the program, the above method is implemented.
[0039] A storage medium, on which a program is stored. When the program is executed, the above method is implemented.
[0040] Compared with the prior art, the present invention has the following beneficial effects:
[0041] 1. Based on the obtained skeleton feature map and body shape feature map, taking the difference with the unified mask, and then cooperating with the weights of each pixel point, it is possible to ensure that the optimized 3D pose detector has relatively accurate 3D pose estimation on the basis of realizing the need for manual annotation of data.
[0042] 2. By means of projection, a two-dimensional skeleton feature map is obtained, so that a relatively mature U-Net network can be used to process and obtain the body shape feature map.
[0043] 3. The pixel values of the skeleton feature map and the body shape feature map are both normalized, which is convenient for setting the deviation coefficient and improving the performance of computer processing.
[0044] 4. A unique pixel deviation is designed, thereby improving the accuracy of 3D pose estimation. Brief Description of the Drawings
[0045] Figure 1 is a schematic diagram of the main step flow of the method of the present invention;
[0046] Figure 2 is a schematic diagram of the 3D pose estimation result of the 3D pose detector optimized according to the solution in the embodiment of the present invention. Detailed Embodiments
[0047] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and the detailed implementation manners and specific operation processes are given, but the protection scope of the present invention is not limited to the following embodiments.
[0048] An optimization method for a 3D pose detector based on a unified mask, as Figure 1 shown, includes:
[0049] Step S1: Obtain the human key point coordinates generated by the 3D pose detector based on the three-dimensional feature map;
[0050] The input of the 3D pose detector is the three-dimensional feature map, and the output is the human key point coordinates. As shown in Table 1, in this embodiment, a total of 18 human key point coordinates are set.
[0051] Table 1
[0052]
[0053]
[0054] Of course, in other embodiments, other different key point schemes can also be adopted.
[0055] Step S2: Based on the obtained human key point coordinates, register them with the pre-configured human skeleton to obtain the axis segments of all bones. Among them, the two ends of the axis segment of the bone respectively correspond to two human key points. Generally, in the human skeleton, the axis segments of two connected bones share a human key point.
[0056] Different human skeletons correspond to different key point schemes. In some embodiments, all adjacent bones are set to have a connection relationship, while in this embodiment, only some effectively connected bones are set to have a connection relationship. Based on the connection relationship configured in this embodiment, the endpoint representations of the axis segments of each bone are shown in Table 2
[0057] Table 2
[0058]
[0059]
[0060] In Table 2, taking Bone 1 and Bone 2 as examples, Bone 1 and Bone 2 have a connection relationship. The axis segments of the two share the human key point of the right hip, that is, Human Key Point 2. Similarly, Bone 2 and Bone 3 also have a connection relationship, and the axis segments of the two share the human key point of the right knee, that is, Human Key Point 3. The same applies to the others. By adopting the design scheme of the human skeleton diagram in this embodiment, the human skeleton features can be expressed in a concise manner, so as to efficiently represent the human body structure information in the skeleton feature diagram.
[0061] Of course, in other embodiments, other existing designs can also be adopted, such as the solutions provided in the following literature: He, Xingzhe, Bastian Wandt, and Helge Rhodin. "Autolink: Self-supervised learning of human skeletons and object outlines by linking keypoints." Advances in Neural Information Processing Systems 35 (2022): 36123 - 36141.
[0062] Step S3: Calculate the distances from all points in the three-dimensional feature map to each bone, and generate the bone feature maps of all bones based on the obtained distances. Specifically, the distance from a point to a bone is the distance from the point to the straight line where the axis segment of the bone is located;
[0063] In some embodiments, a three-dimensional calculation method can be adopted. That is, for the bone feature map of a single bone in Step S3, its generation process includes:
[0064] Calculate the Euclidean distances from all points in the three-dimensional feature map to the bone;
[0065] Generate the pixel values of the corresponding pixel points based on the Euclidean distances from the points to the bone;
[0066] Obtain the bone feature map of the bone based on the pixel values of the points.
[0067] However, in this embodiment, both the skeleton feature map and the bone feature map are two-dimensional images. For the bone feature map of a single bone in Step S3, its generation process includes:
[0068] Project the three-dimensional feature map onto a two-dimensional plane to obtain a two-dimensional feature map, and obtain the projections of the axis segments of each bone in the two-dimensional plane as projection line segments;
[0069] Calculate the Euclidean distances from all points in the two-dimensional feature map to the bone. Specifically, the distance from a point to a bone is the distance from the point to the straight line where the projection line segment of the bone is located;
[0070] Generate the pixel values of the corresponding pixel points based on the Euclidean distance from each point to the skeleton. In this embodiment, the pixel value is specifically exp(-d^2 / sigma^2), where d is the distance from the point to the skeleton, and sigma is a hyperparameter controlling the skeleton width, which is selected as 3e-3 in this embodiment.
[0071] Obtain the skeleton feature map of the skeleton based on the pixel values of each pixel point.
[0072] In this way, a two-dimensional skeleton feature map can be obtained, and then a relatively mature U-Net network can be used to process it to obtain the body shape feature map.
[0073] Step S4: Synthesize the skeleton feature maps of all skeletons into a skeleton feature map. In this embodiment, it specifically includes:
[0074] Register the skeleton feature maps of all skeletons. All the skeleton feature maps have the same size, so the pixel points at the same position have a corresponding relationship.
[0075] Sum the pixel values of the same pixel point in all skeleton feature maps as the pixel value of the corresponding pixel point in the preliminary skeleton feature map.
[0076] Normalize the pixel values of all pixel points in the preliminary skeleton feature map to obtain the final skeleton feature map.
[0077] Step S5: Generate the body shape feature map based on the skeleton feature map. In this embodiment, the body shape feature map is obtained by processing the skeleton feature map through a U-Net network, where the skeleton feature map is used as the input of the U-Net network.
[0078] In this embodiment, since the skeleton feature map is normalized, similarly, the value range of the pixel values of all pixel points in the body shape feature map is 0-1, and it is also normalized, that is, it is necessary to normalize the body shape feature map obtained by processing the U-Net network.
[0079] In addition, in other embodiments, other solutions can also be adopted. For example, for a three-dimensional skeleton feature map, the U-Net network can be improved and adjusted, or other machine learning networks can be developed to adapt to the conversion of three-dimensional images. However, this solution obviously requires relatively large computing power and training data, which is not conducive to promotion.
[0080] Step S6: Obtain the centroid of the pre-stored unified mask. Take the distance from each pixel point in the foreground area of the skeleton feature map to the centroid of the mask as the skeleton weight of this pixel point, and take the distance from each pixel point in the foreground area of the body shape feature map to the centroid of the mask as the body shape weight of this pixel point. Take the shortest distance from each pixel point in the background area of the skeleton feature map to the foreground area as the skeleton weight of this pixel point, and take the shortest distance from each pixel point in the background area of the body shape feature map to the foreground area as the body shape weight of this pixel point; the unified mask among them is a human body mask, and the size of the mask is the same as that of the skeleton feature map and the body shape feature map. The value of each pixel point is 0 or 1. This mask can be obtained by some existing means, so it will not be elaborated here;
[0081] Specifically, due to the increase in motion changes, key points farther from the root of the kinematic tree will be more difficult to be predicted by the detector. At the same time, for smooth optimization, detection points that wrongly fall into the background also need to be given different penalty weights. For this reason, in this application, weights based on geodesic distance are designed to enhance the representation of the skeleton feature map and the body shape feature map. Therefore, for the foreground area, we take the centroid of the mask as the zero point of the geodesic distance and calculate the geodesic distance within this area as the weight; for the background area, all mask parts are set as the zero point of the geodesic distance, and the geodesic distance within this area is also calculated as the weight.
[0082] Step S7: Based on the skeleton feature map and the skeleton weights of all its pixel points, and the body shape feature map and the body shape weights of all its pixel points, combined with the pixel values of each pixel point in the unified mask, calculate the pixel point deviation as the loss function value. Specifically, in this embodiment, the mathematical expression of the pixel point deviation is:
[0083]
[0084] where: L is the pixel point deviation, w_Skel is the skeleton deviation coefficient, taking a constant value, w_Physo is the body shape deviation coefficient, taking a constant value, L_Skel i is the difference between the pixel value of pixel point i in the skeleton feature map and the pixel value of pixel point i in the unified mask, is the skeleton weight of pixel point i, L_Physo i is the difference between the pixel value of pixel point i in the body shape feature map and the pixel value of pixel point i in the unified mask, is the body shape weight of pixel point i, i is the serial number of the pixel point, and M is the number of pixel points.
[0085] In this way, in a rough-to-fine manner, the human body shape is efficiently modeled using the skeleton feature map and the body shape feature map, thereby obtaining accurate supervision signals for the detector and providing a reasonable optimization path. In addition, by assigning different weights to the feature map space, the optimization plane can be effectively smoothed, further improving the optimization process of the detector and providing the possibility of detecting difficult samples.
[0086] Of course, in other embodiments, other settings of loss functions can also be adopted, but the convergence speed and optimization effect may not be as good as those of this embodiment.
[0087] Step S8: Optimize the 3D pose detector based on the obtained loss function value. Since this process belongs to the prior art, for example, in this embodiment, the above method can be applied to the deep learning open-source tool PyTorch to achieve the optimization of the 3D pose detector. The neural network training process is implemented using cascaded optimization. In the first stage, we only use the skeleton feature map to provide supervision for the 3D pose detector and optimize the human torso part except for the arms. When the optimization process approaches convergence, we introduce the body shape feature map to provide supervision and perform joint optimization in the second stage. In the final stage, we will optimize all human key points and continue training until the detection result reaches the optimal. During the testing process, we only need to use the trained 3D pose detector for inference, and the subsequent feature module will no longer be used.
[0088] As shown in Tables 3 and 4, the scheme of this embodiment achieves the best performance among unsupervised methods on the widely used Human3.6M and MPI-INF-3DHP datasets.
[0089] Table 3
[0090]
[0091]
[0092] In Table 3, MPJPE represents the mean point coordinate error, with the unit of millimeters. The lower the value, the better the result.
[0093] The prior art 1 among them adopts the scheme in the following literature: Jose Sosa and David Hogg. Self-supervised 3d human pose estimation from a single image. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 4787–4796, 2023.
[0094] The prior art 2 adopts the solution in the following literature: Jogendra Nath Kundu, Siddharth Seth, MV Rahul, Mugalodi Rakesh, Venkatesh Babu Radhakrishnan, and Anirban Chakraborty. Kinematic-structure-preserved representation for unsupervised 3d human pose estimation. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 11312–11319, 2020.
[0095] Table 4
[0096] Prior art 1 PCK AUC MPJPE Prior art 2 69.6 32.8 \ This embodiment 71.3 42.7 79.3
[0097] In Table 4, MPJPE represents the mean point coordinate error, with the unit of centimeter. The lower the value, the better the result. PCK represents the percentage of correct key points, and AUC represents the area under the curve metric
[0098] The visualization results are as Figure 2 shown. It can be seen that the detected 2D and 3D key point positions are accurate and highly consistent. By comparing the subgraphs horizontally and vertically, it can be found that we can solve the left-right inversion problem existing in most mask-based unsupervised methods.
[0099] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
Claims
1. An optimization method for 3D pose detectors based on a unified mask, characterized in that, Including: Step S1: Obtain the human key point coordinates generated by the 3D pose detector based on the three-dimensional feature map; Step S2: Based on the obtained human key point coordinates, register them with the pre-configured human skeleton to obtain the axis segments of all bones, where the two ends of the axis segment of the bone respectively correspond to two human key points; Step S3: Calculate the distances from all points in the three-dimensional feature map to each bone, and generate the bone feature maps of all bones based on the obtained distances, where the distance from a point to a bone is specifically the distance from the point to the straight line where the axis segment of the bone is located; Step S4: Synthesize the bone feature maps of all bones into a skeleton feature map; Step S5: Generate a body shape feature map based on the skeleton feature map; Step S6: Obtain the centroid of the mask of the pre-stored unified mask, obtain the distance from each pixel point in the foreground area of the skeleton feature map to the centroid of the mask as the skeleton weight of the pixel point, and obtain the distance from each pixel point in the foreground area of the body shape feature map to the centroid of the mask as the body shape weight of the pixel point, Obtain the shortest distance from each pixel point in the background area of the skeleton feature map to the foreground area as the skeleton weight of the pixel point, and obtain the shortest distance from each pixel point in the background area of the body shape feature map to the foreground area as the body shape weight of the pixel point; Step S7: Based on the skeleton feature map and the skeleton weights of all its pixel points, and the body shape feature map and the body shape weights of all its pixel points, combined with the pixel values of each pixel point in the unified mask, calculate the pixel point deviation as the loss function value; Step S8: Optimize the 3D pose detector based on the obtained loss function value; The mathematical expression of the pixel point deviation is: Wherein: L is the pixel deviation, is the skeleton deviation coefficient, taking a constant value, is the body shape deviation coefficient, taking a constant value, is the pixel in the skeleton feature map i and the pixel in the unified mask i the difference in pixel values, is the skeleton weight of the pixel i is the pixel in the body shape feature map i and the pixel in the unified mask i the difference in pixel values, is the body shape weight of the pixel i i is the serial number of the pixel, M is the number of pixels. 2. The optimization method of a 3D pose detector based on a unified mask according to claim 1, wherein, For the generation process of the bone feature map of a single bone in the step S3, it includes: Calculate the Euclidean distances from all points in the three-dimensional feature map to the bone; Generate the pixel values of the corresponding pixel points based on the Euclidean distances from the points to the bone; Obtain the bone feature map of the bone based on the pixel values of the points.
3. An optimization method for a 3D pose detector based on a unified mask according to claim 1, wherein, Both the skeleton feature map and the bone feature map are two-dimensional images. For the generation process of the bone feature map of a single bone in the step S3, it includes: Project the three-dimensional feature map onto a two-dimensional plane to obtain a two-dimensional feature map, and obtain the projection of the axis segment of each bone in the two-dimensional plane as the projection segment; Calculate the Euclidean distances from all points in the two-dimensional feature map to the bone, where the distance from a point to the bone is specifically the distance from the point to the straight line where the projection segment of the bone is located; Generate the pixel values of the corresponding pixel points based on the Euclidean distances from the points to the bone; Obtain the bone feature map of the bone based on the pixel values of the pixel points.
4. The optimization method of a 3D pose detector based on a unified mask according to claim 3, wherein, In the step S5, the body shape feature map is obtained by processing through a U-Net network, where the skeleton feature map is used as the input of the U-Net network.
5. A method for optimizing a 3D pose detector based on a unified mask, according to any one of claims 1-4, characterized in that In the human skeleton, the axis segments of two connected bones share a human key point.
6. A method for optimizing a 3D pose detector based on a unified mask according to any one of claims 1-4, characterized in that, The step S4 specifically includes: Register the bone feature maps of all bones; Sum the pixel values of the same pixel point in all bone feature maps as the pixel value of the corresponding pixel point in the preliminary skeleton feature map; Normalize the pixel values of all pixel points in the preliminary skeleton feature map to obtain the final skeleton feature map.
7. A method for optimizing a 3D pose detector based on a unified mask according to any one of claims 1-4, characterized in that The value range of the pixel values of all pixel points in the body shape feature map is 0-1.
8. An optimization device for a 3D pose detector based on a unified mask, comprising a memory, a processor, and a program stored in the memory, characterized in that, When the processor executes the program, the method described in any one of claims 1-7 is implemented.
9. A storage medium, on which a program is stored, characterized in that, When the program is executed, the method described in any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Mark-point-free visual motion capturing method based on three purposes
CN112819849A
Three-dimensional human body posture estimation method and computer readable storage medium
CN112836618A