A training data generation system, method and non-volatile computer-readable storage medium thereof for human pose recognition

TWI931697BActive Publication Date: 2026-07-11CHUNGHWA TELECOM CO LTD
0 Cites 0 Cited by

Patent Information

Application Number
TW112147252
Authority / Receiving Office
TW · TW
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-12-05
Publication Date
2026-07-11
Estimated Expiration
2043-12-04
Patent Text Reader

Abstract

This invention provides a system, method, and non-volatile computer-readable storage medium for generating training data for human posture recognition. The system includes a mannequin simulation module and a feature rendering module. The mannequin simulation module simulates a person's posture and physique using a set of simulated human body parameters. The feature rendering module renders a nude 3D mannequin image based on the simulated human body parameters and an appearance feature model weight, thereby obtaining a specific motion image dataset. Therefore, this invention can increase the richness of training data and also allows for the customized creation of entirely new image datasets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a training data generation technology for machine learning, and more particularly to a training data generation system, method, and non-volatile computer-readable storage medium for human posture recognition. Prior Technology

[0002] In recent years, with the advancement of technology, machine learning has been applied to a wider range of fields and levels. Training data is a crucial component of machine learning. However, most human image datasets currently available on the market are collected by filming real people. This not only consumes high manpower, time, and equipment costs, but also makes it difficult to collect data on actions that are not suitable for filming, thus hindering related research.

[0003] Furthermore, previous technologies required users to have relevant shooting equipment and hire people to shoot if they wanted to obtain a set of image data of specific actions. If the actions were not suitable for humans to perform, such as fighting, the photographer was easily injured. Acting methods may also fail to present realistic actions, resulting in the inability to provide the required features, which in turn affects the machine learning recognition effect.

[0004] Therefore, how to quickly and conveniently generate a set of high-quality and sufficient video data of specific actions has become an urgent issue for the industry. Summary of the Invention

[0005] To address the aforementioned problems, this invention provides a training data generation system for human posture recognition, comprising: a puppet simulation module that simulates a complex set of simulated human body parameters based on at least one human figure image; and a feature rendering module with a rendering model that is communicatively connected to the puppet simulation module to receive the complex set of simulated human body parameters from the puppet simulation module. The rendering model takes the complex set of simulated human body parameters and an appearance feature model weight as input to render a plurality of nude 3D puppet images with specific actions constructed from the complex set of simulated human body parameters, thereby obtaining a plurality of 3D puppet images with appearance features and specific actions, thus forming a specific action image dataset.

[0006] The present invention further provides a method for generating training data for human posture recognition, comprising: simulating a complex array of simulated human body parameters by a puppet simulation module based on at least one human body image; receiving the complex array of simulated human body parameters from the puppet simulation module by a feature rendering module with a rendering model; and using the rendering model as input the complex array of simulated human body parameters and an appearance feature model weight to render a complex array of nude three-dimensional puppet images with specific actions constructed by the complex array of simulated human body parameters, so as to obtain a complex array of three-dimensional puppet images with appearance features and specific actions, thereby forming a specific action image data set.

[0007] In the foregoing embodiments, the complex array of simulated human body parameters includes at least one set of first simulated human body parameters and a complex array of second simulated human body parameters. The at least one set of first simulated human body parameters is generated by the mascot simulation module based on the at least one human figure image, while the complex array of second simulated human body parameters is generated by the mascot simulation module by a user setting the parameters based on the at least one first simulated human body parameter.

[0008] In the aforementioned embodiments, the dummy simulation module obtains the three-dimensional coordinates of at least one human figure from the at least one human figure image, and converts the three-dimensional coordinates of the at least one human figure into the at least one first simulated human body parameter through posture estimation and body shape estimation.

[0009] The aforementioned embodiment further includes a first database that stores an external feature image dataset and is communicatively connected to the feature rendering module. The feature rendering module uses the external feature image dataset to train the rendering model to obtain the appearance feature model weights.

[0010] The foregoing embodiments further include a second database, which is communicatively connected to the feature rendering module to receive and store specific motion image data sets from the feature rendering module.

[0011] In the aforementioned embodiment, the human simulation module is a Skinned Multi-Person Linear Model (SMPL), and the simulated human body parameters are SMPL parameters.

[0012] As described above, the training data generation system, method, and non-volatile computer-readable storage medium of the present invention for human posture recognition mainly simulate a person's posture and body shape through a set of simulated human body parameters (such as SMPL parameters) to generate nude 3D dummy figures with various postures and body shapes. Then, the nude 3D dummy figures are rendered so that they have appearance features such as clothing, skin color, and facial expressions. Therefore, the present invention can not only quickly and easily increase the richness of training data, but also customize and create entirely new image datasets to make up for the shortcomings of existing human motion datasets that are based on live-action filming. Simple Explanation of the Diagram

[0013] Figure 1 is a schematic diagram of the architecture of the training data generation system for human posture recognition according to the present invention.

[0014] Figure 2 is a flowchart illustrating the method for generating training data for human posture recognition according to the present invention.

[0015] Figures 3A and 3B are schematic diagrams after background removal processing.

[0016] Figure 4 is a schematic diagram of a nude three-dimensional doll.

[0017] Figure 5 is a schematic diagram of a three-dimensional doll image depicting a person engaged in labor.

[0018] Figures 6A and 6B are schematic diagrams of images of people wearing reflective vests and helmets.

[0019] Figure 7 is a schematic diagram of a three-dimensional doll image with appearance features and labor behavior. Implementation

[0020] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification.

[0021] It should be understood that the structures, proportions, sizes, etc., illustrated in the accompanying drawings of this specification are only for the purpose of assisting those skilled in the art in understanding and reading the content disclosed herein, and are not intended to limit the conditions under which the present invention can be implemented. Therefore, they have no substantial technical significance. Any modifications to the structure, changes in the proportional relationships, or adjustments to the size, without affecting the effects and objectives that the present invention can produce, should still fall within the scope of the technical content disclosed herein. Furthermore, the terms such as "a," "first," "second," "above," and "below" used in this specification are merely for clarity of description and are not intended to limit the scope of the present invention. Changes or adjustments to their relative relationships, without substantially altering the technical content, should be considered within the scope of the present invention.

[0022] Figure 1 is a schematic diagram of the architecture of the training data generation system 1 for human posture recognition according to the present invention. As shown in Figure 1, the training data generation system 1 for human posture recognition includes: a puppet simulation module 11, a feature rendering module 12, a first database 21 and a second database 22.

[0023] Specifically, the training data generation system 1 for human posture recognition can be built on the same (or different) servers (such as general-purpose servers, file-based servers, storage unit servers, etc.) and computers, or other electronic devices with appropriate computing mechanisms. Each module in the training data generation system 1 for human posture recognition (such as the dummy simulation module 11 and the feature rendering module 12) can be software, hardware, or firmware. If it is hardware, it can be a processing unit, processor, computer, or server with data processing and computing capabilities. If it is software or firmware, it can include instructions executable by the processing unit, processor, computer, or server, and can be installed on the same hardware device or distributed across multiple different hardware devices. Furthermore, the first database 21 and the second database 22 can be storage devices with hard drives (such as solid-state drives, mechanical hard drives) or magnetic tapes.

[0024] In one embodiment, the puppet simulation module (or: puppet posture and body shape simulation module) 11 can be a Skinned Multi-Person Linear Model (SMPL) to simulate a complex set of simulated human body parameters based on a complex set of human images. The complex set of simulated human body parameters describes the posture (i.e., the action posture of the human body or puppet) and body shape (i.e., the height and weight of the human body or puppet), which serves as the output of the puppet simulation module 11. Subsequently, a complex set of simulated human body parameters can be used to construct a complex set of nude three-dimensional puppet images with specific actions.

[0025] In this embodiment, the plurality of humanoid images can be obtained by taking photos or videos of real people, resulting in a plurality of images with human postures and body shapes. The human postures in these images are determined by the user, and since the images with human postures are obtained by taking photos of real people, they naturally possess the body shapes of real people. Alternatively, the humanoid images can also be generated by a doll posture and body shape setting program, allowing a user to set the desired doll posture and body shape, thus producing a plurality of images with doll postures and body shapes. Furthermore, the doll simulation module 11 uses an object detection algorithm to extract a portion of the humanoid image from the humanoid image, representing the area or location of the human or doll. Then, the doll simulation module 11 uses a feature point detection algorithm to obtain the humanoid coordinates of the human or doll in this portion of the humanoid image, such as: joint position coordinates, limb position coordinates, etc., relative to the human or doll, and is not limited thereto.

[0026] In one embodiment, the puppet's pose and posture setting program includes, but is not limited to, 3D modeling software such as Blender, Maya, and 3ds Max. It should be noted that these software programs themselves do not have the function of simulating SMPL, but related SMPL simulations can be performed by using other programs in conjunction with these software programs.

[0027] In one embodiment, the object detection algorithm includes, but is not limited to, YOLO (You Only Look Once), SSD (Single Shot MultiBox Detector), Faster RCNN, etc.; and the feature point detection algorithm is used to obtain the coordinates of a person's joints, including, but not limited to, OpenPose, PoseNet, etc.

[0028] In one embodiment, the humanoid coordinates are two-dimensional (2D) coordinates. The mascot simulation module 11 uses the camera's intrinsic and extrinsic parameters to convert the two-dimensional coordinates of the humanoid coordinates into three-dimensional (3D) coordinates. Based on the three-dimensional coordinates of the humanoid coordinates, the mascot simulation module 11 uses pose estimation and body posture estimation to convert the three-dimensional coordinates of the humanoid coordinates into the simulated human body parameters. The pose estimation and body posture estimation take into account the joint angles and positions of the human body or mascot, as well as global translation and scaling parameters, to convert the three-dimensional coordinates of the humanoid coordinates into the first simulated human body parameters.

[0029] Furthermore, the dummy simulation module 11 allows the user to set a plurality of second simulated human body parameters based on the first simulated human body parameters. The first simulated human body parameters and the plurality of second simulated human body parameters together form the plurality of simulated human body parameters. The internal and external camera parameters of the plurality of second simulated human body parameters are the same as those of the first simulated human body parameters. Therefore, by allowing the user to set the plurality of second simulated human body parameters based on the first simulated human body parameters through the dummy simulation module 11, various dummy postures and body shapes can be quickly simulated without expending significant manpower and resources to film the movements of real people.

[0030] In one embodiment, the camera's intrinsic and extrinsic parameters are obtained for photographing real people, or are set by the dummy's posture and body setting program.

[0031] In one embodiment, the intrinsic parameters refer to the camera's focal length, principal point position, etc., which are usually fixed after the camera leaves the factory, but different cameras will have different parameters; while the extrinsic parameters refer to the camera's current world coordinate position and the angle of rotation relative to the world benchmark camera, so different shooting conditions (as long as the camera moves or rotates) will produce different extrinsic parameters.

[0032] In one embodiment, the complex array of simulated human body parameters can be Skinned Multi-Person Linear Model (SMPL) parameters, and includes two types: β parameters and θ parameters. The β parameters include 10 parameters for adjusting the body posture of the nude 3D dummy, and the θ parameters include 72 parameters for adjusting the pose of the nude 3D dummy.

[0033] In one embodiment, when the dummy simulation module 11 performs posture estimation and body shape estimation based on the three-dimensional coordinates of the human figure coordinates, the dummy simulation module 11 initializes a preset simulated human body parameter, and then optimizes the preset simulated human body parameter, thereby converting the optimized simulated human body parameter into a candidate three-dimensional coordinate, and evaluating the candidate three-dimensional coordinate with the three-dimensional coordinate of the human figure coordinate to obtain a loss function (such as L-BFGS, conjugate gradient method). In the optimization process of the preset simulated human body parameter, the smaller the value obtained by the loss function, the closer it is to the desired simulated human body parameter, thereby approximating the desired simulated human body parameter by the loss function.

[0034] In detail, this optimization process is an iterative process, with posture and body position being optimized alternately to convert the optimized simulated human body parameters into candidate 3D coordinates one by one. These are then evaluated against the 3D coordinates of the human body, i.e., the loss function is evaluated. Since posture and body position affect each other, after optimizing the posture once, the body position is then optimized until the value of the loss function is less than a preset threshold value, at which point the desired simulated human body parameters can be obtained.

[0035] In one embodiment, the feature rendering module (or: puppet external feature rendering module) 12 has a rendering model 12a and is communicatively connected to the puppet simulation module 11, the first database 21 and the second database 22 to receive the complex array of simulated human body parameters from the puppet simulation module 11. The rendering model 12a takes an appearance feature model weight and the complex array of simulated human body parameters as input to render a plurality of nude 3D puppet images with specific actions constructed from the complex array of simulated human body parameters, thereby obtaining a plurality of 3D puppet images with appearance features and specific actions. The plurality of 3D puppet images with appearance features and specific actions are formed into a specific action image dataset, which is used as a specific action image dataset for training other models.

[0036] In this embodiment, the rendering model 12a is a model based on a neural network (such as a convolutional neural network, CNN). The feature rendering module 12 uses the external feature image dataset stored in the first database 21, the complex array of simulated human body parameters for training, and the intrinsic and extrinsic parameters of the camera to perform deep learning training on the rendering model 12a to obtain the appearance feature model weights. This enables the trained rendering model 12a to quickly and accurately render the appearance features of a nude 3D doll.

[0037] In one embodiment, the rendering model 12a is trained by using an external feature image dataset (containing external feature images at various angles and their corresponding de-masked images), a complex array of simulated human body parameters, and the intrinsic and extrinsic parameters of the camera to train the rendering model 12a and obtain the appearance feature model weights. Specific external features are then rendered onto a doll in a specified pose and body shape using these appearance feature model weights, enabling the rendering model 12a to learn how to represent the doll's external features from various angles. When providing images to the model, the images in the external feature image dataset undergo preprocessing, such as background removal, to prevent the rendering model 12a from learning background-related features, which could affect the rendered image quality.

[0038] Specifically, since the nude 3D doll image lacks visual features, it cannot be directly used as training data for other models in deep learning. Therefore, the rendering model 12a needs to be trained to render facial and clothing features onto the nude 3D doll image. More specifically, a camera is set up approximately 3 meters away from a subject with a specific appearance to capture multiple other human-shaped images of the subject. Alternatively, at least 200 other human-shaped images of the subject can be captured directly, or a video of the subject can be recorded to extract at least 200 other human-shaped images of the subject from the video. These at least 200 other human-shaped images include images of the subject from various angles, and the subject having a specific appearance means that the subject can wear specific clothing as needed.

[0039] Taking the video of the subject as an example, the subject will spin around once, raising their hands and feet during the spin to capture areas that are easily obscured. After the video is recorded, at least 200 other human figures can be extracted from it. In one embodiment, these other human figures may be the same as or different from the human figures in the above embodiment.

[0040] Furthermore, after obtaining at least 200 other human figures, the feature rendering module 12 or other device with a processor first performs background removal processing on the at least 200 other human figures to obtain external feature images and their corresponding background-removed mask images from various angles, thereby obtaining the external feature image dataset. The background removal method can be performed using an artificial intelligence (AI) model or manually, and is not limited to any particular method. In one embodiment, the background-removed mask image can filter out all non-human parts in the other human figures to obtain the external feature image. Therefore, the background removal process does not directly generate an image with the background removed, but rather generates a mask that can filter out the background.

[0041] Subsequently, the feature rendering module 12 or other device with a processor will provide the rendering model 12a with external feature images (as shown in Figure 3A) at various angles and their corresponding de-masked images (as shown in Figure 3B), as well as the training array of simulated human body parameters and the camera's intrinsic and extrinsic parameters, as inputs to obtain the appearance feature model weights.

[0042] In one embodiment, the second database 22 stores the specific action image dataset for use in subsequent deep learning of other types of neural network models, enabling its application in the training of different action recognition or detection systems. For example, it can be applied to fight detection systems, fall detection systems, and behavior recognition systems.

[0043] Figure 2 is a schematic flowchart of the method for generating training data for human posture recognition according to the present invention, and is also illustrated with reference to Figure 1. Furthermore, the similarities between this embodiment and the above embodiments will not be repeated, and this method includes the following steps S21 to S23:

[0044] In step S21, a single doll simulation module 11 simulates a complex set of simulated human body parameters based on a complex set of human body images.

[0045] In step S22, a feature rendering module 12 receives a complex array of simulated human body parameters from the dummy simulation module 11, and the rendering model 12a in the feature rendering module 12 takes the complex array of simulated human body parameters and an appearance feature model weight as input to render a complex number of nude 3D dummy images with specific actions constructed from the complex array of simulated human body parameters, so as to obtain a complex number of 3D dummy images with appearance features and specific actions.

[0046] In step S23, the feature rendering module 12 stores the set of specific action images formed by the plurality of three-dimensional puppet images with appearance features and specific actions in a second database 22 as a set of specific action images, which can be used by other types of neural network models for deep learning.

[0047] Furthermore, this invention discloses a non-volatile computer-readable storage medium applied in a computing device or computer having a processor (e.g., CPU, GPU, etc.) and / or memory, storing instructions, and capable of executing the non-volatile computer-readable storage medium through the processor and / or memory using the computing device or computer to perform the aforementioned methods and steps when executing the non-volatile computer-readable storage medium. In one embodiment, this invention further discloses a non-transitory computer-readable storage medium for performing the aforementioned methods and steps.

[0048] The following example is a specific embodiment of the training data generation system 1 for human posture recognition of the present invention, which is described with reference to Figures 1 and 2. The similarities between this embodiment and the above embodiments will not be repeated.

[0049] First, when a user designs a system for recognizing labor behavior, they must first train an image recognition model through deep learning to recognize the image features of labor. Therefore, a dataset of images containing labor behavior is required, and the quantity and diversity of these images affect the recognition accuracy of the image recognition model. However, obtaining an image dataset for a specific purpose is quite difficult, or the obtained image data may not fully match the user's desired scenario. Therefore, the training data generation system 1 for human posture recognition of this invention can generate an image dataset of specific actions (i.e., labor behavior) required by the user.

[0050] In this embodiment, a puppet simulation module 11 simulates a complex set of simulated human body parameters (such as SMPL parameters) based on a small number of human images of people working. These simulated human body parameters are used to control a naked three-dimensional puppet with 6890 vertices and 24 skeletal points (as shown in Figure 4), so as to obtain a complex set of three-dimensional puppet images with working behavior (as shown in Figure 5). The complex set of three-dimensional puppet images with working behavior lacks external features, such as clothing, facial expressions, and skin color.

[0051] Therefore, a feature rendering module 12 first trains its rendering model 12a using an external feature image dataset provided by a first database 21. After the rendering model 12a completes training, it obtains a set of appearance feature model weights. As shown in Figures 6A and 6B, the external feature image dataset includes images of a person wearing a reflective vest and a safety helmet. Next, the rendering model 12a uses the appearance feature model weights and the complex array of simulated human body parameters as input to generate a complex set of three-dimensional dummy images with appearance features (images of a person wearing a reflective vest and a safety helmet) and work behavior (as shown in Figure 7). This set serves as a specific action image dataset required by the user, which the user can use for subsequent training to train the work behavior recognition system and improve the system's recognition accuracy.

[0052] In summary, the training data generation system, method, and non-volatile computer-readable storage medium of this invention for human posture recognition simulate a person's posture and physique using a set of simulated human parameters (such as SMPL parameters) to generate nude 3D avatars with various postures and physiques. These nude 3D avatars are then rendered to give them appearance features such as clothing, skin color, and facial expressions. Therefore, this invention can not only quickly and easily increase the richness of training data but also customize and create entirely new image datasets to overcome the shortcomings of existing human motion datasets based on live-action filming, such as high filming costs and potential safety issues when filming certain special movements.

[0053] Compared to traditional training data collection methods, this invention effectively increases the diversity of training data while avoiding potential dangers to the human body during data collection. Furthermore, the training data generated by this invention improves the accuracy of human posture recognition and can be applied to various products, possessing high industrial application value and enabling the creation of more diverse and accurate detection service systems.

[0054] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify and alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention should be as set forth in the claims.

[0055]

[0056] 1: Training data generation system for human posture recognition

[0057] 11: Doll Simulation Module

[0058] 12: Feature Rendering Module

[0059] 12a: Rendering Model

[0060] 21: First Database

[0061] 22: Second Database

[0062] S21 to S23: Steps

Claims

1. A system for generating training data for human posture recognition, comprising: A single doll simulation module is a complex array of simulated human body parameters based on at least one human figure image; A feature rendering module with a rendering model is communicatively connected to the dummy simulation module. The feature rendering module receives a complex set of simulated human body parameters from the dummy simulation module. The feature rendering module uses a dataset of complex external feature images from different angles and corresponding complex back-masked images, along with the training dataset of complex simulated human body parameters and the camera's intrinsic and extrinsic parameters, to train the rendering model and obtain appearance feature model weights. The rendering model then uses the complex set of simulated human body parameters from the dummy simulation module, the dataset of complex external feature images from different angles and corresponding complex back-masked images, and the training dataset of complex simulated human body parameters and camera intrinsic and extrinsic parameters to train the rendering model and obtain appearance feature model weights. The weights of the appearance feature model, obtained by training with a complex array of simulated human body parameters, the camera's intrinsic and extrinsic parameters, are used as input. The rendering model of the feature rendering module, based on the complex array of simulated human body parameters, and the external feature image dataset containing a complex set of external feature images from different angles and corresponding complex de-masked images, and the weights of the appearance feature model obtained by training with the complex array of simulated human body parameters, renders a complex set of nude 3D puppet images with different specific actions constructed by the complex array of simulated human body parameters of the puppet simulation module, to obtain a complex set of 3D puppet images with different appearance features and different specific actions, thus forming a specific action image dataset.

2. The training data generation system for human posture recognition as described in claim 1, wherein, The complex array of simulated human body parameters includes at least one set of first simulated human body parameters and a complex array of second simulated human body parameters. The at least one set of first simulated human body parameters is generated by the mascot simulation module based on the at least one human figure image, and the complex array of second simulated human body parameters is generated by the mascot simulation module by a user setting the at least one set of first simulated human body parameters.

3. The training data generation system for human posture recognition as described in claim 2, wherein, The avatar simulation module obtains the three-dimensional coordinates of at least one human figure from the at least one human figure image, and converts the three-dimensional coordinates of the at least one human figure into the at least one set of first simulated human body parameters through posture estimation and body shape estimation.

4. The training data generation system for human posture recognition as described in claim 1 further includes a first database that stores the set of external feature images and is communicatively connected to the feature rendering module.

5. The training data generation system for human posture recognition as described in claim 1 further includes a second database, which is communicatively connected to the feature rendering module to receive and store a specific motion image data set from the feature rendering module.

6. The training data generation system for human posture recognition as described in claim 1, wherein, The human simulation module is a Skinned Multi-Person Linear Model (SMPL), and the simulated human body parameters are SMPL parameters.

7. A method for generating training data for human posture recognition, comprising: A single puppet simulation module simulates multiple sets of simulated human body parameters based on at least one human figure image; A feature rendering module with a rendering model receives complex arrays of simulated human body parameters from the dummy simulation module. The feature rendering module uses a dataset of complex extrinsic feature images from different angles and corresponding complex back-masked images, the complex arrays of simulated human body parameters used for training, and the camera's intrinsic and extrinsic parameters to train the rendering model to obtain appearance feature model weights. The rendering model of the feature rendering module uses the complex arrays of simulated human body parameters from the dummy simulation module, the dataset of complex extrinsic feature images from different angles and corresponding complex back-masked images, the complex arrays of simulated human body parameters used for training, and the camera's intrinsic and extrinsic parameters. The weights of the appearance feature model obtained by training the intrinsic and extrinsic parameters are used as input. The rendering model of the feature rendering module uses the complex array of simulated human body parameters, and the complex extrinsic feature image dataset containing complex extrinsic feature images from different angles and corresponding complex de-masking images, the complex array of simulated human body parameters used for training, and the weights of the appearance feature model obtained by training the camera's intrinsic and extrinsic parameters to render a complex number of nude 3D puppet images with different specific actions constructed by the complex array of simulated human body parameters of the puppet simulation module. This results in a complex number of 3D puppet images with different appearance features and different specific actions, forming a specific action image dataset.

8. The method for generating training data for human posture recognition as described in claim 7 further includes generating at least one set of first simulated human body parameters by the dummy simulation module based on the at least one human figure image, and providing a user with settings based on the at least one set of first simulated human body parameters by the dummy simulation module to generate a plurality of second simulated human body parameters, wherein, The complex array of simulated human body parameters includes at least one set of first simulated human body parameters and the complex array of second simulated human body parameters.

9. The method for generating training data for human posture recognition as described in claim 8 further includes obtaining the three-dimensional coordinates of at least one human figure from the at least one human figure image by the dummy simulation module, and converting the three-dimensional coordinates of the at least one human figure into the at least one set of first simulated human body parameters through posture estimation and body shape estimation.

10. The method for generating training data for human posture recognition as described in claim 7 further includes storing the set of external feature images in a first database.

11. The method for generating training data for human pose recognition as described in claim 7 further includes receiving a specific motion image dataset from the feature rendering module from a second database, and storing the specific motion image dataset.

12. The method for generating training data for human posture recognition as described in claim 7, wherein, The human simulation module is a Skinned Multi-Person Linear Model (SMPL), and the simulated human body parameters are SMPL parameters.

13. A non-volatile computer-readable storage medium, used in a computing device or computer, storing instructions for performing a method for generating training data for human posture recognition as described in any of claims 7 to 12.