Method and Device for Generating Training Data of Human Skeleton Joint Point Extraction Model
By 3D motion modeling and 2D rendering of the target human body, the training data of bone joint nodes is solved, and the automatic labeling of bone joint nodes is achieved and the stability of model performance is improved.
Patent Information
- Application Number
- CN202111644931.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-29
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-12-29
AI Technical Summary
In the prior art, it is time-consuming and labor-intensive to obtain training data from the human skeleton node extraction model, and the collected training data and labeling information are inaccurate and the performance is unstable.
By modeling the target human body 3D motion, 3D motion skin data is generated, the 3D coordinates of the bone joint node are determined using the predefined association relationship, and 2D pixel coordinates are generated through 2D rendering to realize automatic labeling of the bone joint node.
Automatic labeling of bone joint nodes is realized, reducing the cost of manual collection and labeling, improving data quality and stability, reducing system errors, and improving the prediction performance of the model.
Smart Images

Figure CN114359445B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a method and device for generating training data of a human body bone joint point extraction model. Background Art
[0002] The extraction of human body bone joint points is the core technology of artificial intelligence application algorithms such as action recognition, behavior analysis, human body vision measurement, and medical gait analysis. At present, the extraction of bone joint points based on deep learning has achieved good application results and has become the mainstream technical solution for bone joint point extraction. For example, a deep learning-based bone joint point extraction model can be used to extract bone joint points.
[0003] Among them, the deep learning-based bone joint point extraction model can be obtained through the following method: collecting a large amount (or a large scale) of human body image data, performing bone joint point annotation on the collected human body image data to obtain the annotation information of bone joint points, and using the human body image data (or called training data) and the annotation information to train model parameters to obtain a human body bone joint point extraction model.
[0004] As can be seen from the above, the human body bone joint point extraction model needs to be trained relying on a large amount of human body image data and bone joint point annotation information. Currently, artificial collection of human body image data and annotation information is usually adopted, which requires a large amount of human resources and is time-consuming and laborious. In addition, in addition to being time-consuming and laborious, the artificial collection of human body image data and the annotation of bone joint points also have the following problems: it is difficult to adjust the judgment and annotation methods of bone joint points in a timely manner according to requirements, the image quality of the natural image data collected manually is difficult to control, and the systematic errors of manual annotation are difficult to quantify and evaluate, resulting in inaccurate prediction results and unstable performance of the joint point extraction model obtained by training the bone joint extraction model using the image data and annotation information collected manually. Summary of the Invention
[0005] The purpose of the embodiments of this specification is to provide a method and device for generating training data of a human body bone joint point extraction model, so as to solve the problems that the existing acquisition of training data and annotation information is time-consuming and laborious, and the collected training data and annotation information are inaccurate and have unstable performance.
[0006] To solve the above technical problems, the embodiments of the present application are implemented in the following ways:
[0007] In a first aspect, the present application provides a method for generating training data for a human body bone joint point extraction model. The method includes: performing 3D human body motion modeling on a target human body to generate 3D motion skinned data of the target human body; wherein the 3D motion skinned data includes a 3D mesh set corresponding to each 3D motion frame under a coherent motion during the occurrence of the target human body motion; determining the 3D coordinates of each bone joint point in each 3D motion frame of the target human body according to the association relationship between the 3D coordinates of each predefined bone joint point and the corresponding subset of the 3D mesh set, and the 3D motion skinned data of the target human body; wherein the association relationship is predefined according to the task requirements; performing 2D rendering on each 3D motion frame according to the task requirements, generating 2D image data of the target human body according to the rendering result, and performing coordinate transformation on the 3D coordinates of the bone joint points to generate 2D pixel coordinates of the bone joint points.
[0008] In a second aspect, the present application provides a device for generating training data for a human body bone joint point extraction model. The device includes: a first generation module for performing 3D human body motion modeling on a target human body to generate 3D motion skinned data of the target human body; wherein the 3D motion skinned data includes a 3D mesh set corresponding to each 3D motion frame under a coherent motion during the occurrence of the target human body motion; a determination module for determining the 3D coordinates of each bone joint point in each 3D motion frame of the target human body according to the association relationship between the 3D coordinates of each predefined bone joint point and the corresponding subset of the 3D mesh set, and the 3D motion skinned data of the target human body; wherein the association relationship is predefined according to the task requirements; a second generation module for performing 2D rendering on each 3D motion frame according to the task requirements, generating 2D image data of the target human body according to the rendering result, and performing coordinate transformation on the 3D coordinates of the bone joint points to generate 2D pixel coordinates of the bone joint points.
[0009] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method for generating training data for a human body bone joint point extraction model as in the first aspect.
[0010] In a fourth aspect, the present application provides a readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for generating training data for a human body bone joint point extraction model as in the first aspect.
[0011] As can be seen from the technical solutions provided in the embodiments of this specification above, in this solution, a 3D mesh set corresponding to each 3D action frame during the generation of the target human body movement is generated, and the 3D coordinates of the skeletal joint points are obtained according to the association relationship between the predefined subset of the 3D mesh set and the 3D coordinates of each skeletal joint point, completing the automatic annotation of the 3D coordinates of the skeletal joint points. According to the task requirements, the 3D action frames are rendered into 2D images, and the calculation of the 3D coordinates of the skeletal joint points is converted into the pixel coordinates of the 2D image in the corresponding viewing angle, completing the automatic annotation of the 2D pixel coordinates of the skeletal joint points. In this way, not only the automatic annotation of the skeletal joint points is realized, but also this method can be changed in a timely manner according to the requirements, the system error is stable and easy to eliminate. At the same time, the automatically generated annotation data is used as supervision information to train the skeletal joint point extraction model, and the fluctuation range of its prediction results is smaller than that of the manually annotated data, the prediction performance is more stable, and the cost of the image data collection and annotation process is lower than that of the traditional method, the efficiency is higher, and the quality of the image data and the diversity of the samples are easier to control than the natural image collection. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0013] Figure 1 It is a schematic flowchart of the method for generating training data of the human skeletal joint point extraction model provided by this application;
[0014] Figure 2 It is a schematic flowchart of the method for generating 3D action skinning data of the target human body provided by this application;
[0015] Figure 3 It is a schematic diagram of capturing the movement of the target human body provided by this application;
[0016] Figure 4 It is a schematic diagram of constructing a 3D character based on Make human provided by this application;
[0017] Figure 5 It is a schematic diagram of establishing the scene where the target human body movement occurs provided by this application;
[0018] Figure 6 It is a schematic structural diagram of the device for generating training data of the human skeletal joint point extraction model provided by this application;
[0019] Figure 7Schematic structural diagram of the electronic device provided by this application. Detailed implementation manners
[0020] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of this specification.
[0021] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of this application. However, those skilled in the art should understand that this application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of this application.
[0022] Without departing from the scope or spirit of this application, various improvements and changes can be made to the specific implementation manners of this application specification, which are obvious to those skilled in the art. Other implementation manners obtained from the specification of this application are obvious to those skilled in the art. The specification and embodiments of this application are only exemplary.
[0023] Regarding the use of "including", "comprising", "having", "containing", etc. in this article, they are all open-ended terms, that is, they are meant to include but not be limited to. In addition, the "skeletal joint points" in this application can be understood as human skeletal joint points unless otherwise specified.
[0024] In the related art, bone joint points are extracted by a bone joint point extraction model based on deep learning. However, in the process of training a bone joint point extraction model based on the collected relevant human natural image data and the manually labeled bone joint point data, the following three stages are usually required: (1) According to the task requirements, design the acquisition scenario of human images. Such as the camera view, the positions of multiple cameras, the data acquisition site, the composition of the people to be collected (age, height, body type, clothing, gender), the actions to be collected, etc. If it is a multi-camera acquisition scenario, the internal and external parameters of the cameras also need to be calibrated. (2) Data arrangement. Such as data cleaning, resolution unification. If it is a multi-camera acquisition, image frame synchronization and alignment are required. (3) Manual data annotation. According to the task requirements, the annotators need to be trained to identify the joint points to be annotated in the images. Then, the annotators use image annotation tools and, based on subjective experience, judge and mark the pixel points representing the corresponding joint points in the images and annotate them to form an annotation file.
[0025] As can be seen from the above process, in the process of training a bone joint point extraction model, collecting the image data of bone joint points and generating the annotation information of bone joint points are time-consuming and laborious, and the acquisition cost of high-quality annotation data is relatively high. In addition, the following problems also exist:
[0026] Problem (1): It is difficult to adjust the method for judging and annotating bone joint points in a timely manner according to requirements.
[0027] In different applications, the required bone joint point annotation information may be inconsistent, resulting in the need to adjust the definition and judgment rules of bone joint points according to the application. Even for the same batch of images, according to different task requirements, repeated annotation may be required. For example, the bone joint points required by gait analysis and some action recognition models in somatosensory games are different. The former focuses on the bone joint points of the legs, and the latter may focus on the bone joint points of the hands. Therefore, the annotation data of the two are often not interchangeable.
[0028] Problem (2): It is difficult to control the quality of natural image data.
[0029] Due to the constraints of actual conditions, the human body image datasets collected by traditional methods often have problems with uncontrollable sample diversity and quality. Currently, the sources of human body image data are mainly divided into two categories. One is the images of human actions and postures in various scenarios collected from the Internet, and the other is the collection of human body image data under laboratory conditions. For the former, although there are a large number of images, due to the uncontrollable conditions under which the data is generated, there are a large number of low-quality data, such as people being blocked; uneven sample distribution (mainly adults and young people), insufficient image resolution (resulting in the inability to perform high-precision bone joint point annotation), etc. It is not suitable for some tasks that require precise positioning of bone joint points, and the images from the Internet are mainly 2D images, which cannot be used for the training and construction of 3D bone joint point extraction models. While the latter has problems such as insufficient sample diversity and a single imaging background. Training a deep learning model with such data is likely to lead to overfitting of the model.
[0030] Problem (3): The systematic error of manual annotation is difficult to quantify and eliminate.
[0031] Currently, the annotation of large-scale bone joint point data is mainly carried out through multi-person collaboration. The annotation process and quality are affected by many subjective variables such as the experience, working status, and coordination of the annotators. It is very difficult to evaluate the systematic error of the annotation results. This leads to a large fluctuation range in the prediction results of the subsequent trained bone joint point extraction model. For example, often even for a person who is stationary in the video, there will still be an obvious drift phenomenon in the bone joint point positions of the front and back frames predicted by the model, resulting in unstable performance of many subsequent action analysis algorithms or models.
[0032] To solve the above technical problems, according to some current experimental research evidence: The human bone joint points mainly reflect the geometric information of the human body, such as height, weight, body proportion of the head and body, and limbs, etc., and have a low correlation with information such as the color and texture of the human body imaging. Therefore, the embodiment of this application provides a method for generating training data for a human bone joint point extraction model. The method may include: performing three-dimensional (3D) modeling on the target human body to generate 3D action skinned data of the target human body, and determining the 3D coordinates of each bone joint point in each 3D action frame of the target human body according to the association relationship between the 3D coordinates of each predefined bone joint point and the corresponding subset of the 3D mesh set, as well as the 3D action skinned data of the target human body; where the association relationship is predefined according to the task requirements; performing two-dimensional (2D) rendering on each 3D action frame according to the task requirements, generating 2D image data of the target human body according to the rendering result, and performing coordinate transformation on the 3D coordinates of the bone joint points to generate 2D pixel coordinates of the bone joint points.
[0033] That is, the embodiment of the present application proposes a method for automatically generating image data and annotations related to skeletal joint points based on human 3D motion modeling, animation production, and image rendering technologies. This method not only saves time and effort but also meets the training and testing requirements of the vast majority of current skeletal joint point recognition algorithms and models, effectively alleviating problems such as high costs for manual data collection and annotation, long cycles, and the inability to eliminate systematic errors in annotation. At the same time, since each 3D motion frame is rendered in two dimensions (2D) according to the task requirements, that is, during the data generation process, the content of the generated data can be quickly adjusted according to the characteristics of model training and application scenarios. For example, by adjusting parameters such as human body shape, movement, clothing, camera parameters, and environmental background, more abundant human body image data can be obtained compared to natural images collected manually, which is more conducive to improving the training effect and scene adaptability of deep learning models.
[0034] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0035] Refer to Figure 1 , which shows a schematic flowchart of a method for generating training data for a human skeletal joint point extraction model provided by an embodiment of the present application. As Figure 1 shown, the method may include:
[0036] S110. Perform 3D human motion modeling on a target human body to generate 3D motion skinning data of the target human body.
[0037] Among them, the target human body may refer to the human body of the person to be collected, and the person to be collected includes multiple persons. The characteristics of the multiple persons (such as age, height, body shape, clothing, gender, etc.) may be the same or different, without limitation. Specifically, before executing the method of the present application, it is pre-configured which persons are the persons to be collected, so as to collect relevant information of the human bodies of these persons for training to obtain a skeletal joint point extraction model.
[0038] Among them, the 3D motion skinning data of the target human body can be called 3D mesh data or 3D motion frame sequence. The 3D mesh data can include a 3D mesh set corresponding to each 3D motion frame during the occurrence of the target human body's motion. The 3D mesh set can be called a 3D human body mesh set or a human body 3D mesh set. The 3D mesh set can represent the human body surface in a 3D space. A 3D mesh set can include one or more mesh blocks (or called 3D mesh blocks), and a mesh block is a triangular plane determined by three vertices in 3D space. The relative positions of the mesh blocks included in the 3D mesh set are related to the skeletal joint points of the human body and the human body pose parameters. The human body pose parameters can include human body motion parameters and / or body shape parameters. The human body shape can be called the human body build, and the body shape parameters specifically include parameters such as the head-to-body ratio of the human body, the height, weight, and build of the human body, and clothing.
[0039] A motion of the human body can be represented by the relative position relationship formed by a set of specific human skeletal joint points. A change in the human body's pose will cause a change in the relative positions of the respective mesh blocks in the 3D mesh set. At the same time, a change in the human body shape, such as a change in height, weight, and build, will also cause a change in the relative positions of the respective mesh blocks in the 3D mesh set. Specifically, through a statistical machine learning method, the mapping relationship between the human body pose parameters (such as the data structure of the relative position relationship of the human skeletal joint points and the human body shape) and the corresponding subset of the 3D mesh set can be learned. The corresponding subset of the 3D mesh set can refer to a part of the mesh blocks included in the 3D mesh set. For example, the corresponding subset of the 3D mesh set can be the mesh blocks formed by the relative positions of the human skeletal joint points in multiple key poses during the continuous motion of the target human body. As shown in the following formula (1), it shows the mapping relationship between the human skeletal joint points and the human body pose and the corresponding subset M of the 3D mesh set:
[0040] , formula (1)
[0041] Among them, θ in formula (1) is the human body motion parameter of the target human body. θ represents the angles of each skeletal joint point in the 3D coordinate system with the center point root of the human body as the origin, and θ reflects the human body motion. β in formula (1) is the body shape parameter of the target human body. β represents the distances of each skeletal joint point relative to the center point root of the human body, and β reflects the head-to-body ratio, height, weight, and build of the target human body. represents the pose mapping model parameters. Among them, the center point root of the human body in this application can be one of the skeletal joint points of the human body. Specifically, which skeletal joint point it is can be set as needed. For example, it is the center point of the human pelvis, without limitation.
[0042] In the embodiments of the present application, the human 3D motion modeling of the target human body essentially represents the human body surface of the target human body in a 3D space using a 3D mesh set. The human 3D motion modeling of the target human body may refer to designing the actions of the target human body according to task requirements, such as designing a set of actions, and performing character modeling on the target human body, such as setting body parameters such as gender, age, muscle ratio, weight, height, and race for the 3D character model corresponding to the target human body.
[0043] Exemplarily, S110 may include S1101-S1103 as Figure 2 shown:
[0044] S1101. According to task requirements, perform motion modeling on the target human body, set the relative positions of the human skeletal joints at multiple key postures in a coherent motion, and obtain a sequence of pose parameters of the target human body.
[0045] Among them, the actions of the target human body are preset according to task requirements. The sequence of pose parameters includes the human motion parameters θ corresponding to each 3D motion frame during the occurrence of the actions of the target human body. Specifically, S1101 may include the following steps 1)-step 5):
[0046] Step 1). Mark the skeletal joints of the target human body on the model. For example, as Figure 3 shown, the model can be made to wear a specific / specially made wearable device, and the skeletal joints of the model are displayed and marked on the wearable device. Optionally, the skeletal joints can be marked with small balls with infrared light sources.
[0047] Step 2). Use multiple infrared cameras (or called high-speed infrared light cameras) to collect action images of the model from different angles. Among them, the action images may refer to the images generated during the occurrence of the actions of the model human body. The action images may also be called 3D animation images. The action images may include multiple action image frames. One action image frame may correspond to one acquisition moment, corresponding to one action of the model. The action image frame may be called a 3D motion frame or a 3D animation image frame, without limitation. In the present application, the 3D motion frame is taken as an example for description, and this is uniformly described here and will not be repeated later.
[0048] Since infrared cameras are used for collection, and the skeletal joints of the model human body are marked with marks that can emit infrared light sources, the foreground pixels of the collected action images represent the skeletal joints.
[0049] Step 3) Align the action images collected by multiple infrared cameras. For example, align the action images collected by multiple infrared cameras at the same moment to obtain the 3D action frames corresponding to the model human body at that moment. Traverse each moment and perform the frame alignment operation to obtain the 3D action frames corresponding to all the collection moments during the action occurrence process of the model human body.
[0050] Step 4) Using the camera imaging principle, based on the internal and external parameters of multiple infrared cameras and the 3D action frames obtained after performing Step 3) at each collection moment, calculate the coordinate sequence of the skeletal joint points during the action occurrence process of the model human body , where represents the 3D coordinates of each skeletal joint point in the infrared camera space at the i-th moment (which can be called the collection moment), and the value range of i is [1, t].
[0051] Step 5) Perform coordinate translation transformation on the 3D coordinates of each skeletal joint point in the infrared camera space to convert them into 3D coordinates in a 3D coordinate system with the human body center point root as the coordinate origin, and calculate the posture parameter sequence of the action occurrence process of the target human body . Where represents the coordinates (which can be called 3D coordinates) of the skeletal joint point in the 3D coordinate system with the human body center point root as the coordinate origin at the i-th moment (which can be called the collection moment), the value range of i is [1, t], and t is an integer greater than 1.
[0052] S1102. According to the task requirements, perform character modeling on the target human body through 3D character modeling software to obtain the body posture parameters of the target human body.
[0053] Among them, the posture parameters include the body posture parameters β corresponding to each 3D action frame during the action occurrence process of the target human body. It can be understood that S1102 can also be described as performing character modeling (or constructing a 3D character) on the target human body to obtain the body posture parameters of the target human body.
[0054] In the embodiments of the present application, performing character modeling on the target human body through 3D character modeling software can refer to setting parameters such as the gender, age, clothing, and body type of the target human body.
[0055] Specifically, S1102 can be as Figure 4As shown in the figure, through 3D character modeling software, set the body parameters of the target human body such as gender, age, muscle proportion, weight, height, and race according to the task requirements, and obtain the body parameters of a 3D nude human body with a specific body shape. Based on the 3D nude human body model and according to the task requirements, use the 3D character modeling software to add clothing, shoes, hats and other clothing decorations to the 3D nude human body model, and obtain the 3D body parameters β integrated with specific clothing decorations.
[0056] Among them, the 3D character modeling software of the present application may include Make human. Make human is an open-source 3D character modeling software based on a large amount of anthropological morphological feature data. Make human can quickly form male and female face and limb models of different ages, and adjust the local body shapes of the formed male and female face and limb models.
[0057] It should be noted that the present application does not limit the execution order of S1101 and S1102. It can execute S1101 first and then S1102, or execute S1102 first and then S1101, or execute S1101 and S1102 simultaneously, without limitation.
[0058] S1103. Generate the 3D action skinning data of the target human body based on the body pose parameters of the target human body.
[0059] Among them, as above, the body pose parameters of the target human body may include the pose parameter sequence of the target human body obtained in S1101 and the body parameters of the target human body obtained in S1102.
[0060] Specifically, according to the task requirements, set parameters such as action duration and frequency, and use the linear interpolation method to interpolate between the body pose parameters at two adjacent sampling times ( , ), and ( , ) to obtain the body pose sequence of a coherent action (such as a coherent 3D action). . Input each body pose parameter in the body pose sequence into the mapping model
[0061] in turn, calculate the 3D mesh set corresponding to each 3D action frame under the coherent action of the target human body, and sort the 3D mesh sets corresponding to each 3D action frame to generate the 3D action skinning data of the target human body.
[0062] S120. Determine the 3D coordinates of the skeletal joint points in each 3D action frame according to the association relationship between the 3D coordinates of each predefined skeletal joint point and the corresponding subset of the 3D mesh set, and the 3D action skinning data of the target human body generated in S110.
[0063] In the embodiments of the present application, the 3D coordinates of the skeletal joint points in the 3D action frame can be the annotation information of the skeletal joint points in the 3D action frame, and the 3D coordinates are the coordinates of the skeletal joint points in the 3D action frame in the three-dimensional coordinate system with the center point of the human body as the origin. The 3D coordinates of the skeletal joint points of the target human body can be referred to / understood as the annotation information of the skeletal joint points of the target human body.
[0064] In the embodiments of the present application, the association relationship between the 3D coordinates of each skeletal joint point and the corresponding subset of the 3D mesh set can be defined / determined in advance according to the task requirements. Specifically, the 3D mesh sets corresponding to each 3D action frame of the target human body can be uniformly numbered. For example, each mesh block included in the 3D mesh set can be encoded, and each mesh block has a unique number, and the same mesh blocks corresponding to different 3D action frames use the same number. Further, according to the task requirements, the 3D coordinates of each skeletal joint point to be annotated are defined / determined based on the corresponding subset of the 3D mesh set. For example, the center point of the surface formed by the mesh blocks included in the corresponding subset of the 3D mesh set can be used as the 3D coordinates of the skeletal joint point to be annotated.
[0065] For example, as shown in the following formula (2), the 3D coordinates of the skeletal joint point can be expressed as:
[0066] Formula (2)
[0067] wherein, in formula (2), represents the th mesh block in the corresponding subset of the 3D mesh set, and the mesh block is a vector containing 3 3D coordinates. represents the total number of mesh blocks included in the corresponding subset of the 3D mesh set.
[0068] It should be noted that the task requirements in the embodiments of the present application can be replaced and described as task requirements, and the task requirements are mainly used to design the image acquisition scenario of the skeletal joint points. For example, the task requirements mainly specify the camera view angle corresponding to the person to be collected, the camera positions of multiple cameras, the data collection site, the composition of the person to be collected (age, height, body type, clothing, gender), the actions to be collected, etc. Optionally, if it is a multi-camera and multi-camera acquisition scenario, the calibration of the internal and external parameters of the camera also needs to be performed.
[0069] S130. Render each 3D action frame into 2D according to the task requirements, and obtain the 2D image data of the target human body and the 2D pixel coordinates of the skeletal joints of the target human body based on the rendering results.
[0070] The annotation information of the skeletal joints of the target human body may include the 2D pixel coordinates of the skeletal joints of the target human body. In other words, the 2D pixel coordinates of the skeletal joints of the target human body can be referred to / understood as the annotation information of the skeletal joints of the target human body.
[0071] Specifically, S130 may include: (1) Based on a general 3D action modeling software (such as any one of Blender, Unity3D, 3DMax, etc.), import / input the 3D mesh set corresponding to each 3D action frame of the target human body's coherent actions into the general 3D action modeling software to generate a 3D image of the target human body during the occurrence of coherent actions. (2) Establish the scene where the target human body's actions occur according to the task requirements. For example, establish the scene where the target human body's actions occur as shown in Figure 5 According to the task requirements, add multiple cameras to the scene where the target human body's actions occur to form multiple perspectives and set the internal and external parameters of the cameras (including focal length, lens size, etc.). (3) According to the task requirements, establish the background, other scenery, etc. in the scene where the target human body's actions occur; according to the task requirements, set and adjust the position and intensity of the light source in the scene where the target human body's actions occur. (4) Set and load an image rendering engine in the general 3D action modeling software, perform 2D rendering on each 3D action frame of the target human body from different camera perspectives in the scene where the target human body's actions occur, and save the rendering results after 2D rendering as the 2D image data of the target human body. (5) According to the internal and external parameters of the camera and the camera imaging principle formula, convert the 3D coordinates of the skeletal joints in each 3D action frame into 2D pixel coordinates in the corresponding camera perspective and save them in a specific annotation file format. Thus, the annotation of the 2D pixel coordinates of the skeletal joints in the image is automatically completed.
[0072] Among them, the 3D coordinates of the skeletal joints in each 3D action frame can be obtained with reference to S120. For example, for each 3D action frame, according to the association relationship between the pre-defined 3D coordinates of each skeletal joint and the corresponding subset of the 3D mesh set, obtain the 3D coordinates of each skeletal joint. The specific process can refer to S120 and will not be elaborated here.
[0073] It should be noted that the 2D image data of the target human body and the annotation information of the skeletal joint points generated in this application (such as the 3D coordinates of the skeletal joint points and the 2D pixel coordinates of the skeletal joint points) can be referred to as the training data of the human skeletal joint point extraction model and can be used to train the human skeletal joint point extraction model. Specifically, how to use the generated image data and annotation information to train the human skeletal joint point extraction model is not discussed in this application. Specifically, it can refer to the existing training process and will not be elaborated.
[0074] Based on Figure 1 the method shown, a 3D mesh set corresponding to each 3D action frame during the occurrence of the target human body action is generated. According to the association relationship between the predefined subset of the 3D mesh set and the 3D coordinates of the corresponding skeletal joint points, the 3D coordinates of the skeletal joint points are obtained, and the automatic annotation of the 3D coordinates of the skeletal joint points is completed. According to the task requirements, the 3D action frames are rendered into 2D images from each camera perspective, and the 3D coordinates of the skeletal joint points are calculated and transformed into the pixel coordinates of the 2D images from the corresponding perspectives through the camera optical imaging principle, and the automatic annotation of the 2D pixel coordinates of the skeletal joint points is completed. In this way, not only the automatic annotation of the skeletal joint points is realized, but also this method can be changed in a timely manner according to the task requirements. The system error is stable and easy to eliminate. At the same time, the automatically generated annotation data is used as the supervision information to train the skeletal joint point extraction model. The fluctuation range of its prediction results is smaller than that of the manually annotated data, the prediction performance is more stable, and the cost of the image data acquisition and annotation process is lower than that of the traditional method, the efficiency is higher, and the quality of the image data and the diversity of the samples are easier to control than the natural image acquisition.
[0075] For example, experiments show that: the human skeletal joint points mainly reflect the geometric information of the human body, such as height, weight, head-to-body and limb proportions, etc., and have a low correlation with the information such as color and texture of the human body imaging. The method for automatically generating 2D image data and annotation information of skeletal joint points based on processes such as human 3D action modeling, animation production, and 2D image rendering proposed in this embodiment can meet the training and testing requirements of most current skeletal joint point recognition algorithms and models. At the same time, due to the stable system error of the automatic annotation method, using the automatically generated annotation data as the supervision information to train the skeletal joint point extraction model, the fluctuation range of its prediction results is smaller than that of the model trained with manually annotated data, and the performance is more stable.
[0076] Referring to Figure 6 , which shows a schematic structural diagram of a training data generation device for a human skeletal joint point extraction model described according to an embodiment of the present application. As Figure 6 shown, the training data generation device 600 for the human skeletal joint point extraction model may include: a first generation module 601, a determination module 602, and a second generation module 603.
[0077] The first generation module 601 is used to perform 3D human body motion modeling on the target human body to generate 3D motion skinned data of the target human body; wherein, the 3D motion skinned data includes a 3D mesh set corresponding to each 3D motion frame under continuous actions during the occurrence of the target human body motion.
[0078] The determination module 602 is used to determine the 3D coordinates of each bone joint point in each 3D motion frame of the target human body according to the predefined association relationship between each bone joint point and the corresponding subset of the 3D mesh set, and the 3D motion skinned data of the target human body; wherein, the association relationship is predefined according to the task requirements.
[0079] The second generation module 603 is used to perform 2D rendering on each 3D motion frame according to the task requirements, generate 2D image data of the target human body according to the rendering result, and perform coordinate transformation on the 3D coordinates of the bone joint points to generate 2D pixel coordinates of the bone joint points.
[0080] Optionally, the first generation module 601 is further used for:
[0081] Perform 3D human body motion modeling on the target human body to obtain the human body pose parameters of the target human body. The human body pose parameters include the pose parameter sequence of the target human body and the body shape parameters of the target human body. The pose parameter sequence of the target human body includes the pose parameters corresponding to multiple key poses in the continuous actions of the target human body. The pose parameters of the target human body are used to represent the relative positions of the bone joint points of the target human body; the body shape parameters of the target human body are used to represent the human body shape characteristics of the target human body.
[0082] According to the task requirements, interpolate between the human body pose parameters at adjacent sampling moments in sequence to obtain the human body pose sequence of the target human body. The human body pose sequence includes the human body pose parameter sequence and the body shape parameters of the target human body under the continuous actions of the target human body.
[0083] Input each human body pose parameter included in the human body pose sequence into the mapping model to generate 3D motion skinned data of the target human body; wherein, the mapping model is a model representing the mapping relationship between the human body pose parameters and the 3D mesh set.
[0084] Optionally, the first generation module 601 is further used for:
[0085] According to the task requirements, perform motion modeling on the target human body, set the relative positions of the bone joint points of multiple key poses in the continuous actions, and obtain the pose parameter sequence of the target human body.
[0086] According to the task requirements, perform character modeling on the target human body through 3D character modeling software to obtain the body shape parameters of the target human body.
[0087] Optionally, the subset of the 3D mesh set includes multiple mesh blocks, and the association relationship is predefined according to the task requirements, including:
[0088] According to the task requirements, the center point of the surface formed by the mesh blocks included in the subset corresponding to the 3D mesh set is used as the 3D coordinates of the skeletal joint point to be labeled.
[0089] Optionally, the 3D coordinates of the skeletal joint point to be labeled are expressed as: ;
[0090] wherein, represents the th mesh block in the subset corresponding to the 3D mesh set, and the mesh block is a vector containing 3 3D coordinates; represents the total number of mesh blocks included in the subset corresponding to the 3D mesh set.
[0091] Optionally, the second generation module 603 is further configured to:
[0092] Establish the scene where the target human body movement occurs according to the task requirements;
[0093] Set and load an image rendering engine in a general 3D motion modeling software, perform 2D rendering on each 3D motion frame of the target human body under different camera perspectives, and save the rendered result after 2D rendering as the 2D image data of the target human body;
[0094] According to the internal and external parameters of the camera and the camera imaging principle formula, convert the 3D coordinates of the skeletal joint points in each 3D motion frame into 2D pixel coordinates in the corresponding camera perspective.
[0095] Optionally, the second generation module 603 is further configured to:
[0096] Add multiple cameras in the scene where the target human body movement occurs according to the task requirements to form multiple perspectives, and set the internal and external parameters of the cameras;
[0097] Establish the background, scenery in the scene where the target human body movement occurs according to the task requirements, and set and adjust the position and intensity of the light source in the scene where the target human body movement occurs.
[0098] The training data generation device for a human skeletal joint point extraction model provided in this embodiment can execute the embodiments of the above method, and its implementation principle and technical effects are similar, and will not be described in detail here.
[0099] Figure 7 This is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. AsFigure 7 As shown, a schematic structural diagram of an electronic device 700 suitable for implementing the embodiments of the present application is shown.
[0100] As Figure 7 shown, the electronic device 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage section 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 are also stored. The CPU 701, ROM 702, and RAM 703 are connected to each other via a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.
[0101] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. The drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed so that a computer program read from it can be installed into the storage section 708 as needed.
[0102] Specifically, according to the embodiments of the present disclosure, the process described above with reference to Figure 1 can be implemented as a computer software program. For example, the embodiments of the present disclosure include a computer program product, which includes a computer program tangibly contained on a machine-readable medium, and the computer program includes program code for executing the above-mentioned method for generating training data of the human skeleton joint point extraction model. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 709, and / or installed from the removable medium 711.
[0103] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0104] The units or modules described in the embodiments of the present application can be implemented in software or in hardware. The described units or modules can also be provided in a processor. The names of these units or modules do not, in some cases, constitute a limitation on the units or modules themselves.
[0105] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a mobile phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0106] On the other hand, the present application also provides a storage medium. The storage medium can be the storage medium included in the aforementioned device in the above embodiments; or it can exist separately and be not assembled into the device. The storage medium stores one or more programs, and the aforementioned programs are used by one or more processors to execute the method for generating training data of the human skeletal joint point extraction model described in the present application.
[0107] A storage medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0108] It should be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0109] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the description of the method embodiment.
Claims
1. A method for generating training data of a human body bone joint point extraction model, characterized in that The method includes: Performing human 3D motion modeling on a target human body to generate 3D motion skinning data of the target human body; wherein, the 3D motion skinning data includes a 3D mesh set corresponding to each 3D motion frame under a continuous motion during the occurrence of the motion of the target human body. Determine the 3D coordinates of each bone joint point in each 3D action frame of the target human body according to the association relationship between each predefined bone joint point and the corresponding subset of the 3D mesh set, and the 3D action skinning data of the target human body; wherein, the corresponding subset of the 3D mesh set includes multiple mesh blocks, and the association relationship is predefined according to the task requirements, including: according to the task requirements, using the center point of the surface formed by the mesh blocks included in the corresponding subset of the 3D mesh set as the 3D coordinates of the bone joint point to be labeled, and the 3D coordinates of the bone joint point to be labeled is expressed as: , the represents the th mesh block in the corresponding subset of the 3D mesh set, and the mesh block is a vector containing 3 3D coordinates; the represents the total number of mesh blocks included in the corresponding subset of the 3D mesh set; Performing 2D rendering on each 3D motion frame according to the task requirements, generating 2D image data of the target human body according to the rendering result, and performing coordinate transformation on the 3D coordinates of the bone joint points to generate 2D pixel coordinates of the bone joint points.
2. The method according to claim 1, characterized in that The performing human 3D motion modeling on a target human body to generate 3D motion skinning data of the target human body includes: Performing human 3D motion modeling on the target human body to obtain human pose parameters of the target human body, where the human pose parameters include a sequence of pose parameters of the target human body and body shape parameters of the target human body, the sequence of pose parameters of the target human body includes pose parameters corresponding to multiple key poses in the continuous motion of the target human body, and the pose parameters of the target human body are used to represent the relative positions of the bone joint points of the target human body; the body shape parameters of the target human body are used to represent the human body shape characteristics of the target human body. Interpolating between the human pose parameters at adjacent sampling moments in sequence according to the task requirements to obtain a human pose sequence of the target human body, where the human pose sequence includes a sequence of human pose parameters under the continuous motion of the target human body and the body shape parameters of the target human body. Inputting each human pose parameter included in the human pose sequence into a mapping model to generate 3D motion skinning data of the target human body; wherein, the mapping model is a model representing the mapping relationship between human pose parameters and a 3D mesh set.
3. The method according to claim 2, wherein The human pose parameters include a sequence of pose parameters of the target human body and body shape parameters of the target human body; the performing human 3D motion modeling on the target human body to obtain human pose parameters of the target human body includes: Performing motion modeling on the target human body according to the task requirements, setting the relative positions of the human bone joint points of multiple key poses in the continuous motion, and obtaining a sequence of pose parameters of the target human body. Performing character modeling on the target human body through 3D character modeling software according to the task requirements to obtain the body shape parameters of the target human body.
4. The method according to claim 1, characterized in that, The performing 2D rendering on each 3D motion frame according to the task requirements, generating 2D image data of the target human body according to the rendering result, and performing coordinate transformation on the 3D coordinates of the bone joint points to generate 2D pixel coordinates of the bone joint points includes: Establishing a scene where the motion of the target human body occurs according to the task requirements. Setting and loading an image rendering engine in a general 3D motion modeling software, performing 2D rendering on each 3D motion frame of the target human body from different camera perspectives in the scene, and saving the rendering result after 2D rendering as 2D image data of the target human body. According to the internal and external parameters of the camera and the camera imaging principle formula, the 3D coordinates of the bone joint points in each 3D action frame are converted into 2D pixel coordinates in the corresponding camera view.
5. The method according to claim 4, wherein The establishment of the scene where the target human action occurs according to the task requirements includes: According to the task requirements, multiple cameras are added to the scene where the target human action occurs to form multiple views, and the internal and external parameters of the cameras are set. According to the task requirements, a background, scenery are established in the scene where the target human action occurs, and the position and intensity of the light source in the scene where the target human action occurs are set and adjusted.
6. A training data generation device for a human body bone joint point extraction model, characterized in that, The device includes: A first generation module for performing 3D human action modeling on the target human body to generate 3D action skinning data of the target human body; wherein, the 3D action skinning data includes a 3D mesh set corresponding to each 3D action frame in the continuous action during the movement of the target human body. A determination module, configured to determine the 3D coordinates of each skeletal joint point in each 3D action frame of the target human body according to the association relationship between the 3D coordinates of each predefined skeletal joint point and the corresponding subset of the 3D mesh set, and the 3D action skinning data of the target human body; wherein, the corresponding subset of the 3D mesh set includes a plurality of mesh blocks, and the association relationship is predefined according to the task requirements, including: according to the task requirements, using the center point of the surface formed by the mesh blocks included in the corresponding subset of the 3D mesh set as the 3D coordinates of the skeletal joint point to be labeled, and the 3D coordinates of the skeletal joint point to be labeled is expressed as: , the represents the th mesh block in the corresponding subset of the 3D mesh set, and the mesh block is a vector containing 3 3D coordinates; the represents the total number of mesh blocks included in the corresponding subset of the 3D mesh set; A second generation module for performing 2D rendering on each 3D action frame according to the task requirements, generating 2D image data of the target human body according to the rendering result, and converting the 3D coordinates of the bone joint points to generate 2D pixel coordinates of the bone joint points.
7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the training data generation method of the human bone joint point extraction model as described in any one of claims 1-5.
8. A readable storage medium, on which a computer program is stored, characterized in that, When the program is executed by the processor, it implements the training data generation method of the human bone joint point extraction model as described in any one of claims 1-5.
Citation Information
Patent Citations
Skeletal animation vertex correction method and model learning method, device and apparatus
CN113554736A
Apparatus and method for generating synthetic training data for motion recognition
US20190295278A1