Human body posture estimation data generation method and electronic equipment
The parameterized human body model deformation and motion migration technology generates diverse human body data, which solves the limitations of data acquisition and labeling in human body posture estimation, and achieves efficient and accurate human body posture estimation.
Patent Information
- Application Number
- CN202510486980.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-08-15
AI Technical Summary
In the prior art, in human posture estimation, data acquisition costs are high, it is difficult to cover complex body shapes and extreme movements, and manual labeling introduces errors, resulting in insufficient generalization of the model.
Generate diverse human data through parameterized human model deformation and human movement migration technology, combining 3D to 2D projection and automatic node annotation to reduce labeling costs and errors.
It improves the diversity and accuracy of human body data, covers complex body shapes and extreme movements, reduces labeling costs and errors, and improves the accuracy and generalization of human body posture estimation.
Smart Images

Figure CN120496167A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of computer vision and provides a data generation method and electronic device for human posture estimation. Background Art
[0002] With the development of artificial intelligence (AI) and deep learning technologies, AI applications (such as fitness apps and somatosensory games) are rapidly becoming popular on electronic devices such as smart TVs, mobile phones, and AR / VR devices. Vision-based human pose estimation (HPE) technology has become the core technology for AI applications to achieve natural human interaction.
[0003] Human pose estimation aims to detect key points of the human body from images or videos and construct a representation of human motion. It is a core technology for tasks such as action recognition, virtual human driving, and motion analysis. The detection performance is highly dependent on the diversity and annotation quality of the training data.
[0004] Currently, data collection for human posture estimation generally relies on wearable sensors or multi-camera systems. However, the cost of single-experiment collection is high, and it can only cover a limited number of action types in laboratory environments. In addition, long-tail scenarios (such as medical rehabilitation) are missing. At the same time, rare body shapes and high-risk actions (such as falls, extreme sports, high flexibility, etc.) are difficult to collect and cover. Subsequent data labeling will also introduce labeling errors, resulting in the human posture estimation model being unable to accurately detect human posture, which in turn affects the response quality of the application. Summary of the Invention
[0005] The embodiments of the present application provide a data generation method and electronic device for human posture estimation, which are used to address the limitations of human body data collection and labeling.
[0006] In a first aspect, an embodiment of the present application provides a method for generating data for human posture estimation, comprising:
[0007] Deforming the parameterized human body model according to a plurality of shape parameters to obtain a plurality of target human body models;
[0008] For each target human model, perform the following operations:
[0009] According to a pre-generated human motion sequence, the target human body model is driven to move to obtain a human body mesh model;
[0010] Setting lighting information in a virtual environment where the human mesh model resides, and placing occluders for the human mesh model based on the lighting information, the occluders including shadows of the human mesh model and obstacles with shadows in the virtual space;
[0011] For each frame of human motion, the human mesh model and the target skeleton of the human mesh model are projected from multiple perspectives to obtain multiple 2D human images and the 2D coordinates of the first joint point of the target skeleton visible in each human image. The multiple human images are fused with the scene image respectively, and the 2D coordinates of the first joint points in the multiple human images are combined and stored in a human image dataset for human posture estimation.
[0012] The beneficial effects of the above technical solution are: by dynamically adjusting the shape parameters of the parameterized human body model, target human body models of different body shapes are obtained, so that the generated human body image dataset can cover people with complex body shapes, thereby improving the diversity of human body data; and, by migrating the human body motion sequence to each target human body model, the diversity of human body motions of different body shapes is expanded, further improving the diversity of human body data. In this way, for human body mesh models of different body shapes driven by each frame of human body motion, after 3D to 2D projection of the human body mesh model and the corresponding target skeleton at different perspectives, multiple human body images and the 2D coordinates of the first visible joint point in each human body image can be automatically obtained, without the need for manual joint point labeling, reducing the labeling cost while reducing the labeling error, thereby ensuring that the human body image dataset obtained after fusion with the scene image is used for human posture estimation, which can effectively improve the accuracy of human posture estimation.
[0013] On the other hand, it supports the automatic configuration of environmental parameters such as lighting information and obstructions, thereby achieving low-cost, highly controllable generation of diversified human body data while ensuring the authenticity of the data.
[0014] Optionally, setting lighting information in the virtual environment where the human body mesh model is located, and placing occluders for the human body mesh model according to the lighting information, includes:
[0015] Sampling is performed in the spatial area of the virtual environment, the position of each sampling point is regarded as a light source, and a light intensity is generated for each light source to obtain the light information;
[0016] Calculating a bounding box of the human body mesh model, and an offset radius between a center of the generated bounding box and the obstacle;
[0017] Determining, based on the illumination information, a first direction of a shadow of the human mesh model and a second direction of an obstacle with a shadow;
[0018] A shadow of the human mesh model is generated according to the first direction, and an obstacle with the shadow is placed according to the second direction and the offset radius.
[0019] The beneficial effects of the above technical solution are: determining the direction of the occluder by simulating the lighting information of the light source in the virtual environment, and forming the occluder of the human body model based on the shadow of the human body mesh model and the obstacles with shadows in the virtual space, enhancing the complexity of light and shadow interaction, obtaining reasonable lighting and shadow effects, and thus improving the authenticity of the data.
[0020] Optionally, projecting the human body mesh model and the target skeleton of the human body mesh model from multiple perspectives to obtain multiple 2D human body images and the 2D coordinates of the first joint point of the target skeleton visible in each human body image includes:
[0021] For any one of the multiple perspectives, perform the following operations:
[0022] Projecting the surface vertices of the human body mesh model onto a plane according to calibration parameters of the virtual camera corresponding to the viewing angle to obtain a human body image;
[0023] According to the calibration parameters of the virtual camera corresponding to the viewing angle, the first joint point visible within the viewing angle on the target skeleton of the human mesh model is detected, and the visible first joint point is projected onto the human body image to obtain the 2D coordinates of the visible first joint point in the human body image.
[0024] The beneficial effects of the above technical solution are: by projecting the human body mesh model through the calibration parameters of virtual cameras at different perspectives, human body images at different perspectives can be automatically obtained without the need for camera acquisition, and the amount of human body data is increased. Moreover, by projecting the first visible joint point on the target skeleton through the calibration parameters of virtual cameras at different perspectives, the first visible joint point in each human body image is obtained, without the need for manual labeling, which reduces the labeling cost and reduces the labeling error.
[0025] Optionally, the visibility detection process of each first joint point on the target skeleton includes:
[0026] determining the optical center of the virtual camera according to calibration parameters of the virtual camera corresponding to the viewing angle;
[0027] For any 3D point in each of the first joint points, perform the following operations:
[0028] Establishing a ray with the optical center pointing to a 3D point, and obtaining an intersection point between the ray and an object in the virtual environment;
[0029] When the intersection point is located in a preset value interval, it is determined that the 3D point is invisible; wherein the value interval is determined according to the second distance between the 3D point and the optical center and a preset skin distance threshold.
[0030] The beneficial effect of the above technical solution is: by performing visibility detection on the first joint point in the target skeleton, automatic labeling of the joint point is achieved.
[0031] Optionally, at least one source skeleton corresponding to the human motion sequence and the target skeleton respectively include multiple joint point pairs, and each joint point pair includes: a first joint point in the target human skeleton and a second joint point in the source skeleton;
[0032] Then, the target human body model is driven to move according to the pre-generated human body motion sequence to obtain a human body mesh model, including:
[0033] For each frame of human body action, perform the following steps:
[0034] Determining a target rotation parameter of a first joint point in the plurality of joint point pairs according to a source rotation parameter of a second joint point corresponding to the first joint point in the plurality of joint point pairs under a corresponding human body motion, in combination with a predetermined transformation matrix between the source skeleton and the target skeleton;
[0035] Determining a target displacement parameter of the first joint point according to a source displacement parameter of the second joint point corresponding to the first root node in the target skeleton under the corresponding human body motion and in combination with a predetermined displacement difference between the source skeleton and the target skeleton;
[0036] The surface vertices of the target human body model are deformed according to the skin weights of the respective first joint points to obtain a human body mesh model.
[0037] The beneficial effect of the above technical solution is: by migrating the human body motion sequence of the source skeleton to the target skeleton of the target human body model of different body types, the difficulty of collecting some extreme movements is reduced, the diversity of human body postures is increased, and the generalization of subsequent human body posture estimation is improved.
[0038] Optionally, when the number of second joint points included in the source skeleton is less than the number of first joint points included in the target skeleton, for a first joint point in the target skeleton that does not have a corresponding second joint point in the source skeleton, perform the following operations:
[0039] The target rotation parameter of the first joint point is generated based on the target rotation parameter of at least one adjacent joint point of the first joint point, and the target displacement parameter of the first joint point is generated based on the target displacement parameter of at least one adjacent joint point of the first joint point; wherein each adjacent joint point has a corresponding second joint point in the source skeleton.
[0040] The beneficial effect of the above technical solution is: when the number of second joint points in the source skeleton is less than the number of first joint points in the target skeleton, for the first light node that does not have a corresponding second joint point, its rotation and displacement parameters are determined through at least one adjacent joint point, thereby achieving normal driving of the first light node that does not have a corresponding second joint point, and improving the driving accuracy of the human target model.
[0041] Optionally, the transformation matrix and the displacement difference are obtained in the following manner:
[0042] When the source skeleton and the target skeleton maintain the same posture, calculating a transformation matrix for converting rotation parameters of second joint points in a plurality of joint point pairs included in the source skeleton and the target skeleton into rotation parameters of corresponding first joint points;
[0043] When the source skeleton and the target skeleton maintain the same posture, the displacement difference is calculated according to the displacement parameters of the second root node of the source skeleton and the first root node of the target skeleton, and the scale difference between the source skeleton and the target skeleton.
[0044] The beneficial effect of the above technical solution is: when the source skeleton and the target skeleton are in the same posture, the transformation matrix is calculated based on the rotation parameters of multiple joint point pairs in the two skeletons, and the displacement difference is calculated based on the displacement parameters of the root nodes in the two skeletons, thereby achieving the alignment of the coordinate systems of the source skeleton and the target skeleton, so that the target human body model can be driven by the human body movements of the source skeleton.
[0045] In a second aspect, an embodiment of the present application provides an electronic device, comprising a processor, a memory, and a display screen, wherein the display screen, the memory, and the processor are connected via a bus;
[0046] The display screen is used to display a model and image of a human body;
[0047] The memory stores a computer program, and the processor performs the following operations according to the computer program:
[0048] Deforming the parameterized human body model according to a plurality of shape parameters to obtain a plurality of target human body models;
[0049] For each target human model, perform the following operations:
[0050] According to a pre-generated human motion sequence, the target human body model is driven to move to obtain a human body mesh model;
[0051] Setting lighting information in a virtual environment where the human mesh model resides, and placing occluders for the human mesh model based on the lighting information, the occluders including shadows of the human mesh model and obstacles with shadows in the virtual environment;
[0052] For each frame of human motion, the human mesh model and the target skeleton of the human mesh model are projected from multiple perspectives to obtain multiple 2D human images and the 2D coordinates of the first joint point of the target skeleton visible in each human image. The multiple human images are fused with the scene image respectively, and the 2D coordinates of the first joint points in the multiple human images are combined and stored in a human image dataset for human posture estimation.
[0053] Optionally, the processor sets lighting information in the virtual environment where the human body mesh model is located, and places occluders for the human body mesh model according to the lighting information, specifically by:
[0054] Sampling is performed in the spatial area of the virtual environment, the position of each sampling point is regarded as a light source, and a light intensity is generated for each light source to obtain the light information;
[0055] Calculating a bounding box of the human body mesh model, and an offset radius between a center of the generated bounding box and the obstacle;
[0056] Determining, based on the illumination information, a first direction of a shadow of the human mesh model and a second direction of an obstacle with a shadow;
[0057] A shadow of the human mesh model is generated according to the first direction, and an obstacle with the shadow is placed according to the second direction and the offset radius.
[0058] Optionally, the processor projects the human body mesh model and the target skeleton of the human body mesh model from multiple perspectives to obtain multiple 2D human body images and the 2D coordinates of the first joint point of the target skeleton visible in each human body image, specifically by:
[0059] For any one of the multiple perspectives, perform the following operations:
[0060] Projecting the surface vertices of the human body mesh model onto a plane according to calibration parameters of the virtual camera corresponding to the viewing angle to obtain a human body image;
[0061] According to the calibration parameters of the virtual camera corresponding to the viewing angle, the first joint point visible within the viewing angle on the target skeleton of the human mesh model is detected, and the visible first joint point is projected onto the human body image to obtain the 2D coordinates of the visible first joint point in the human body image.
[0062] Optionally, the processor detects a first visible joint point on the target skeleton in the following manner:
[0063] determining the optical center of the virtual camera according to calibration parameters of the virtual camera corresponding to the viewing angle;
[0064] For any 3D point in each of the first joint points, perform the following operations:
[0065] Establishing a ray with the optical center pointing to a 3D point, and obtaining an intersection point between the ray and an object in the virtual environment;
[0066] When the intersection point is located in a preset value interval, it is determined that the 3D point is invisible; wherein the value interval is determined according to the second distance between the 3D point and the optical center and a preset skin distance threshold.
[0067] Optionally, at least one source skeleton corresponding to the human motion sequence and the target skeleton respectively include multiple joint point pairs, and each joint point pair includes: a first joint point in the target human skeleton and a second joint point in the source skeleton;
[0068] The processor drives the target human body model to move according to the pre-generated human body motion sequence to obtain a human body mesh model. The specific operations are:
[0069] For each frame of human body action, perform the following steps:
[0070] Determining a target rotation parameter of a first joint point in the plurality of joint point pairs according to a source rotation parameter of a second joint point corresponding to the first joint point in the plurality of joint point pairs under a corresponding human body motion, in combination with a predetermined transformation matrix between the source skeleton and the target skeleton;
[0071] Determining a target displacement parameter of the first joint point according to a source displacement parameter of the second joint point corresponding to the first root node in the target skeleton under the corresponding human body motion and in combination with a predetermined displacement difference between the source skeleton and the target skeleton;
[0072] The surface vertices of the target human body model are deformed according to the skin weights of the respective first joint points to obtain a human body mesh model.
[0073] Optionally, when the number of second joint points included in the source skeleton is less than the number of first joint points included in the target skeleton, the processor performs the following operations for a first joint point in the target skeleton that does not have a corresponding second joint point in the source skeleton:
[0074] The target rotation parameter of the first joint point is generated based on the target rotation parameter of at least one adjacent joint point of the first joint point, and the target displacement parameter of the first joint point is generated based on the target displacement parameter of at least one adjacent joint point of the first joint point; wherein each adjacent joint point has a corresponding second joint point in the source skeleton.
[0075] Optionally, the processor obtains the transformation matrix and the displacement difference in the following manner:
[0076] When the source skeleton and the target skeleton maintain the same posture, calculating a transformation matrix for converting rotation parameters of second joint points in a plurality of joint point pairs included in the source skeleton and the target skeleton into rotation parameters of corresponding first joint points;
[0077] When the source skeleton and the target skeleton maintain the same posture, the displacement difference is calculated according to the displacement parameters of the second root node of the source skeleton and the first root node of the target skeleton, and the scale difference between the source skeleton and the target skeleton.
[0078] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of any one of the above-mentioned methods for generating data for estimating human motion are implemented.
[0079] The technical effects brought about by any implementation method in the second to third aspects can refer to the technical effects brought about by the corresponding implementation method in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0081] Figure 1 A flowchart of a method for generating data for estimating human motion provided in an embodiment of the present application;
[0082] Figure 2This is a diagram showing the deformation effect of the parametric human body model controlled by shape parameters provided in an embodiment of the present application;
[0083] Figure 3 A flow chart of a method for migrating human body movements per frame provided in an embodiment of the present application;
[0084] Figure 4 This is an effect diagram of human body motion migration from the source skeleton to the target skeleton provided in the embodiment of the present application;
[0085] Figure 5 A schematic diagram of the process of setting lighting and shadows provided in an embodiment of the present application;
[0086] Figure 6 A schematic diagram of the 3D to 2D projection process provided in an embodiment of the present application;
[0087] Figure 7A Schematic diagram of a four-channel RGBA human body image after projection;
[0088] Figure 7B A schematic diagram of the annotation of the first joint point visible in the projected human body image;
[0089] Figure 8 This is the fusion effect diagram of the human body image and the scene image;
[0090] Figure 9 A system architecture diagram of a method for generating data for estimating human motion provided in an embodiment of the present application;
[0091] Figure 10 This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0092] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of the technical solutions of this application, but not all of them. Based on the embodiments described in this application document, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the technical solutions of this application.
[0093] Based on the exemplary embodiments shown in this application, all other embodiments obtained by persons of ordinary skill in the art without inventive effort are within the scope of protection of this application. In addition, although the disclosure in this application is presented based on one or several exemplary examples, it should be understood that each aspect of the disclosure can independently constitute a complete technical solution.
[0094] In addition, the terms "comprises" and "comprising" and any variations thereof are intended to cover but not exclude inclusion, for example, a product or device comprising a list of components is not necessarily limited to those components expressly listed but may include other components not expressly listed or inherent to such product or device.
[0095] The term "module" as used in this application refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.
[0096] Human pose estimation combined with deep learning models has become an important direction in the field of computer vision. The powerful characteristics of deep learning have significantly improved the accuracy and efficiency of human pose estimation, and have achieved good results in real-time and posture recognition in complex backgrounds, making human pose estimation technology widely used in various fields.
[0097] Human posture estimation technology based on deep learning detects human joints (such as head, shoulders, elbows, knees, etc.) in real time, analyzes movement standardization and motion trajectory, and thus provides functions such as movement correction, training feedback, and human-computer interaction.
[0098] Currently, human pose estimation models primarily rely on 2D images captured by monocular cameras for detection. This requires massive amounts of multi-perspective human body annotation data (e.g., the same action requires multiple planar projections, including front, side, and top views) to compensate for depth loss, leading to exponentially increasing data acquisition and annotation costs. Furthermore, pixel-level deviations caused by perspective ambiguity (e.g., limb shortening effects) during manual annotation can lead to inconsistencies in the spatial characteristics of joints (e.g., elbow coordinates exceeding the shoulder-wrist line), further amplifying model training errors. Furthermore, in long-tail scenarios such as medical rehabilitation and high-risk actions, the difficulty in obtaining annotated data on humans with complex body shapes and movements severely restricts the application and promotion of models.
[0099] In view of this, an embodiment of the present application provides a data generation method for human posture estimation, which automatically generates human body data of various human body movements of different body shapes through parametric human body model deformation technology and human body motion migration technology, effectively improving the diversity of human body data, thereby covering human body data of rare body shapes and extreme movements (such as: simulating the gait of scoliosis patients or the unbalanced posture in fall detection), and improving the problem of insufficient generalization of human posture estimation models caused by incomplete traditional data collection. In addition, while obtaining the imaging of the human mesh model in the virtual camera under different perspectives through 3D to 2D projection technology, the human joint points contained in the human body image are also automatically obtained, and the pixel coordinates and human joint point positions are generated simultaneously, thereby avoiding the labeling errors caused by manual interpretation deviations from the source, improving labeling efficiency, and reducing labeling costs.
[0100] On the other hand, it supports automatic configuration of environmental parameters such as lighting information and occlusions, thereby achieving low-cost, highly controllable generation of diverse human body data while ensuring the authenticity of the data.
[0101] Exemplary embodiments of the present application are described below with reference to the accompanying drawings.
[0102] See also Figure 1 , is a flow chart of a method for generating data for human motion estimation provided in an embodiment of the present application, the process mainly includes the following steps:
[0103] S101: deforming a parameterized human body model according to a plurality of shape parameters to obtain a plurality of target human body models.
[0104] The parametric human body model is loaded into the rendering engine, and the parametric human body model is deformed by randomly generating a plurality of human body shape parameters to obtain a plurality of target human body models, wherein each shape parameter corresponds to a target human body model.
[0105] Taking the parameterized human body model as the SMPL model, the formula is expressed as:
[0106]
[0107] Among them, β∈R n is the n-dimensional shape parameter, θ∈R m*3 is the posture parameter expressed by m human joints in 3D, Standard human body mesh template, B s and B p are the deformation functions of shape and posture, β~D shape ,θ~D pose , D shape represents a mixed distribution containing regular body types and long-tailed body types, D pose Represents the angle constraint distribution of joint points based on biomechanics.
[0108] In specific implementations, the standard deviation of the randomly generated shape parameters can range from -3 to 3, generating diverse target human models with heights ranging from 1.4m to 2.1m and BMIs from 16 to 35. The long-tail distribution of shape parameters is sampled (e.g., μ = 0, σ = 1.5) to cover special body types (e.g., those with scoliosis). After completing shape parameter customization, the target skeleton corresponding to each shape parameter is automatically assigned skinning weights to the Mesh model through the skeletal binding panel. The target human model of the corresponding body type is then exported as an FBX file containing the target skeleton's skeletal information (e.g., the coordinates and index numbers of each first light node).
[0109] like Figure 2As shown in FIG, a schematic diagram of the effect of controlling the deformation of the parametric human body model by β. The target human body model after deformation and the parametric human body model before deformation have different attributes such as height, posture, and body proportion.
[0110] After obtaining multiple target human body models, the correction tool can be used to optimize the clothing model to fit each target human body model so that texture mapping can be performed on the driven target human body model later.
[0111] It should be noted that the embodiment of the present application does not impose any restrictive requirements on the values of the randomly generated shape parameters, which can be set according to actual scenarios.
[0112] S102: For each target human body model, drive the target human body model to move according to a pre-generated human body motion sequence to obtain a human body mesh model.
[0113] Among them, each frame of human motion in the human motion sequence can be expressed by human joint points.
[0114] In some embodiments, a human motion sequence may be expressed using human joints of at least one source skeleton, and the topological structure of each source skeleton and the target skeleton used by the target human body model may be the same or different.
[0115] Before migrating a human motion sequence to a target human model, a bone topology mapping relationship and bone coordinate system alignment between at least one source skeleton and the target skeleton may be established in advance.
[0116] Taking a source skeleton as an example, the process of establishing the skeleton topology mapping relationship between the source skeleton and the target skeleton includes: T ={j t1 ,j t2 ,…,j tm} (m is the number of first joints), and the set J consisting of each second joint in the source skeleton S ={j s1 ,j s2 ,…,j sn} (n is the number of second joints) as input, and a mapping function M:J is constructed based on the human skeleton structure. S →J T , clarify the correspondence between the first joint point and the second joint point (e.g., the left shoulder joint point in the source skeleton corresponds to the left shoulder joint point in the target skeleton), and obtain multiple joint point pairs, each joint point pair includes: a first joint point in the target human skeleton and a second joint point in the source skeleton.
[0117] The process of aligning the bone coordinate systems between the source skeleton and the target skeleton includes two steps: determining the transformation matrix of the rotation parameters and determining the displacement difference of the displacement parameters.
[0118] Taking a source skeleton as an example, the process of aligning the bone coordinate systems between the source skeleton and the target skeleton includes:
[0119] First, when the source skeleton and the target skeleton maintain the same posture, multiple joint point pairs contained in the source skeleton and the target skeleton are calculated, and the rotation parameters of the second joint points in the multiple joint point pairs are converted into the transformation matrix of the rotation parameters of the corresponding first nodes.
[0120] In the specific implementation, the source skeleton and the target skeleton are initialized to a T-shaped posture (arms flat, legs upright), and each joint point pair (j s ,j t ) in the T-shaped posture (including the orientation (i.e., rotation parameters) and position (i.e., displacement parameters) of the joint points on the XYZ axes of the skeleton coordinate system), and solve the transformation matrix that converts the rotation parameters of the second joint point in multiple joint point pairs into the rotation parameters of the corresponding first joint point through least squares fitting or singular value decomposition method, denoted as R s→t ∈SO(3).
[0121] For example, if the second joint point points to the left on the Y axis, and the corresponding first joint point points to the front on the Y axis, then R s→t It must include a 90° vertical rotation around the Y axis.
[0122] In some embodiments, when the source skeleton and the target skeleton are different in the same posture, a quaternion rotation compensation algorithm can be used to align the skeleton coordinate system, and the rotation direction mismatch problem can be solved through the skeleton axial transformation matrix (XYZ→YZX).
[0123] Then, the displacement difference is calculated based on the displacement parameters of the second root node of the source skeleton and the first root node of the target skeleton in the same posture, and the scale difference between the source skeleton and the target skeleton.
[0124] In the specific implementation, assuming that the second root node is the hip joint point in the source skeleton and the first root node is the hip joint point in the target skeleton, then in the T-shaped posture, the difference in the displacement parameters of the two root nodes is calculated, and the difference is dynamically scaled according to the proportional difference between the source skeleton and the target skeleton to obtain the displacement difference Δd between the source skeleton and the target skeleton. root , thereby using the displacement difference of the root node to compensate for the global position offset of subsequent skeletal animation.
[0125] It should be noted that the embodiment of the present application does not impose any restrictive requirements on the posture maintained when the skeletal coordinate system is aligned. For example, the source skeleton and the target skeleton can also be initialized to an A-shaped posture.
[0126] When the source skeleton and the target skeleton are in the same posture, the transformation matrix is calculated based on the rotation parameters of multiple joint point pairs in the two skeletons, and the displacement difference is calculated based on the displacement parameters of the root nodes in the two skeletons and the proportional difference between the two skeletons, so as to achieve the alignment of the coordinate systems of the source skeleton and the target skeleton, so that the target human body model can be driven by the human body motion of the source skeleton to achieve motion transfer.
[0127] In some embodiments, the bone length of the target skeleton can be predefined during the modeling phase to ensure that after rotational migration, the limb extension ratio of the target human body model matches its own bone length, such as a long-arm model swinging its arms with a natural amplitude.
[0128] After determining the bone topology mapping relationship and bone coordinate system alignment between the source skeleton and the target skeleton, the human motion sequence is imported into the rendering engine to form a skeletal animation. The rendering engine migrates the human motion sequence to multiple target human models based on the joint point mapping relationship between the source skeleton and the target skeleton, thereby obtaining human motions of different body shapes.
[0129] See also Figure 3 The migration process of each frame of human body motion provided in the embodiment of the present application mainly includes the following steps:
[0130] S1021: Determine the target rotation parameters of the first joint point among the multiple joint point pairs based on the source rotation parameters of the second joint point corresponding to the first joint point among the multiple joint point pairs under the corresponding human body movement, combined with the predetermined transformation matrix between the source skeleton and the target skeleton.
[0131] For each frame of human motion, we traverse the corresponding source skeleton and target skeleton's multiple joint point pairs frame by frame, and calculate the target rotation parameter of the first joint point to achieve local coordinate alignment. The formula is as follows:
[0132]
[0133] in, The matrix representing the target rotation parameters of the first joint point in multiple joint point pairs, The matrix representing the source rotation parameters of the second joint point in multiple joint point pairs, k represents the kth frame in the human action sequence.
[0134] In some embodiments, when the number of second joints included in a source skeleton of a human motion sequence is less than the number of first joints included in the target skeleton, some first joints in the target skeleton do not have corresponding second joints in the source skeleton. At this time, for the first joint that does not have a corresponding second joint, its target rotation parameter can be determined by the target rotation parameter of at least one adjacent joint that has a corresponding second joint in the source skeleton. Specifically. For the first joint that does not have a corresponding second joint, the target rotation parameter of the first joint is generated based on the target rotation parameter of at least one adjacent joint of the first joint, and the target displacement parameter of the first joint is generated based on the target displacement parameter of at least one adjacent joint of the first joint.
[0135] Optionally, the target rotation parameter of a first joint point that does not have a corresponding second joint point can be the average of the target rotation parameters of at least one adjacent joint point, or the interpolation of the target rotation parameters of at least one adjacent joint point; similarly, the target displacement parameter of a first joint point that does not have a corresponding second joint point can be the average of the target displacement parameters of at least one adjacent joint point, or the interpolation of the target displacement parameters of at least one adjacent joint point.
[0136] When the number of second joint points in the source skeleton is less than the number of first joint points in the target skeleton, for the first light node that does not have a corresponding second joint point, its rotation and displacement parameters are determined through at least one adjacent joint point, thereby achieving normal driving of the first joint point that does not have a corresponding second joint point and improving the driving accuracy of the human target model.
[0137] S1022: For each frame of human motion, determine the target displacement parameter of the first root node in the target skeleton based on the source displacement parameter of the second root node in the source skeleton under the corresponding human motion and the predetermined displacement difference between the source skeleton and the target skeleton.
[0138] For each frame of human motion, the target displacement parameters of the first root node in the target skeleton are calculated frame by frame to perform global motion compensation (such as running, walking, etc.). The formula is as follows:
[0139]
[0140] in, Represents the global target displacement parameter of the first root node in the target skeleton, represents the global source displacement parameter of the second root node in the source skeleton, and k represents the kth frame in the human motion sequence.
[0141] S1023: Deforming the surface vertices of the target human body model according to the skin weights of the respective first joint points to obtain a human body mesh model.
[0142] After the first joints of the target skeleton are locally changed and globally aligned with the root node, the target skeleton and the source skeleton have the same human body motion, such as Figure 4 As shown, further, for each frame of human body action, the corresponding surface vertices of the target human body model are deformed according to the skin weights of the respective first joints, thereby obtaining target human body models of different actions.
[0143] By migrating the human motion sequence of the source skeleton to the target skeleton of target human models of different body shapes, the difficulty of collecting some extreme motions is reduced, the diversity of human postures is increased, and the generalization of subsequent human posture estimation is improved.
[0144] It should be noted that the embodiments of the present application do not impose any restrictive requirements on the method of generating a human motion sequence set. For example, it can be collected from text or video based on a motion extraction model built based on deep learning, or it can be collected using a motion capture system, or it can be a fusion of multiple collection methods.
[0145] S103: Setting lighting information in the virtual environment where the human body mesh model is located, and placing occluders of the human body mesh model according to the lighting information.
[0146] For the three-dimensional virtual environment, multiple highly randomized light sources are generated in the rendering engine to simulate the real ambient light, and occluders are placed on the human mesh model to generate reasonable lighting and shadow effects, thereby improving the authenticity of the human body data.
[0147] In some embodiments, the occluders of the human mesh model include the shadow of the human mesh model and obstacles with shadows in the virtual environment. Figure 5 , which is the process of setting up lighting and shadows, mainly includes the following steps:
[0148] S1031: Sampling is performed in the spatial area of the virtual environment, the position of each sampling point is regarded as a light source, and light intensity noise is generated for each light source to obtain light information.
[0149] The lighting information includes the position and intensity of the light source.
[0150] Generally, a virtual environment is a three-dimensional space area (such as a three-dimensional room or a spherical outdoor area). Sampling is performed within this space area, supporting sampling at any height and direction, and the position of each sampling point is used as a light source to obtain a dynamically distributed light source position.
[0151] Taking uniform sampling as an example, the formula for the position distribution of the light source is expressed as:
[0152] P light~Uniform(Ω 3D )Formula 4
[0153] in, Define the three-dimensional space area of the virtual environment, P light Indicates the position of the light source.
[0154] It should be noted that the embodiments of the present application do not impose any restrictive requirements on the sampling method. For example, random sampling, Gaussian sampling, etc. can also be used.
[0155] Furthermore, a light intensity is assigned to each light source obtained by sampling.
[0156] Assume that the light intensity is normally distributed I(light)~Distribution(μ I ,σ I ), I(light)=μ I +∈(light),∈(light)~Ν(0,σ I 2 ), μ I and σ I represent mean and standard deviation respectively.
[0157] In some embodiments, the rendering engine may be divided into a multi-layer lighting system for multiple light sources at different positions to improve the realism of the light sources.
[0158] Taking the three-layer lighting system as an example, it includes the main light source lighting system, the auxiliary light source lighting system and the contour light lighting system. Among them, the lighting intensity of the main light source in the main light source lighting system varies randomly in the range of 800-1200 lux, and the auxiliary light source in the auxiliary light source lighting system is composed of multiple groups (such as 4 groups) of movable point light sources with a color temperature in the range of 2700K-6500K. The contour light in the contour light lighting system uses a directional spotlight with a cone angle in the range of 15°-60°.
[0159] S1032: Calculate the bounding box of the human body mesh model, and the offset radius between the center of the generated bounding box and the obstacle.
[0160] In the rendering engine, an occlusion generator is set around the human mesh model (e.g., within a range of 2 meters) to randomly generate a certain number of obstacles that are adapted to the virtual environment. The material system automatically assigns translucent materials such as frosted glass and metal mesh to the obstacles to ensure that the occlusion area occupies 15%-40% of the screen.
[0161] For example, when the virtual environment is an indoor environment, 5-10 household obstacles such as sofas, coffee tables, cups, murals, and potted plants are set up.
[0162] In order to improve the realism of the virtual environment, these obstacles need to be associated with the human body mesh model. In the specific implementation, an AABB bounding box (Axis-Aligned Bounding Box, AABB) is generated for the human body mesh model, and the center of the bounding box is denoted as P human , then randomly generate obstacles to P human Offset radius R offset , where the offset radius is used to control the horizontal distance range between the occluder and the human mesh model.
[0163] Optionally, the offset radius is 1.5 meters and can be adjusted according to the actual scenario.
[0164] S1033: Determine a first shadow direction of the human body mesh model and a second shadow direction of the obstacle according to the illumination information.
[0165] According to the position distribution characteristics of the light source Φ light , determining a first direction of the shadow of the human mesh model and a second direction of the obstacle with the shadow. This is because when the direction of the obstacle is determined, the direction of the shadow of the obstacle is also determined.
[0166] For example, when the light source is mainly located in front of the human mesh model, the obstacles tend to be placed on the back side of the human mesh model to avoid complete shading. When the light source is evenly distributed, the directions of the obstacles are evenly sampled in the three-dimensional space.
[0167] S1034: Generate a shadow of the human body mesh model according to the first direction, and place an obstacle with the shadow according to the second direction and the offset radius.
[0168] After determining the second direction of the obstacle, we also need to determine the position of the obstacle relative to the human mesh model. Specifically, we can randomly generate the position of the occluder based on the center point of the bounding box of the human mesh model. The formula is as follows:
[0169] P object =P human +δ Formula 5
[0170] Where, δ=(r x ,r y ,r z ), satisfying ||δ||≤R offset , r z ∈[z min ,z max ],z min and z max It refers to the vertical range (e.g. from the ground to a height of 2 meters).
[0171] Finally, the location distribution of obstacles can be expressed as:
[0172] object~Placement(P human ,R offset ,Φ light ) Formula 6
[0173] Typically, the orientation of the light source determines the space between the human mesh model and the light source. In some embodiments, this space can be partitioned into smaller subspaces, ensuring that the distribution of occluders within these subspaces is random, rather than having only one occluder per light source position. By introducing a spatial partitioning strategy, the distribution of occluders along the light source's projection direction is more uniform, resulting in greater realism.
[0174] In an embodiment of the present application, by automatically configuring a simulated light source in an environment, determining the direction of an occluder based on the lighting information of the light source, and forming an occluder of the human body model based on the shadow of the human body mesh model and obstacles with shadows in the virtual space to enhance the complexity of light and shadow interaction, reasonable lighting and shadow effects are obtained, thereby achieving low-cost, highly controllable generation of diversified human body data while ensuring the authenticity of the data.
[0175] S104: For each frame of human motion, project the human mesh model and the target skeleton of the human mesh model from multiple perspectives to obtain multiple 2D human images and the first joint point of the target skeleton visible in each human image.
[0176] In the rendering engine, multiple (e.g., 8) virtual cameras are arranged around the human body mesh model to form a spherical array (e.g., horizontal interval of 45°, pitch angle of ±30°), thereby obtaining multiple acquisition perspectives, and multiple virtual cameras are used to acquire images at a set sampling rate (e.g., 512), and the acquired images are subjected to noise reduction processing.
[0177] Before projection, for each frame of human action, the rendering engine uses a preset clothing library to texture only the human mesh model, ignoring scene objects and background, and setting a transparent background image, that is, the transparency of the background image's Alpha channel is 0.
[0178] After mapping, for any of the multiple perspectives, the 3D to 2D projection process is as follows Figure 6 As shown, it mainly includes the following steps:
[0179] S1041: Projecting the surface vertices of the human body mesh model onto a plane according to calibration parameters of the virtual camera corresponding to the viewing angle to obtain a human body image.
[0180] The calibration parameters include intrinsic parameters K and extrinsic parameters (including rotation matrix R and translation vector t).
[0181] In practice, for each frame of the human body mesh model, the 3D coordinates of any surface vertex on the human body mesh model are transformed from the world coordinate system to the camera coordinate system using the rotation matrix R and the translation vector t. The 3D coordinates of the surface vertex in the camera coordinate system are then transformed onto a plane using the intrinsic parameter K to obtain a 2D pixel point. If the pixel point exceeds the image boundary, it indicates that the surface vertex is invisible and is directly discarded. After completing the projection of all surface vertices, a 2D human image is obtained.
[0182] like Figure 7A As shown, it is a schematic diagram of the human body image of RGBA four channels after projection.
[0183] It should be noted that the embodiment of the present application does not impose any restrictive requirements on the size and type of the human body image. For example, the human body image is a 1920*1080 PNG image.
[0184] S1042: Detecting a first joint point on a target skeleton of a human mesh model that is visible within the viewing angle according to calibration parameters of the virtual camera corresponding to the viewing angle.
[0185] After human body animation, the joints seen from different perspectives are different. Therefore, it is necessary to perform visibility detection on each first joint point in the target skeleton. The specific detection process is as follows:
[0186] S1042_1: Determine the optical center of the virtual camera according to calibration parameters of the virtual camera corresponding to the viewing angle.
[0187] Specifically, the optical center O of the virtual camera can be expressed in the world coordinate system as:
[0188] O=-R T ·t Formula 7
[0189] Among them, R is the rotation matrix of the virtual camera, and t is the translation vector of the virtual camera.
[0190] S1042_2: For any 3D point in each first joint point, establish a ray with the optical center pointing to the 3D point, and obtain the intersection point of the ray and the object in the virtual environment.
[0191] Specifically, the direction of the ray can be expressed as:
[0192]
[0193] The ray can be expressed as:
[0194]
[0195] Where l represents the distance from the point on the ray to the optical center of the virtual camera.
[0196] In some embodiments, the intersection point of the ray and the object in the virtual environment may be the intersection point of the ray and the obstacle, or the intersection point of the ray and the human body.
[0197] S1042_3: Determine whether the intersection point is within a preset value range. If so, execute S1042_4; otherwise, execute S1042_5.
[0198] Among them, the value range is [0, l max -τ] is determined based on the second distance between the first joint point and the optical center and the preset skin distance threshold τ. The second distance l max The formula is expressed as:
[0199] l max =|P 3D -O| Formula 10
[0200] Optionally, the skin distance threshold τ is usually equal to the radius of the joint point, for example, when the 3D point is a knee joint point, τ=5 cm.
[0201] S1042_4: Determine that the 3D point is invisible.
[0202] S1042_5: Determine whether the 3D point is visible.
[0203] S1043: For each visible first joint point, project the first joint point onto the human body image, and obtain the 2D coordinates of the first joint point in the human body image.
[0204] Assume that the 3D coordinates of the first visible h-th joint point in the world coordinate system are By rotating the matrix R and translating the vector t, the hth first joint point is positioned in the world coordinate system. Converted to the camera coordinate system, the formula is expressed as:
[0205]
[0206] Furthermore, the intrinsic parameter K is used to position the hth first joint point in the camera coordinate system. Convert it to the human body image and obtain the 2D coordinates of the hth first joint point in the human body image. The formula is expressed as:
[0207]
[0208] in,
[0209] like Figure 7B , which is a schematic diagram of the annotation of the first joint point visible in the projected human body image.
[0210] In some embodiments, when obtaining each first joint point in a human body image, a commonly used COCO (Common Objects in Context) dataset format containing 17 joint points may be used.
[0211] In some embodiments, the 2D coordinates of the first joint point visible in each human body image may be stored in a file in JSON format.
[0212] It should be noted that the embodiment of the present application does not impose any restrictive requirements on the output format of the 2D coordinates of the first joint point visible in the human body image. For example, it can also be stored in a table format.
[0213] By projecting the human mesh model using the calibration parameters of virtual cameras at different perspectives, human body images from different perspectives can be automatically obtained without camera acquisition, which increases the amount of human body data. In addition, by projecting the first visible joint point on the target skeleton using the calibration parameters of virtual cameras at different perspectives, the 2D coordinates of the first visible joint point in each human body image can be obtained without manual labeling, which reduces labeling costs and reduces labeling errors.
[0214] S105: For each frame of the human action sequence, multiple human images are fused with the scene image respectively, and the 2D coordinates of each first joint point in the multiple human images are combined and stored in a human image dataset for human posture estimation.
[0215] The scene image may be a scene image of a virtual environment or a scene image of a real environment, and may be adjusted according to the application scenario of subsequent human posture estimation.
[0216] For any human image from multiple perspectives corresponding to each frame of human action sequence, the human image is used as the foreground image and the scene image as the background image. The pixels in the two images are fused pixel by pixel according to different transparency levels. The fusion formula is expressed as:
[0217] I final (x,y)=α·I fore (x,y)+(1-α)·I back (x,y) Formula 13
[0218] Among them, α∈[0,1] represents the transparency of the human body map, I fore (x,y) represents the human body image, I back (x,y) represents the scene image.
[0219] like Figure 8 As shown in FIG, this is the fusion effect diagram of the human body image and the scene image.
[0220] For each frame of human action sequence, the fused image of the human body image and the scene image from each perspective is associated with the 2D coordinates of the first joint point visible in the corresponding human body image and stored in the human body image dataset for subsequent human posture estimation.
[0221] In some embodiments, during the 3D to 2D projection process of the human body mesh model, a mask map and a depth map of the human body may also be generated.
[0222] Assuming that there are 500 parameters in total for shape parameters, human motion and light source, by randomly combining these 500 parameters, we can obtain a human image dataset containing approximately 100,000 sample data.
[0223] See also Figure 9 , which is a system architecture diagram of the data generation method for human motion estimation provided in an embodiment of the present application, mainly including an input layer, a parameterized modeling layer, a motion migration layer, a random scene generation layer, a synchronous rendering and annotation layer, and an output layer. Among them, the input layer is used to prepare the human motion sequence of skeletal animation and randomly generate shape parameters for controlling the body shape; the parameterized modeling layer is used to load the SMPL model, sample the shape parameters of the SMPL model, and constrain the posture parameters; the motion migration layer is used to establish a skeletal topology mapping between the source skeleton and the target skeleton, align the coordinate systems of the source skeleton and the target skeleton, migrate the rotation parameters of the human motion sequence, and compensate for the root node displacement; the random scene generation layer is used to randomly generate the position and light intensity of the light source, and place occluders for the human body; the synchronous rendering and annotation layer is used to project the human mesh model from 3D to 2D, perform visibility detection on the first joint point of the target skeleton from multiple perspectives, fuse the human image with the scene image, and annotate the 2D coordinates of the visible first joint point in the projected human image; the output layer is used to output a fused image of the human image and the scene image and the 2D coordinates of the visible first joint point.
[0224] In an embodiment of the present application, by dynamically adjusting the shape parameters of a parameterized human body model, target human body models of different body shapes are obtained, so that the generated human body image dataset can cover people with complex body shapes, thereby improving the diversity of human body data. In addition, by migrating human motion sequences to each target human body model, the diversity of human body motions of different body shapes is expanded, further improving the diversity of human body data. In this way, for each frame of human body motion driven human body mesh models of different body shapes, after the human body mesh model and the corresponding target skeleton are projected from 3D to 2D at different perspectives, multiple human body images and the 2D coordinates of the first visible joint point in each human body image can be automatically obtained. This eliminates the need for manual joint point annotation, reducing annotation costs and annotation errors, thereby ensuring that the human body image dataset obtained after fusion with the scene image can effectively improve the accuracy of human body pose estimation when used for human pose estimation. On the other hand, by supporting the automated configuration of environmental parameters such as lighting information and occlusions, low-cost, highly controllable generation of diverse human body data is achieved while ensuring the authenticity of the data.
[0225] Based on the same technical concept, an embodiment of the present application provides an electronic device, which can be a laptop computer, a desktop computer, a tablet, a mobile phone, etc., which can implement the steps of the above-mentioned data generation method for human posture estimation and achieve the same technical effect.
[0226] See also Figure 10 , the electronic device includes a processor 1001, a memory 1002 and a display screen 1003, and the display screen 1003, the memory 1002 and the processor 1001 are connected via a bus 1004;
[0227] The display screen 1003 is used to display a model and image of the human body;
[0228] The memory 1002 stores a computer program, and the processor 1001 executes the computer program according to the computer program. Figure 1 The steps of the data generation method for human pose estimation are shown.
[0229] The above-mentioned memory 1002 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, and programs required to run instant messaging functions, etc.; the data storage area may store various instant messaging information and operating instruction sets, etc. The memory 1002 may be a volatile memory (volatile memory), such as a random-access memory (RAM); the memory 1002 may also be a non-volatile memory (non-volatile memory), such as a read-only memory, a flash memory (flash memory), a hard disk drive (HDD) or a solid-state drive (SSD); or the memory 1002 may be any other medium that can be used to carry or store a desired computer program in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1002 may be a combination of the above-mentioned memories. The processor 1001 may include one or more central processing units (CPUs), GPUs or digital processing units, etc. The processor 1001 is configured to implement the steps of any of the above-mentioned methods for generating data for estimating human posture when calling the computer program stored in the memory 1002 .
[0230] It should be noted that Figure 10 This is merely an example of the hardware necessary for an electronic device to execute the steps of the data generation method for human posture estimation provided in the embodiments of the present application. If not shown, the electronic device may also include conventional hardware such as a pickup, a microphone, a communication interface, a power supply, and a remote control device.
[0231] In the embodiment of the present application, the specific connection medium between the display screen 1003, the memory 1002 and the processor 1001 is not limited. In the embodiment of the present application, the display screen 1003 is connected to the bus 1004 between the memory 1002 and the processor 1001. Figure 10 The connections between the other components are shown in bold lines for illustration only and are not intended to be limiting. The bus 1004 can be divided into an address bus, a data bus, a control bus, etc. For ease of description, Figure 10 The diagram shows a single thick line, but this does not indicate that there is only one bus or one type of bus.
[0232] For the convenience of description, the electronic device can be divided into modules (or units) according to their functions and described separately. Of course, when implementing this application, the functions of each module (or unit) can be implemented in the same or multiple software or hardware.
[0233] Those skilled in the art will appreciate that various aspects of the present application can be implemented as systems, methods, or program products. Therefore, various aspects of the present application can be specifically implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation that combines hardware and software aspects, which may be collectively referred to herein as a "circuit," "module," or "system."
[0234] An embodiment of the present application also provides a computer-readable storage medium for storing some instructions. When these instructions are executed, the steps of any one of the data generation methods for human body posture estimation in the aforementioned embodiments can be completed.
[0235] An embodiment of the present application also provides a computer program product for storing a computer program, which is used to execute the steps of any one of the methods for generating data for human posture estimation in the aforementioned embodiments.
[0236] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0237] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0238] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device that implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0239] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0240] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
Claims
1. A data generation method for human motion estimation, characterized in that: The method comprises: Deforming the parameterized human body model according to a plurality of shape parameters to obtain a plurality of target human body models; For each target human model, perform the following operations: According to a pre-generated human motion sequence, the target human body model is driven to move to obtain a human body mesh model; Setting lighting information in a virtual environment where the human mesh model resides, and placing occluders for the human mesh model based on the lighting information, the occluders including shadows of the human mesh model and obstacles with shadows in the virtual environment; For each frame of human motion, the human mesh model and the target skeleton of the human mesh model are projected from multiple perspectives to obtain multiple 2D human images and the 2D coordinates of the first joint point of the target skeleton visible in each human image. The multiple human images are fused with the scene image respectively, and the 2D coordinates of the first joint points in the multiple human images are combined and stored in a human image dataset for human posture estimation.
2. The method according to claim 1, wherein Setting lighting information in a virtual environment where the human body mesh model is located, and placing occluders of the human body mesh model according to the lighting information, including: Sampling is performed in the spatial area of the virtual environment, the position of each sampling point is regarded as a light source, and a light intensity is generated for each light source to obtain the light information; Calculating a bounding box of the human body mesh model, and an offset radius between a center of the generated bounding box and the obstacle; Determining, based on the illumination information, a first direction of a shadow of the human mesh model and a second direction of an obstacle with a shadow; A shadow of the human mesh model is generated according to the first direction, and an obstacle with the shadow is placed according to the second direction and the offset radius.
3. The method according to claim 1, wherein The projecting of the human body mesh model and the target skeleton of the human body mesh model from multiple perspectives to obtain multiple 2D human body images and the 2D coordinates of the first joint point of the target skeleton visible in each human body image includes: For any one of the multiple perspectives, perform the following operations: Projecting the surface vertices of the human body mesh model onto a plane according to calibration parameters of the virtual camera corresponding to the viewing angle to obtain a human body image; According to the calibration parameters of the virtual camera corresponding to the viewing angle, the first joint point visible within the viewing angle on the target skeleton of the human mesh model is detected, and the visible first joint point is projected onto the human body image to obtain the 2D coordinates of the visible first joint point in the human body image.
4. The method according to claim 3, wherein The visibility detection process of each first joint point on the target skeleton includes: determining the optical center of the virtual camera according to calibration parameters of the virtual camera corresponding to the viewing angle; For any 3D point in each of the first joint points, perform the following operations: Establishing a ray with the optical center pointing to a 3D point, and obtaining an intersection point between the ray and an object in the virtual environment; When the intersection point is located in a preset value interval, it is determined that the 3D point is invisible; wherein the value interval is determined according to the second distance between the 3D point and the optical center and a preset skin distance threshold.
5. The method according to claim 1, wherein At least one source skeleton corresponding to the human motion sequence and the target skeleton respectively include multiple joint point pairs, each joint point pair includes: a first joint point in the target human skeleton and a second joint point in the source skeleton; Then, the target human body model is driven to move according to the pre-generated human body motion sequence to obtain a human body mesh model, including: For each frame of human body action, perform the following steps: Determining a target rotation parameter of a first joint point in the plurality of joint point pairs according to a source rotation parameter of a second joint point corresponding to the first joint point in the plurality of joint point pairs under a corresponding human body motion, in combination with a predetermined transformation matrix between the source skeleton and the target skeleton; Determining a target displacement parameter of the first joint point according to a source displacement parameter of the second joint point corresponding to the first root node in the target skeleton under the corresponding human body motion and in combination with a predetermined displacement difference between the source skeleton and the target skeleton; The surface vertices of the target human body model are deformed according to the skin weights of the respective first joint points to obtain a human body mesh model.
6. The method according to claim 5, wherein When the number of second joint points included in the source skeleton is less than the number of first joint points included in the target skeleton, for a first joint point in the target skeleton that does not have a corresponding second joint point in the source skeleton, perform the following operations: The target rotation parameter of the first joint point is generated based on the target rotation parameter of at least one adjacent joint point of the first joint point, and the target displacement parameter of the first joint point is generated based on the target displacement parameter of at least one adjacent joint point of the first joint point; wherein each adjacent joint point has a corresponding second joint point in the source skeleton.
7. The method according to claim 5 or 6, wherein: The transformation matrix and the displacement difference are obtained in the following manner: When the source skeleton and the target skeleton maintain the same posture, calculating a transformation matrix for converting rotation parameters of second joint points in a plurality of joint point pairs included in the source skeleton and the target skeleton into rotation parameters of corresponding first joint points; When the source skeleton and the target skeleton maintain the same posture, the displacement difference is calculated according to the displacement parameters of the second root node of the source skeleton and the first root node of the target skeleton, and the scale difference between the source skeleton and the target skeleton.
8. An electronic device, characterized in that: It includes a processor, a memory and a display screen, wherein the display screen, the memory and the processor are connected via a bus; The display screen is used to display a model and image of a human body; The memory stores a computer program, and the processor performs the following operations according to the computer program: Deforming the parameterized human body model according to a plurality of shape parameters to obtain a plurality of target human body models; For each target human model, perform the following operations: According to a pre-generated human motion sequence, the target human body model is driven to move to obtain a human body mesh model; Setting lighting information in a virtual environment where the human mesh model resides, and placing occluders for the human mesh model based on the lighting information, the occluders including shadows of the human mesh model and obstacles with shadows in the virtual environment; For each frame of human motion, the human mesh model and the target skeleton of the human mesh model are projected from multiple perspectives to obtain multiple 2D human images and the 2D coordinates of the first joint point of the target skeleton visible in each human image. The multiple human images are fused with the scene image respectively, and the 2D coordinates of the first joint points in the multiple human images are combined and stored in a human image dataset for human posture estimation.
9. The electronic device according to claim 8, wherein The processor sets illumination information in a virtual environment where the human body mesh model is located, and places an occluder of the human body mesh model according to the illumination information, specifically by: Sampling is performed in the spatial area of the virtual environment, the position of each sampling point is regarded as a light source, and a light intensity is generated for each light source to obtain the light information; Calculating a bounding box of the human body mesh model, and an offset radius between a center of the generated bounding box and the obstacle; Determining, based on the illumination information, a first direction of a shadow of the human mesh model and a second direction of an obstacle with a shadow; A shadow of the human mesh model is generated according to the first direction, and an obstacle with the shadow is placed according to the second direction and the offset radius.
10. The electronic device according to claim 8, wherein The processor projects the human body mesh model and the target skeleton of the human body mesh model from multiple perspectives to obtain multiple 2D human body images and the 2D coordinates of the first joint point of the target skeleton visible in each human body image, specifically by: For any one of the multiple perspectives, perform the following operations: Projecting the surface vertices of the human body mesh model onto a plane according to calibration parameters of the virtual camera corresponding to the viewing angle to obtain a human body image; According to the calibration parameters of the virtual camera corresponding to the viewing angle, the first joint point visible within the viewing angle on the target skeleton of the human mesh model is detected, and the visible first joint point is projected onto the human body image to obtain the 2D coordinates of the visible first joint point in the human body image.