Multi-person motion capture real-time digital human driving method based on skeleton-muscle interaction model
By introducing bone-muscle interaction model and multi-dimensional feature reasoning in multi-person motion capture technology, the problem of real-time driving motion capture in multi-person occlusion scenarios is solved, and the accurate and vivid action performance of digital people in complex scenarios is achieved.
Patent Information
- Application Number
- CN202510092293.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
In multi-person occlusion scenarios, it is difficult for the existing technology to realize real-time driving of multi-person motion capture, especially in complex scenarios. The occlusion phenomenon causes the accuracy of key point recognition to decrease, and it is difficult for the system to accurately judge the posture of each human body.
The method based on the bone-muscle interaction model is adopted, and the three-dimensional coordinates are reconstructed through two-stage dynamic occlusion point compensation and multi-dimensional proprietary-shared feature reasoning, and the "real person-skeleton model-digital person" real-time mapping method is proposed to combine mechanics and joint movement of the skeletal system to update the vertex position of the digital person in real time.
Real-time multiplayer motion capture drive in multiplayer occlusion scenarios is realized. Digital people can accurately reflect real moving postures, and their movements are smoother and more vivid, overcoming the problem of degradation of recognition accuracy caused by occlusion and diversity.
Smart Images

Figure CN120014705A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence and pan-metaverse, and relates to a multi-person motion capture real-time driving digital human method based on a skeleton-muscle interaction model. Background Art
[0002] With the rapid development of virtual reality technology, motion capture technology for virtual digital humans has gradually become a research hotspot year by year. According to different physical principles, it can be divided into mechanical, acoustic, electromagnetic, optical and inertial types. In the current industry, motion capture usually uses two main technologies: inertial and optical. Although inertial devices perform well in overcoming the common occlusion problems of VR devices and can achieve multi-target capture, due to their reliance on the physical properties of IMU (inertial measurement unit) and MEMS (micro-electromechanical system) devices, inertial motion capture systems have limitations in accurately tracking human postures for a long time. Optical motion capture devices have gradually become the mainstream of motion capture technology due to their higher capture accuracy and device stability. However, in actual production and life, people prefer to use markerless motion capture methods, which mainly rely on image recognition and analysis technology. Users do not need to wear any equipment, and computers use modern computer graphics theory and artificial intelligence technology to directly analyze the captured images to obtain virtual characters. At the same time, specific algorithms can also be implemented and continuously optimized for the accurate capture of different character movements in multi-person scenes. Therefore, in the future, markerless real-time driven digital human multi-person motion capture technology will have broad application prospects in virtual reality related fields.
[0003] Nowadays, how to drive digital humans to perform real-time actions corresponding to real people has become a hot issue, especially in complex scenes. For example, in environments where people gather or multiple people interact, there will be many occlusions between people and objects, which will significantly reduce the recognition accuracy of key points and make it difficult for the system to accurately judge the posture of each human body. In addition, the diversity of different body shapes, clothing, and postures also increases the difficulty of training the model, requiring the algorithm to have higher robustness. Therefore, the technology of multi-person motion capture to drive digital humans in real time is still a challenge to be solved. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide a method for real-time driving of digital human by motion capture of multiple people based on the skeleton-muscle interaction model. According to the analysis of human motion joints and occlusion conditions in multi-person occlusion scenes, a two-stage dynamic occlusion point compensation method is used to obtain the position of the occlusion point. Then, a multi-dimensional proprietary-shared feature reasoning method is used to reconstruct the three-dimensional coordinates, and a "real person-skeleton model-digital human" real-time mapping method is proposed. By combining the driving method of the skeleton-muscle interaction model, the position of the digital human vertex is updated in real time in combination with the action of mechanics and the joint movement of the skeleton system, and the layer-by-layer transmission from the vertex to the joint, skeleton, and muscle enables the digital human to accurately reflect the real motion posture after receiving the driving signal, thereby realizing the process of real-time driving of the digital human by motion capture of multiple people.
[0005] In order to achieve the above object, the present invention provides the following technical solutions:
[0006] A method for real-time driving digital human by multi-person motion capture based on a skeleton-muscle interaction model, the method comprising the following steps:
[0007] S1: Use YOLO-Pose to detect the position information and global confidence of all 17 skeleton key points of the human body, divide the key points into visible points and occluded points, and use a two-stage dynamic occluded point compensation method that integrates visible poses and inter-frame consistency to obtain the position of the occluded points;
[0008] S2: A 3D keypoint deep extractor using multi-dimensional proprietary-shared feature reasoning combined with the MTransPose model to output the 3D coordinates of the key points of the human body;
[0009] S3: Map the 3D key points of the human body obtained in S2 to the standard human skeleton model, bind the digital human character to the skeleton model through the weight mapping algorithm, and based on the bone-muscle interaction model, through the mechanical effects of the bone rigid body dynamics model and the muscle force line model, make the digital human accurately reflect the real movement posture, and realize the process of multi-person motion capture and real-time driving of the digital human.
[0010] Furthermore, the S1 specifically includes the following steps:
[0011] S11: Using the YOLO-Pose model, we obtain the key point distribution and global confidence scores of all people from top to bottom for the 17 skeletal joints of the human body; the position P of each key point i It is obtained by maximizing the response value of the key point heat map through the YOLO-Pose model, and the confidence f of each key point i It is determined by the maximum response value in the heat map of the point; if a key point i among the 17 key points of each person is not detected in the current frame or its confidence f i If it is less than the predetermined threshold θ, then the point is an occlusion point pi o ; If the confidence level f detected by the key point i If the value is higher than the threshold θ, the point p is considered visible. i v ;
[0012] S12: The first stage: compensation of occlusion points based on visible postures; first, the approximate position p of the occlusion point is inferred from the visible postures j o , and then calculate each potential visible point p i v With the current occlusion point p j o Distance d ij To determine whether it is a neighboring point, further use the angle relationship θ between the occluded point and the surrounding visible points ij Determine the number of visible points n adjacent to the occlusion point, and finally use distance, skeleton and geometric angle constraints to infer the position of the occlusion point, so as to obtain the compensated j-th occlusion point position p j o ′, as shown in equations (1) and (2);
[0013]
[0014] in Represents the visible point p i v and the occlusion point p j o If the Euclidean distance between them is less than the distance threshold d h =0.5m, then p i v are potential nearby visible points, represents the occlusion point p j o With visible point p i v If the angle between h =10°, the geometric constraint between the occluded point and the visible point is considered to be satisfied, indicating that the relative positions and viewing angles of the two are geometrically related. n represents the number of valid visible points near the occluded point, m represents the total number of potential visible points, and I(·) represents the indicator function. If the condition is satisfied, 1 is returned, otherwise 0 is returned. j o represents the position of the jth occlusion point to be compensated, p i v represents the position of the i-th visible point, p i s represents the position of the i-th standard skeleton reference point, w iRepresents the weight coefficient of each visible point, which is inversely proportional to the distance. The closer the visible point is to the occluded point, the greater the influence on the compensation result. λ1 represents the adjustment factor of the skeleton constraint, which is used to control the influence of the skeleton reference point on the objective function. cos(θ) represents the angle constraint, which reflects the geometric relationship between the occluded point and the visible point. γ1 represents the adjustment factor of the angle constraint, which is used to control the intensity of the angle optimization.
[0015] S13: The second stage: occlusion point compensation based on inter-frame consistency; first record the confidence f of the occlusion point i in the current frame t i (t) is greater than the threshold θ, the point changes from the occluded state to the visible state; then the displacement vector v of the point between consecutive frames is calculated o , which is used to describe the displacement of the point in frame t relative to frame tn; finally, the actual position p of the occlusion point is inferred through speed change and confidence correction i o″ (t), as shown in formula (3);
[0016]
[0017] where p i o″ (t) represents the position of the occlusion point i after compensation, p i o ′(tn) represents the position of the occlusion point i in the previous n frames, v i represents the displacement vector between the current frame t and the previous n frames, The correction term for velocity change reflects the change in velocity of the occlusion point, γ a represents the speed adjustment coefficient, which is used to adjust the impact of speed on compensation. Δt represents the time interval between the current frame t and the previous n frames. f i (t) represents the confidence of the occluded point i in the current frame t, γ c Represents the confidence adjustment coefficient. If the confidence is high, it means that the visibility of the occluded point in the current frame is strong.
[0018] Further, the S2 specifically includes the following steps:
[0019] S21: Local reasoning of proprietary features; first, the proprietary features are obtained based on the position, motion information and posture angle information of each key point i; specifically, the two-dimensional spatial coordinates of the key point i provide the position in the image, while the relative motion information describes the motion trajectory of the key point between different frames, and the posture angle information describes the relative rotation between joints; then the proprietary features are input into the deep neural network DNN, and the depth value of the key point is predicted by learning the mapping relationship between the proprietary features of the key point and the depth value
[0020] S22: Global reasoning of shared features: Based on geometric constraints and skeleton anatomical rules, shared features are extracted by obtaining the spatial relative position and topological structure relationship between each key point i and its neighboring point k; then, the shared features are combined with the preliminary depth value. The spatial dependency between key points is captured by the graph neural network GNN to ensure that the depth value meets the proportion and symmetry requirements of the body, thereby obtaining the depth value of each key point i in the global space.
[0021] S23: MTransPose modeling; deep feature fusion is performed through joint modeling of multi-layer perceptron and Transformer to predict accurate 3D coordinates; first, the local depth value obtained by the proprietary feature And the global depth value obtained by sharing features Fusion, through weighted averaging to obtain the final fused depth features; then, the fused depth features are input into the MTransPose network. On the one hand, the MLP network is responsible for local depth fine-tuning and adjusting the depth values that deviate abnormally between key points. On the other hand, the Transformer network processes the spatial relationship between key points through the self-attention mechanism and optimizes the depth prediction of each key point, thereby obtaining the true three-dimensional coordinates of each key point.
[0022] Further, the S3 specifically includes the following steps:
[0023] S31: Construction of human skeleton model; First, the length of the bone segments between the key points is calculated through each human 3D key point obtained in S2 and mapped to each joint of the human standard skeleton model. The human skeleton is an undirected graph G(V,E) consisting of n bone segments and m joints, where V represents the joint set and E represents the bone segment connection set; Then, the motion range of the skeleton model is set using the bone segment length and the joint motion angle, completing the preliminary construction of the human skeleton model;
[0024] S32: Adaptive matching of human skeleton model based on body shape features: first, the body shape features of the human body, such as height H, shoulder width W and waist circumference M, are extracted by capturing the three-dimensional joint positions of the human body, and then the proportions of each part of the skeleton model are dynamically adjusted using proportion calculation, and finally, the position P of the adjusted joint point i is updated in real time by capturing the dynamic body shape changes. i a , as shown in formula (4);
[0025]
[0026] Where P i s is the position of the standard bone joint point, P i(t) is the joint position of the current frame t, α i =(α H +α W +α M )β i is the adjustment scale factor of the bone segment joint point i, and Represent the proportional factors of height, shoulder width and waist circumference, H s , W s 、M s Respectively represent the height, shoulder width and waist circumference characteristics of the standard skeleton model, β i represents the body shape feature adjustment coefficient of joint point i, and m represents the set of all joint points;
[0027] S33: Digital human character binding: First, the joint points in the human skeleton model are paired with the digital human vertices, and the distance d between each joint point i and vertex j is calculated. ij (t) and its relative motion Δd ij (t), set the weight matrix w ij (t) and assign it to each vertex i; then, by dynamically adjusting the rotation and displacement parameters of the skeletal joints and updating the vertex j in real time based on the current joint weights, the real-time mapping of the real person-skeletal model-digital human in different postures of consecutive frames in a multi-person occlusion scene is completed, as shown in formula (5);
[0028]
[0029] where w ij (t) is the weight of the i-th joint point to the j-th vertex, which changes dynamically with time t; d ij (t) is the Euclidean distance between the jth vertex and the i-th joint point, Δd ij (t) is the relative motion between the jth vertex and the ith joint point, indicating the position difference between the two in consecutive frames, R i (t) is the rotation matrix of the i-th joint point, indicating the degree of rotation of the joint at time t; ||R i (t)|| represents the norm of rotation, reflecting the influence of joint rotation on vertex j; v i (t) is the velocity vector of the i-th joint point, indicating the movement speed of the joint at time t; ||v i (t)|| is the norm of the velocity of the i-th joint point, which is used to measure the influence of the velocity of the joint movement on the weight; λ1, λ2 are attenuation factors, which are used to control the attenuation speed of the influence of distance and relative motion on the weight; α1, α2, α3 and β are adjustment factors, which respectively represent the influence of adjusting distance, relative motion, rotation and speed;
[0030] S34: Skeleton multi-rigid body dynamic model; Through the human multi-rigid body dynamics method, the human skeleton is regarded as a number of independent rigid bodies connected to each other by hinges, and the human body is simplified into a multi-rigid body system with limited degrees of freedom; to describe the relationship between force and motion, each rigid body has the following equilibrium equation;
[0031]
[0032] Among them, L i is the dynamic moment of rigid body i relative to the center of mass, m i and I i are the mass and moment of inertia of rigid body i about the simplified center, s i and ω i are the linear acceleration and angular acceleration of the rigid body i, ∑F i and ∑M i are the principal vectors of the force system acting on the rigid body i and the principal moments about the simplified center. The force system includes both external forces and moments, as well as skeletal forces and moments between and within the rigid bodies. Suppose the rigid body i with the center of mass at point P is subjected to a skeletal force F at point O. i p and bone moment At point P, an external force F is applied i e , and is subjected to external torque Then establish the coordinate system at point O. According to the Newton-Euler formula, we have:
[0033]
[0034] Among them, s i is the linear acceleration of the rigid body i, α, β, and γ are the angular accelerations of the rigid body rotating around the x-axis, y-axis, and z-axis respectively, I x ,I y ,I z are the moments of inertia of the rigid body about the centroid around the x-axis, y-axis, and z-axis, respectively. i F is the force F acting on rigid body i i e and the bone force F i p And the additional torque generated by the inertial force about the moment center; therefore, for any rigid body i, there is a matrix form of the dynamic equation:
[0035]
[0036] where I=[I x I y I z ],ω=[α β γ] T ;
[0037] S35: Muscle force line model; First, a muscle force line model is established based on the characteristics of the length and speed changes of muscles during contraction or stretching; the length change and contraction speed of muscles are the key factors affecting muscle force. By using the kinematic information of bones, the length L of the muscle is obtained by calculating the starting point of each muscle in three-dimensional space, thereby obtaining the contraction speed v of the muscle; then, the optimal solution of muscle force is obtained by analyzing the force-length and force-speed relationships of the muscles; the relationship between muscle force and length is nonlinear, and the force-length relationship of each muscle i is expressed by the function f i l (L) indicates that when the muscle contracts faster, the force it produces is smaller. The force-velocity relationship of each muscle i uses a negative exponential decay model f i s (v) is expressed as shown in formula (9);
[0038]
[0039] Among them, f i l (L) represents the force-length function of muscle i. The muscle force is adjusted according to the deviation of the muscle length from the natural length. ξ represents the sensitivity of controlling the change of muscle length. represents the natural length of muscle i, represents the maximum length of muscle i when the force is maximum; f i s (v) represents the force-velocity function of muscle i, which describes the negative correlation between muscle contraction velocity and muscle force, indicating that as the velocity increases, the force generated by the muscle gradually decreases. δ represents the attenuation coefficient that controls the influence of contraction velocity. v i represents the contraction speed of muscle i, ||v i || is the norm of the contraction velocity of muscle i, and exp is used to describe the nonlinear relationship between muscle force and muscle length and velocity;
[0040] Taking into account the effects of force-length and force-speed, the final muscle force F i m It is expressed as shown in formula (10);
[0041]
[0042] in, It is a nonlinear transformation matrix that represents the effect of muscle force in different directions. It is used to adjust the final muscle force vector so that it conforms to the directional constraint in the muscle space. Each element of R represents the factor of muscle force direction adjustment. i maxRepresents the maximum limit of muscle force, W1 and W2 represent weight matrices, which are used to control the intensity of force-length and force-velocity effects. Represents the Hadamard product, which is used for element-wise multiplication;
[0043] S36: Under the joint action of the skeleton model and the muscle model, the new position of the digital human vertex is determined by the joint transformation and mechanics, thereby affecting the deformation of the digital human model to drive the digital human in real time, including the following steps:
[0044] Step 1: Calculate joint torque;
[0045] According to rigid body kinematics, the first step in simulating the movement of joints and skeletal systems is to calculate the torque of 17 joints of the human body and deduce the movement trend and angle change of the joints, thereby affecting the posture of the entire digital human;
[0046] Step 2: Update joint angles and accelerations;
[0047] The change of angle simulates the rotation process of the digital human. In natural movement, the joint angle changes continuously with the contraction or extension of the muscles. The acceleration is updated to obtain the change of joint angle, which in turn promotes the overall change of the skeletal system.
[0048] Step 3: Update the bone rigid body transformation;
[0049] The rigid body change of bones ensures that the skeletal system undergoes corresponding spatial changes as the joint angle and acceleration change. The position and posture of the bones are updated through rigid body transformation to ensure the consistency of the digital human's skeletal structure and joint angle changes.
[0050] Step 4: Bone-muscle mechanical interaction;
[0051] Muscles generate force through contraction, which is transmitted to the attachment points on the bones through tendons, causing the bones to rotate and affecting the changes in joint angles. The mechanical interaction between muscles and bones is simulated through the skeleton rigid body dynamics model and muscle force line model to more realistically restore the skeleton's reaction, thereby driving the digital human to present corresponding action postures.
[0052] Step 5: Update the vertex position of the digital human;
[0053] The update of the vertex position of the digital human is the core step to drive the movement of the digital human. It is achieved through the joint movement of the skeletal system and combined with the muscle force F i m Feedback and bone force F i p The role of the digital human vertex position; first, the vertex update is based on the rigid body transformation matrix T between the bone joints i (θ i ,Li ), the position v of each vertex j It is obtained by interpolating the transformation matrix of the bones; secondly, the bone force F is fed back by the mechanical interaction between the bones. i m Further adjust the position of the vertex, muscle force F i p It will also affect the vertex position through force feedback; during the vertex update process, the bone force and muscle force interact with each other, and they jointly affect the position v′ of the digital human vertex j , as shown in formula (11);
[0054]
[0055] in is the updated vertex position, is the original vertex position, w ij is the influence weight of joint point i on vertex j, T i (θ i ,L i ) is the joint transformation matrix, representing the joint angle θ i and bone length L i The effect of rotation and translation on the bone reflects the geometric properties of the skeleton, ΔT i p is the combined bone force F i m The joint transformation correction matrix caused by i , bone length L i and bone force F i m The mutual influence of i m is the combined muscle force F i p The joint transformation correction matrix caused by i , bone length L i and muscle force F i p the mutual influence of
[0056] Step 6: Vertex - Movement transmission of digital human limbs;
[0057] Based on the update of the digital human's vertex position, the digital human's limbs are finally corrected according to the state of the previous frame, and the action postures of various parts of the digital human's limbs change accordingly, allowing the digital human to perform continuous motion simulation on the timeline, thereby synchronizing the digital human's movements in real time and realizing the process of multi-person motion capture driving the digital human in real time.
[0058] The beneficial effects of the present invention are:
[0059] (1) For the 2D key points occluded by multiple people, the present invention adopts a two-stage occlusion point compensation method. On the one hand, the occlusion points are preliminarily compensated through the human body kinematic constraints and the skeleton structure relationship. Based on the prior knowledge of the human body topology, the positions of the occlusion points are roughly estimated using visible points. On the other hand, based on the principle of motion consistency, the positions of the occlusion points are dynamically corrected considering the displacement and posture changes between adjacent frames, thereby obtaining the 2D key points of all human bodies.
[0060] (2) Compared with the existing skeleton-driven deformation that directly acts on the vertex, the interaction between bones and muscles is usually ignored, and the muscle force, bone force and the mutual influence between the two are not considered. The present invention introduces a bone-muscle interaction model, takes into account the mechanical effect of muscles and the kinematic constraints of bones, and establishes a force transmission mechanism between muscles and bones. The force generated when muscles contract is transmitted to the bones through tendons to drive the movement of the bones. This interactive mechanism makes the deformation of the digital human model not only affected by the bones, but also regulated by muscle force, making the driving process of the digital human more in line with the real movement mechanism of the human body. Moreover, this method makes the movement smoother and more vivid through the layer-by-layer transmission from the vertex to the joints, bones, and muscles. The flow of information between different levels makes the movement of the digital human not only the movement of the skeleton, but a complete dynamic performance.
[0061] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:
[0063] Figure 1 This is a main flow chart of the multi-person motion capture method according to an embodiment of the present invention;
[0064] Figure 2 A diagram of 17 skeletal joints of a human body according to an embodiment of the invention;
[0065] Figure 3 It is an occlusion point compensation map for a multi-person scene according to an embodiment of the invention;
[0066] Figure 4 The "skeleton-muscle" interaction model driving digital human flow chart described in the embodiment of the invention. DETAILED DESCRIPTION
[0067] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0068] Among them, the drawings are only used for illustrative explanations, and they only represent schematic diagrams rather than actual pictures, and should not be understood as limitations on the present invention. In order to better illustrate the embodiments of the present invention, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the size of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0069] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "upper", "lower", "left", "right", "front", "rear", etc. indicate the orientation or position relationship, they are based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation. Therefore, the terms describing the position relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0070] Specific implementation method Figure 1 As shown in the figure, firstly, the 17 key points of the human body are divided into visible points and occluded points according to the human body movement behavior and human body topological structure, then the precise position of the occluded point is obtained through occluded point compensation, and the spatial depth value is further inferred through proprietary and shared features. Finally, a human skeleton model is constructed and a digital human is bound. By establishing a force transmission mechanism between muscles and bones, the force generated by muscle contraction is transmitted to the bones through tendons, driving the movement of the bones to realize a multi-person motion capture real-time driving digital human method based on the bone-muscle interaction model, which specifically includes the following steps:
[0071] Step 1: Use the top-down model YOLO-Pose to detect the position information and global confidence of all 17 skeleton key points of the human body, divide the key points into visible points and occluded points, and then propose a two-stage dynamic occluded point compensation method that integrates visible posture and inter-frame consistency: Stage 1: Preliminary compensation of occluded points through human kinematic constraints and skeleton structure relationships, based on prior knowledge of human topological structure, use visible points to roughly infer the position of occluded points; Stage 2: Based on the principle of motion consistency, consider the displacement and posture changes between adjacent frames, and dynamically correct the position of occluded points, thereby obtaining all two-dimensional key points of the human body;
[0072] 101: Using the YOLO-Pose model, we obtain the key point distribution and global confidence scores of all people from top to bottom for the 17 skeletal joints of the human body, such as Figure 2 As shown. The position P of each key point i It is obtained by maximizing the response value of the key point heat map through the YOLO-Pose model, and the confidence f of each key point i It is determined by the maximum response value in the heat map of the point. If a key point i among the 17 key points of each person is not detected in the current frame or its confidence f i If it is less than the predetermined threshold θ, then the point is an occlusion point p i o ; If the confidence level f detected by the key point i If the value is higher than the threshold θ, the point p is considered visible. i v .
[0073] 102: As Figure 3 As shown in the first stage: compensation of occlusion points based on visible posture. First, the approximate position p of the occlusion point is inferred from the visible posture. j o , and then calculate each potential visible point p i v With the current occlusion point p j o Distance d ij To determine whether it is a neighboring point, further use the angle relationship θ between the occluded point and the surrounding visible points ij Determine the number of visible points n adjacent to the occlusion point, and finally use distance, skeleton and geometric angle constraints to infer the position of the occlusion point, so as to obtain the compensated j-th occlusion point position p j o ′, as shown in equations (1) and (2).
[0074]
[0075] in Represents the visible point p iv and the occlusion point p j o If the Euclidean distance between them is less than the distance threshold d h =0.5m, then p i v are potential nearby visible points, represents the occlusion point p j o With visible point p i v If the angle between h =10°, the geometric constraint between the occluded point and the visible point is considered to be satisfied, indicating that the relative positions and viewing angles of the two are geometrically related. n represents the number of valid visible points near the occluded point, m represents the total number of potential visible points, and I(·) represents the indicator function. If the condition is satisfied, 1 is returned, otherwise 0 is returned. j o represents the position of the jth occlusion point to be compensated, p i v represents the position of the i-th visible point, p i s represents the position of the i-th standard skeleton reference point, w i Represents the weight coefficient of each visible point, which is inversely proportional to the distance. The closer the visible point is to the occluded point, the greater the influence on the compensation result. λ1 represents the adjustment factor of the skeleton constraint, which is used to control the influence of the skeleton reference point on the objective function. cos(θ) represents the angle constraint, which reflects the geometric relationship between the occluded point and the visible point. γ1 represents the adjustment factor of the angle constraint, which is used to control the intensity of the angle optimization.
[0076] 103: The second stage: occlusion point compensation based on inter-frame consistency. First, record the confidence f of the occlusion point i in the current frame t. i (t) is greater than the threshold θ, the point changes from the occluded state to the visible state. Then calculate the displacement vector v of the point between consecutive frames o , which is used to describe the displacement of the point in frame t relative to frame tn. Finally, the actual position p of the occlusion point is inferred through speed change and confidence correction. i o″ (t), as shown in formula (3).
[0077]
[0078] where p i o″ (t) represents the position of the occlusion point i after compensation, p i o ′(tn) represents the position of the occlusion point i in the previous n frames, v irepresents the displacement vector between the current frame t and the previous n frames, The correction term for velocity change reflects the change in velocity of the occlusion point, γ a represents the speed adjustment coefficient, which is used to adjust the impact of speed on compensation. Δt represents the time interval between the current frame t and the previous n frames. f i (t) represents the confidence of the occluded point i in the current frame t, γ c Represents the confidence adjustment coefficient. If the confidence is high, it means that the visibility of the occluded point in the current frame is strong.
[0079] Step 2: Propose a 3D keypoint depth extractor based on multi-dimensional proprietary-shared feature reasoning. First, for each 2D keypoint in S1, preliminary depth information (z coordinate) is inferred locally based on proprietary features. Then, based on shared features, the depth information of each keypoint is adjusted using global spatial constraints to infer the relative positions of multiple keypoints in 3D space, avoiding inconsistency of the overall 3D model due to local errors of a single keypoint. Finally, the MTransPose model is used to jointly model proprietary and shared features and output the 3D coordinates of the keypoints of the human body.
[0080] 201: Local reasoning of proprietary features. First, the proprietary features are obtained based on the position, motion information and posture angle information of each key point i. Specifically, the two-dimensional spatial coordinates of the key point i provide the position in the image, while the relative motion information describes the motion trajectory of the key point between different frames, and the posture angle information describes the relative rotation between joints. The proprietary features are then input into the deep neural network DNN to predict the depth value of the key point by learning the mapping relationship between the proprietary features of the key point and the depth value.
[0081] 202: Global reasoning of shared features. Based on geometric constraints and skeleton anatomical rules, shared features are extracted by obtaining the spatial relative position and topological structure relationship between each key point i and its neighboring point k. Then, the shared features are combined with the preliminary depth value. The spatial dependency between key points is captured by the graph neural network GNN to ensure that the depth value meets the proportion and symmetry requirements of the body, thereby obtaining the depth value of each key point i in the global space.
[0082] 203: MTransPose modeling. Through the joint modeling of multi-layer perceptron and Transformer, deep feature fusion is used to predict accurate 3D coordinates. First, the local depth value obtained by the proprietary feature is And the global depth value obtained by sharing features Fusion, through weighted averaging to obtain the final fused depth features. Then, the fused depth features are input into the MTransPose network. On the one hand, the MLP network is responsible for local depth fine-tuning and adjusting the depth values that deviate abnormally between key points. On the other hand, the Transformer network processes the spatial relationship between key points through the self-attention mechanism and optimizes the depth prediction of each key point, thereby obtaining the true three-dimensional coordinates of each key point.
[0083] Step 3: Propose a real-time mapping and driving method of "real person-bone model-digital human". First, map the three-dimensional key points of the human body obtained in S2 to the joint positions of the standard human skeleton model to ensure that the various parts of the skeleton are correctly connected through the joints, establish a complete skeleton structure, and design an adaptive skeleton model that can adapt to different heights, body shapes and skeleton proportions based on the standard proportions of the human body. Then, through the weight drawing algorithm, connect each vertex of the digital human to the joints in the skeleton model to bind the digital human character. Finally, a digital human driving method based on the skeleton-muscle interaction model is proposed. The skeleton force is obtained through the skeleton rigid body dynamics model, and the muscle force is obtained using the muscle force line model. Then, combined with the joint movement of the skeletal system, based on the update of the vertex position of the digital human, the digital human is transmitted from the vertex to the joints, bones, and muscles layer by layer, so that the digital human accurately reflects the real movement posture after receiving the driving signal, thereby realizing the process of multi-person motion capture and real-time driving of the digital human.
[0084] 301: Construction of human skeleton model. First, the length of the bone segments between the key points is calculated through each human 3D key point obtained in S2 and mapped to each joint of the human standard skeleton model. The human skeleton is an undirected graph G(V,E) consisting of n bone segments and m joints, where V represents the joint set and E represents the bone segment connection set. Then, the motion range of the skeleton model is set using the bone segment length and joint motion angle to complete the preliminary construction of the human skeleton model.
[0085] 302: Adaptive matching of human skeleton model based on body shape features. First, the body shape features of the human body, such as height H, shoulder width W and waist circumference M, are extracted by capturing the three-dimensional joint positions of the human body. Then, the proportions of each part of the skeleton model are dynamically adjusted using proportion calculation. Finally, the position P of the adjusted joint point i is updated in real time by capturing the dynamic body shape changes. i a , as shown in formula (4).
[0086]
[0087] Where P i s is the position of the standard bone joint point, P i (t) is the joint position of the current frame t, αi =(α H +α W +α M )β i is the adjustment scale factor of the bone segment joint point i, and Represent the proportional factors of height, shoulder width and waist circumference, H s , W s 、M s Respectively represent the height, shoulder width and waist circumference characteristics of the standard skeleton model, β i represents the body shape feature adjustment coefficient of joint point i, and m represents the set of all joint points.
[0088] 303: Digital human character binding. First, the joint points in the human skeleton model are paired with the digital human vertices. By calculating the distance d between each joint point i and vertex j ij (t) and its relative motion Δd ij (t), set the weight matrix w ij (t) and assigned to each vertex i. Then, by dynamically adjusting the rotation and displacement parameters of the skeletal joints and updating the vertex j in real time based on the current joint weights, the real-time mapping of the real person-skeletal model-digital human in different postures of consecutive frames in a multi-person occlusion scene is completed, as shown in formula (5).
[0089]
[0090] where w ij (t) is the weight of the i-th joint point to the j-th vertex, which changes dynamically with time t. ij (t) is the Euclidean distance between the jth vertex and the i-th joint point, Δd ij (t) is the relative motion between the jth vertex and the ith joint point, indicating the position difference between the two in consecutive frames, R i (t) is the rotation matrix of the i-th joint point, indicating the degree of rotation of the joint at time t. ||R i (t)|| represents the norm of the rotation, reflecting the influence of the joint rotation on vertex j. i (t) is the velocity vector of the i-th joint point, indicating the movement speed of the joint at time t. ||v i (t)|| is the norm of the velocity of the i-th joint point, which is used to measure the influence of the speed of joint movement on the weight. λ1 and λ2 are attenuation factors, which are used to control the attenuation speed of the influence of distance and relative movement on the weight. α1, α2, α3 and β are adjustment factors, which respectively represent the influence of adjusting distance, relative movement, rotation and speed.
[0091] 304: Skeleton multi-rigid body dynamic model. Through the human multi-rigid body dynamics method, the human skeleton is regarded as a number of independent rigid bodies connected to each other by hinges, that is, the human body is simplified into a multi-rigid body system with limited degrees of freedom. To describe the relationship between force and motion, each rigid body has the following equilibrium equation.
[0092]
[0093] Among them, L i is the dynamic moment of rigid body i relative to the center of mass, m i and I i are the mass and moment of inertia of rigid body i about the simplified center, s i and ω i are the linear acceleration and angular acceleration of the rigid body i, ∑F i and ∑M i are the principal vectors of the force system acting on the rigid body i and the principal moments about the simplified center. The force system includes both external forces and moments, as well as skeletal forces and moments between and within the rigid bodies, such as Figure 4 As shown, suppose the rigid body i with the center of mass at point P is subjected to the bone force F at point O. i p and bone moment At point P, an external force F is applied i e , and is subjected to external torque Then establish the coordinate system at point O. According to the Newton-Euler formula, we have:
[0094]
[0095] Among them, s i is the linear acceleration of the rigid body i, α, β, and γ are the angular accelerations of the rigid body rotating around the x-axis, y-axis, and z-axis respectively, I x ,I y ,I z are the moments of inertia of the rigid body about the centroid around the x-axis, y-axis, and z-axis, respectively. i F is the force F acting on rigid body i i e and the bone force F i p And the additional torque generated by the inertial force about the moment center. Therefore, for any rigid body i, there is a matrix form of dynamic equations:
[0096]
[0097] where I=[I x I y I z ],ω=[αβγ]T .
[0098] 305: Muscle force line model. First, a muscle force line model is established based on the characteristics of the length and speed changes of muscles during contraction or stretching. The length change and contraction speed of muscles are the key factors affecting muscle force. By using the kinematic information of bones, the length L of the muscle is obtained by calculating the starting point of each muscle in three-dimensional space, and then the contraction speed v of the muscle is obtained. Then, the optimal solution for muscle force is obtained by analyzing the force-length and force-speed relationships of the muscles. The relationship between muscle force and length is usually nonlinear. When the length of the muscle deviates from the optimal range (usually 90%-110% of the natural length of the muscle), the degree of stretching of the muscle is inappropriate, and the tension of the muscle will decrease, resulting in the force generated being lower than normal, and the force generated is about 50%-70% of the maximum force. This reduction in force makes it impossible for the muscle to generate maximum force. The force-length relationship of each muscle i is expressed by the function f i l (L) indicates that when the muscle contracts faster, the force it produces is smaller. The force-velocity relationship of each muscle i can be expressed using the negative exponential decay model f i s (v) is expressed as shown in formula (9).
[0099]
[0100] Among them, f i l (L) represents the force-length function of muscle i. The muscle force is adjusted according to the deviation of the muscle length from the natural length. ξ represents the sensitivity of controlling the change of muscle length. L i o represents the natural length of muscle i, L i max represents the maximum length of muscle i when the force is maximum. i s (v) represents the force-velocity function of muscle i, which describes the negative correlation between muscle contraction velocity and muscle force, indicating that as the velocity increases, the force generated by the muscle gradually decreases. δ represents the attenuation coefficient that controls the influence of contraction velocity. v i represents the contraction speed of muscle i, ||v i || is the norm of the contraction velocity of muscle i, and exp is used to describe the nonlinear relationship between muscle force and muscle length and velocity.
[0101] Taking into account the effects of force-length and force-speed, the final muscle force F i m It can be expressed as shown in formula (10).
[0102]
[0103] in, It is a nonlinear transformation matrix that represents the effect of muscle force in different directions. It is used to adjust the final muscle force vector so that it conforms to the directional constraint in the muscle space. Each element of R represents the factor of muscle force direction adjustment. i max Represents the maximum limit of muscle force, W1 and W2 represent weight matrices, which are used to control the intensity of force-length and force-velocity effects. Represents the Hadamard product, which is used for element-wise multiplication.
[0104] 306: Under the joint action of the skeleton model and the muscle model, the new position of the digital human vertex is determined by the joint transformation and mechanics, thereby affecting the deformation of the digital human model to drive the digital human in real time, including the following steps:
[0105] Step 1: Calculate joint torque.
[0106] According to rigid body kinematics, the first step in simulating the movement of joints and skeletal systems is to calculate the torque of the 17 joints of the human body and deduce the movement trend and angle change of the joints, thereby affecting the posture of the entire digital human.
[0107] Step 2: Update joint angles and accelerations.
[0108] The change in angle simulates the rotation process of the digital human. In natural movement, the joint angle changes continuously with the contraction or extension of the muscles. Updating the acceleration can derive the change in joint angle, thereby promoting the overall change of the skeletal system.
[0109] Step 3: Update the bone rigid body transformation.
[0110] The rigid body transformation of bones ensures that the skeletal system changes in space accordingly with the changes in joint angles and acceleration. The position and posture of bones are updated through rigid body transformation to ensure the consistency of the digital human's skeletal structure and joint angle changes.
[0111] Step 4: Bone-muscle mechanical interaction.
[0112] Muscles generate force through contraction, which is transmitted to the attachment points on the bones through tendons, causing the bones to rotate and affecting the changes in joint angles. The mechanical interaction between muscles and bones is simulated through the skeleton rigid body dynamics model and muscle force line model to more realistically restore the skeleton's reaction, thereby driving the digital human to present corresponding action postures.
[0113] Step 5: Update the vertex position of the digital human.
[0114] The update of the vertex position of the digital human is the core step to drive the movement of the digital human. It is achieved through the joint movement of the skeletal system and combined with the muscle force F i m Feedback and bone force F i p The role of the digital human vertex position is to affect the vertex position of the digital human. First, the vertex update is based on the rigid body transformation matrix T between the bone joints. i (θ i ,L i ), the position v of each vertex j It is obtained by interpolating the transformation matrix of the bones. Secondly, the bone force F is fed back by the mechanical interaction between the bones. i m Further adjust the position of the vertex, muscle force F i p It will also affect the vertex position through force feedback. During the vertex update process, the bone force and muscle force interact with each other, and they jointly affect the position v′ of the digital human vertex. j , as shown in formula (11).
[0115]
[0116] in is the updated vertex position, is the original vertex position, w ij is the influence weight of joint point i on vertex j, T i (θ i ,L i ) is the joint transformation matrix, representing the joint angle θ i and bone length L i The effect of rotation and translation on the bone reflects the geometric properties of the skeleton, ΔT i p is the combined bone force F i m The joint transformation correction matrix caused by i , bone length L i and bone force F i m The mutual influence of i m is the combined muscle force F i p The joint transformation correction matrix caused by i , bone length L i and muscle force F i p mutual influence.
[0117] Step 6: Vertex - Movement transmission of digital human limbs.
[0118] Based on the update of the digital human's vertex position, the digital human's limbs are finally corrected according to the state of the previous frame, and the action postures of various parts of the digital human's limbs change accordingly, allowing the digital human to perform continuous motion simulation on the timeline, thereby synchronizing the digital human's movements in real time and realizing the process of multi-person motion capture driving the digital human in real time.
[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.
Claims
1. A method for real-time driving of digital human using multi-person motion capture based on a skeleton-muscle interaction model, characterized by: The method comprises the following steps: S1: Use YOLO-Pose to detect the position information and global confidence of all 17 skeleton key points of the human body, divide the key points into visible points and occluded points, and use a two-stage dynamic occluded point compensation method that integrates visible poses and inter-frame consistency to obtain the position of the occluded points; S2: A 3D keypoint deep extractor using multi-dimensional proprietary-shared feature reasoning combined with the MTransPose model to output the 3D coordinates of the key points of the human body; S3: Map the 3D key points of the human body obtained in S2 to the standard human skeleton model, bind the digital human character to the skeleton model through the weight mapping algorithm, and based on the bone-muscle interaction model, through the mechanical effects of the bone rigid body dynamics model and the muscle force line model, make the digital human accurately reflect the real movement posture, and realize the process of multi-person motion capture and real-time driving of the digital human.
2. The method for real-time driving of digital human by multi-person motion capture based on the skeleton-muscle interaction model according to claim 1 is characterized in that: The S1 specifically includes the following steps: S11: Using the YOLO-Pose model, we obtain the key point distribution and global confidence scores of all people from top to bottom for the 17 skeletal joints of the human body; the position P of each key point i It is obtained by maximizing the response value of the key point heat map through the YOLO-Pose model, and the confidence f of each key point i It is determined by the maximum response value in the heat map of the point; if a key point i among the 17 key points of each person is not detected in the current frame or its confidence f i If it is less than the predetermined threshold θ, then the point is an occlusion point p i o ; If the confidence f detected by the key point i If the value is higher than the threshold θ, the point p is considered visible. i v ; S12: The first stage: compensation of occlusion points based on visible postures; first, the approximate position p of the occlusion point is inferred from the visible postures j o , and then calculate each potential visible point p i v With the current occlusion point p j o Distance d ij To determine whether it is a neighboring point, further use the angle relationship θ between the occluded point and the surrounding visible points ij Determine the number of visible points n adjacent to the occlusion point, and finally use distance, skeleton and geometric angle constraints to infer the position of the occlusion point, so as to obtain the compensated j-th occlusion point position p j o ′, as shown in equations (1) and (2); in Represents the visible point p i v and the occlusion point p j o If the Euclidean distance between them is less than the distance threshold d h =0.5m, then p i v are potential nearby visible points, represents the occlusion point p j o With visible point p i v If the angle between h =10°, the geometric constraint between the occluded point and the visible point is considered to be satisfied, indicating that the relative positions and viewing angles of the two are geometrically related. n represents the number of valid visible points near the occluded point, m represents the total number of potential visible points, and I(·) represents the indicator function. If the condition is satisfied, 1 is returned, otherwise 0 is returned. j o represents the position of the jth occlusion point to be compensated, p i v represents the position of the i-th visible point, p i s represents the position of the i-th standard skeleton reference point, w i Represents the weight coefficient of each visible point, which is inversely proportional to the distance. The closer the visible point is to the occluded point, the greater the influence on the compensation result. λ1 represents the adjustment factor of the skeleton constraint, which is used to control the influence of the skeleton reference point on the objective function. cos(θ) represents the angle constraint, which reflects the geometric relationship between the occluded point and the visible point. γ1 represents the adjustment factor of the angle constraint, which is used to control the intensity of the angle optimization. S13: The second stage: occlusion point compensation based on inter-frame consistency; first record the confidence f of the occlusion point i in the current frame t i (t) is greater than the threshold θ, the point changes from the occluded state to the visible state; then the displacement vector v of the point between consecutive frames is calculated o , which is used to describe the displacement of the point in frame t relative to frame tn; finally, the actual position p of the occlusion point is inferred through speed change and confidence correction i o″ (t), as shown in formula (3); where p i o″ (t) represents the position of the occlusion point i after compensation, p i o′ (tn) represents the position of the occlusion point i in the previous n frames, v i represents the displacement vector between the current frame t and the previous n frames, The correction term for velocity change reflects the change in velocity of the occlusion point, γ a represents the speed adjustment coefficient, which is used to adjust the impact of speed on compensation. Δt represents the time interval between the current frame t and the previous n frames. f i (t) represents the confidence of the occluded point i in the current frame t, γ c Represents the confidence adjustment coefficient. If the confidence is high, it means that the visibility of the occluded point in the current frame is strong.
3. The method for real-time driving of digital human by multi-person motion capture based on the skeleton-muscle interaction model according to claim 2 is characterized by: The S2 specifically includes the following steps: S21: Local reasoning of proprietary features; first, the proprietary features are obtained based on the position, motion information and posture angle information of each key point i; specifically, the two-dimensional spatial coordinates of the key point i provide the position in the image, while the relative motion information describes the motion trajectory of the key point between different frames, and the posture angle information describes the relative rotation between joints; then the proprietary features are input into the deep neural network DNN, and the depth value of the key point is predicted by learning the mapping relationship between the proprietary features of the key point and the depth value S22: Global reasoning of shared features: Based on geometric constraints and skeleton anatomical rules, shared features are extracted by obtaining the spatial relative position and topological structure relationship between each key point i and its neighboring point k; then, the shared features are combined with the preliminary depth value. The spatial dependency between key points is captured by the graph neural network GNN to ensure that the depth value meets the proportion and symmetry requirements of the body, thereby obtaining the depth value of each key point i in the global space. S23: MTransPose modeling; deep feature fusion is performed through joint modeling of multi-layer perceptron and Transformer to predict accurate 3D coordinates; first, the local depth value obtained by the proprietary feature And the global depth value obtained by sharing features Fusion, through weighted averaging to obtain the final fused depth features; then, the fused depth features are input into the MTransPose network. On the one hand, the MLP network is responsible for local depth fine-tuning and adjusting the depth values that deviate abnormally between key points. On the other hand, the Transformer network processes the spatial relationship between key points through the self-attention mechanism and optimizes the depth prediction of each key point, thereby obtaining the true three-dimensional coordinates of each key point.
4. The method for real-time driving of digital human by multi-person motion capture based on the skeleton-muscle interaction model according to claim 3 is characterized in that: The S3 specifically includes the following steps: S31: Construction of human skeleton model; First, the length of the bone segments between the key points is calculated through each human 3D key point obtained in S2 and mapped to each joint of the human standard skeleton model. The human skeleton is an undirected graph G(V,E) consisting of n bone segments and m joints, where V represents the joint set and E represents the bone segment connection set; Then, the motion range of the skeleton model is set using the bone segment length and the joint motion angle, completing the preliminary construction of the human skeleton model; S32: Adaptive matching of human skeleton model based on body shape features: first, the body shape features of the human body, such as height H, shoulder width W and waist circumference M, are extracted by capturing the three-dimensional joint positions of the human body, and then the proportions of each part of the skeleton model are dynamically adjusted using proportion calculation, and finally, the position P of the adjusted joint point i is updated in real time by capturing the dynamic body shape changes. i a , as shown in formula (4); Where P i s is the position of the standard bone joint point, P i (t) is the joint position of the current frame t, α i =(α H +α W +α M )β i is the adjustment scale factor of the bone segment joint point i, and Represent the proportional factors of height, shoulder width and waist circumference, H s , W s 、M s Respectively represent the height, shoulder width and waist circumference characteristics of the standard skeleton model, β i represents the body shape feature adjustment coefficient of joint point i, and m represents the set of all joint points; S33: Digital human character binding: First, the joint points in the human skeleton model are paired with the digital human vertices, and the distance d between each joint point i and vertex j is calculated. ij (t) and its relative motion Δd ij (t), set the weight matrix w ij (t) and assign it to each vertex i; then, by dynamically adjusting the rotation and displacement parameters of the skeletal joints and updating the vertex j in real time based on the current joint weights, the real-time mapping of the real person-skeletal model-digital human in different postures of consecutive frames in a multi-person occlusion scene is completed, as shown in formula (5); where w ij (t) is the weight of the i-th joint point to the j-th vertex, which changes dynamically with time t; d ij (t) is the Euclidean distance between the jth vertex and the i-th joint point, Δd ij (t) is the relative motion between the jth vertex and the ith joint point, indicating the position difference between the two in consecutive frames, R i (t) is the rotation matrix of the i-th joint point, indicating the degree of rotation of the joint at time t; ||R i (t)|| represents the norm of rotation, reflecting the influence of joint rotation on vertex j; v i (t) is the velocity vector of the i-th joint point, indicating the movement speed of the joint at time t; ||v i (t)|| is the norm of the velocity of the i-th joint point, which is used to measure the influence of the velocity of the joint movement on the weight; λ1, λ2 are attenuation factors, which are used to control the attenuation speed of the influence of distance and relative motion on the weight; α1, α2, α3 and β are adjustment factors, which respectively represent the influence of adjusting distance, relative motion, rotation and speed; S34: Skeleton multi-rigid body dynamic model; Through the human multi-rigid body dynamics method, the human skeleton is regarded as a number of independent rigid bodies connected to each other by hinges, and the human body is simplified into a multi-rigid body system with limited degrees of freedom; to describe the relationship between force and motion, each rigid body has the following equilibrium equation; Among them, L i is the dynamic moment of rigid body i relative to the center of mass, m i and I i are the mass and moment of inertia of rigid body i about the simplified center, s i and ω i are the linear acceleration and angular acceleration of the rigid body i, ∑F i and ∑M i are the principal vectors of the force system acting on the rigid body i and the principal moments about the simplified center. The force system includes both external forces and moments, as well as skeletal forces and moments between and within the rigid bodies. Suppose the rigid body i with the center of mass at point P is subjected to a skeletal force F at point O. i p and bone moment At point P, an external force F is applied i e , and is subjected to external torque Then establish the coordinate system at point O. According to the Newton-Euler formula, we have: Among them, s i is the linear acceleration of the rigid body i, α, β, and γ are the angular accelerations of the rigid body rotating around the x-axis, y-axis, and z-axis respectively, I x ,I y ,I z are the moments of inertia of the rigid body about the centroid around the x-axis, y-axis, and z-axis, respectively, i F is the force F acting on rigid body i i e and the bone force F i p And the additional torque generated by the inertial force about the moment center; therefore, for any rigid body i, there is a matrix form of the dynamic equation: where I = [I x I y I z , ω = [αβγ] T ; S35: Muscle force line model; First, a muscle force line model is established based on the characteristics of the length and speed changes of muscles during contraction or stretching; the length change and contraction speed of muscles are the key factors affecting muscle force. By using the kinematic information of bones, the length L of the muscle is obtained by calculating the starting point of each muscle in three-dimensional space, thereby obtaining the contraction speed v of the muscle; then, the optimal solution of muscle force is obtained by analyzing the force-length and force-speed relationships of the muscles; the relationship between muscle force and length is nonlinear, and the force-length relationship of each muscle i is expressed by the function f i l (L) indicates that when the muscle contracts faster, the force it produces is smaller. The force-velocity relationship of each muscle i uses a negative exponential decay model f i s (v) is expressed as shown in formula (9); Among them, f i l (L) represents the force-length function of muscle i. The muscle force is adjusted according to the deviation of the muscle length from the natural length. ξ represents the sensitivity of controlling the change of muscle length. represents the natural length of muscle i, represents the maximum length of muscle i when the force is maximum; f i s (v) represents the force-velocity function of muscle i, which describes the negative correlation between muscle contraction velocity and muscle force, indicating that as the velocity increases, the force generated by the muscle gradually decreases. δ represents the attenuation coefficient that controls the influence of contraction velocity. v i represents the contraction speed of muscle i, ||v i || is the norm of the contraction velocity of muscle i, and exp is used to describe the nonlinear relationship between muscle force and muscle length and velocity; Taking into account the effects of force-length and force-speed, the final muscle force F i m It is expressed as shown in formula (10); in, It is a nonlinear transformation matrix that represents the effect of muscle force in different directions. It is used to adjust the final muscle force vector so that it conforms to the directional constraint in the muscle space. Each element of R represents the factor of muscle force direction adjustment. i max Represents the maximum limit of muscle force, W1 and W2 represent weight matrices, which are used to control the intensity of force-length and force-velocity effects. Represents the Hadamard product, which is used for element-wise multiplication; S36: Under the joint action of the skeleton model and the muscle model, the new position of the digital human vertex is determined by the joint transformation and mechanics, thereby affecting the deformation of the digital human model to drive the digital human in real time, including the following steps: Step 1: Calculate joint torque; According to rigid body kinematics, the first step in simulating the movement of joints and skeletal systems is to calculate the torque of 17 joints of the human body and deduce the movement trend and angle change of the joints, thereby affecting the posture of the entire digital human; Step 2: Update joint angles and accelerations; The change of angle simulates the rotation process of the digital human. In natural movement, the joint angle changes continuously with the contraction or extension of the muscles. The acceleration is updated to obtain the change of joint angle, which in turn promotes the overall change of the skeletal system. Step 3: Update the bone rigid body transformation; The rigid body change of bones ensures that the skeletal system undergoes corresponding spatial changes as the joint angle and acceleration change. The position and posture of the bones are updated through rigid body transformation to ensure the consistency of the digital human's skeletal structure and joint angle changes. Step 4: Bone-muscle mechanical interaction; Muscles generate force through contraction, which is transmitted to the attachment points on the bones through tendons, causing the bones to rotate and affecting the changes in joint angles. The mechanical interaction between muscles and bones is simulated through the skeleton rigid body dynamics model and muscle force line model to more realistically restore the skeleton's reaction, thereby driving the digital human to present corresponding action postures. Step 5: Update the vertex position of the digital human; The update of the vertex position of the digital human is the core step to drive the movement of the digital human. It is achieved through the joint movement of the skeletal system and combined with the muscle force F i m Feedback and bone force F i p The role of the digital human vertex position; first, the vertex update is based on the rigid body transformation matrix T between the bone joints i (θ i ,L i ), the position v of each vertex j It is obtained by interpolating the transformation matrix of the bones; secondly, the bone force F is fed back by the mechanical interaction between the bones. i m Further adjust the position of the vertex, muscle force F i p It will also affect the vertex position through force feedback; during the vertex update process, the bone force and muscle force interact with each other, and they jointly affect the position v′ of the digital human vertex j , as shown in formula (11); in is the updated vertex position, is the original vertex position, w ij is the influence weight of joint point i on vertex j, T i (θ i ,L i ) is the joint transformation matrix, representing the joint angle θ i and bone length L i The effect of rotation and translation on the bone reflects the geometric properties of the skeleton, ΔT i p is the combined bone force F i m The joint transformation correction matrix caused by i , bone length L i and bone force F i m The mutual influence of i m is the combined muscle force F i p The joint transformation correction matrix caused by i , bone length L i and muscle force F i p the mutual influence of Step 6: Vertex - Movement transmission of digital human limbs; Based on the update of the digital human's vertex position, the digital human's limbs are finally corrected according to the state of the previous frame, and the action postures of various parts of the digital human's limbs change accordingly, allowing the digital human to perform continuous motion simulation on the timeline, thereby synchronizing the digital human's movements in real time and realizing the process of multi-person motion capture driving the digital human in real time.
Citation Information
Cited By
Mark-free action acquisition and analysis method and system
CN120766338A
A markerless motion acquisition and analysis method and system
CN120766338B
Unmarked kinetic analysis method and system for real-time tracking shooting of unmanned aerial vehicle
CN120766339A
A markerless dynamics analysis method and system for real-time tracking and photographing of a drone
CN120766339B
Key frame rapid correction method and system based on action recognition
CN121033925A