Virtual avatar driving method, apparatus, device and readable storage medium
By using real-time motion capture and multi-network prediction technology, the problem of synchronizing virtual avatars with human body movements has been solved, achieving accurate synchronization and efficient driving of virtual avatars and human body movements.
Patent Information
- Application Number
- CN202211578602.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-06
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2042-12-06
AI Technical Summary
In existing technologies, virtual images and human movements cannot be accurately synchronized, resulting in high prediction difficulty. Existing technologies also face synchronization challenges when driving virtual images.
Images are acquired through real-time motion capture. Joint prediction networks, limb prediction networks, and hand prediction networks are used to predict the position and posture data of human joints, respectively. Hand images are cropped for accurate prediction to drive the virtual avatar.
It achieves accurate synchronization between virtual avatars and human body movements, simplifies the prediction process, improves prediction efficiency and matching accuracy, reduces prediction complexity, and enhances the real-time performance and reliability of the drive.
Smart Images

Figure CN115798050B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and more specifically, to a virtual avatar driving method, apparatus, device, and readable storage medium. Background Technology
[0002] Virtual avatars are a technology that connects information from the real world and the virtual world. Virtual avatar synchronization refers to the ability of a virtual avatar to synchronize with the movements, gestures, and limbs of a real human body. Through virtual avatars, people can conceal and protect their real-world information, and through virtual avatar synchronization, they can experience events impossible in the real world. However, current technologies for virtual avatar synchronization face challenges in predicting the entire human body as a whole, leading to inaccurate synchronization issues. Summary of the Invention
[0003] In view of this, this application provides a virtual avatar driving method, apparatus, device and readable storage medium to solve the shortcomings of the prior art in that virtual avatars and human body movements cannot be accurately synchronized.
[0004] To achieve the above objectives, the following solution is proposed:
[0005] A virtual avatar driving method, comprising:
[0006] Real-time motion capture of the human body driving the virtual avatar is performed to obtain captured images;
[0007] Obtain the joint prediction network and limb prediction network;
[0008] The captured image is input into the joint prediction network to obtain the position data of multiple human joints output by the joint prediction network and the joint confidence level corresponding to each human joint.
[0009] The position data of each human joint is input into the limb prediction network to obtain the posture data corresponding to each body joint output by the limb prediction network.
[0010] If the location data of each human joint includes the location data of the hand joint and the confidence of the joint corresponding to the hand joint is greater than a preset threshold, obtain the hand prediction network.
[0011] The captured image is cropped to obtain a hand capture image;
[0012] The hand capture image is input into the hand prediction network to obtain the pose data corresponding to multiple hand joints output by the hand prediction network;
[0013] The virtual avatar is driven based on the position data of each human body joint, the posture data corresponding to each torso joint, and the posture data corresponding to each hand joint.
[0014] Optionally, the acquisition of the keypoint prediction network includes:
[0015] An initial joint prediction network and multiple human images are obtained, with the location data of human joints marked on each human image.
[0016] Based on the location data of the human joints marked on each human body image, draw the thermal image of the joints corresponding to the human body image;
[0017] Each human body image and its corresponding joint point thermal image are used as one of the training data for training the initial joint point prediction network, resulting in multiple training data.
[0018] The initial joint prediction network is trained using the multiple training data, and the initial joint prediction network obtained after training is used as the latest joint prediction network.
[0019] Optionally, the limb prediction network includes:
[0020] Obtain the initial limb prediction network;
[0021] Set the pose and shape parameters of the 3D human body model to obtain the position data of multiple body joints and the pose data of multiple body joints corresponding to different parameters.
[0022] The position data of each body joint corresponding to each parameter and the posture data of each body joint are used as one of the training data for training the initial limb prediction network, resulting in multiple training data.
[0023] The initial limb prediction network is trained using multiple training data sets, and the trained initial limb prediction network is used as the latest limb prediction network.
[0024] Optionally, the process of obtaining the hand prediction network includes:
[0025] Obtain the initial hand prediction network;
[0026] The initial hand prediction network is trained using a mixed data approach, and the trained initial hand prediction network is used as the latest hand prediction network.
[0027] Optionally, the initial hand prediction network can be trained using a mixed data approach, including:
[0028] Multiple three-dimensional hand data are acquired, wherein each three-dimensional hand data is generated based on multiple two-dimensional images of a real human body in the same pose, and each three-dimensional hand data contains a hand image with two-dimensional position data of multiple hand joints labeled, three-dimensional position data of multiple hand joints, and pose data corresponding to multiple hand joints.
[0029] Multiple two-dimensional hand data are acquired, wherein each two-dimensional hand data is generated based on a real two-dimensional image, and each two-dimensional hand data includes a hand image with two-dimensional position data and posture data of multiple hand joints labeled.
[0030] Multiple simulation data are acquired. Each simulation data is obtained by setting the posture parameters and morphological parameters of a three-dimensional hand model. Each simulation data includes a hand image with two-dimensional position data of multiple hand joints labeled, three-dimensional position data of multiple hand joints, and posture data corresponding to multiple hand joints.
[0031] The training set is composed of multiple 3D hand data, multiple 2D hand data, and multiple simulation data.
[0032] The initial hand prediction network is trained using the training set.
[0033] Optionally, the captured image is cropped to obtain a hand capture image, including:
[0034] Based on the position data of multiple human joints output by the joint prediction network, a rectangular region of the face in the captured image is determined, and the rectangular region of the face contains the face image of the human body.
[0035] The rectangular area of the face is enlarged according to a preset ratio to obtain a rectangular frame;
[0036] Based on the position data of the hand joints, the rectangular frame is moved in the captured image to obtain a rectangular area of the hand;
[0037] A rectangular area of the hand is cropped from the captured image, and the cropped rectangular area of the hand is used as the captured image of the hand.
[0038] Optionally, driving the virtual avatar based on the position data of each human body joint, the posture data corresponding to each torso joint, and the posture data corresponding to each hand joint includes:
[0039] Determine the root node position data of the human body image in the captured image, and draw the motion chain corresponding to the human hand with the root node position data as the starting point. The motion chain contains multiple target body joints.
[0040] The posture data corresponding to the wrist joint is determined from the posture data of each posture data, and the posture data corresponding to each target body joint is determined from the posture data corresponding to each body joint.
[0041] The posture data corresponding to the wrist joint is updated based on the posture data corresponding to each target body joint and the posture data corresponding to the wrist joint.
[0042] The virtual avatar is driven based on the updated posture data of the wrist joint, the posture data of each hand joint other than the wrist joint, the position data of each human body joint, and the posture data corresponding to each torso joint.
[0043] A virtual avatar driving device, comprising:
[0044] A motion capture unit is used to capture the motion of the human body driving the virtual avatar in real time to obtain captured images;
[0045] The network acquisition unit is used to acquire the joint prediction network and the limb prediction network.
[0046] An image input unit is used to input the captured image into the joint prediction network to obtain the position data of multiple human joints output by the joint prediction network and the joint confidence level corresponding to each human joint.
[0047] The network utilization unit is used to input the position data of each human joint point into the limb prediction network to obtain the posture data corresponding to each body joint point output by the limb prediction network.
[0048] The threshold comparison unit is used to obtain the hand prediction network if the position data of each human joint includes the position data of the hand joint and the confidence of the joint corresponding to the hand joint is greater than a preset threshold.
[0049] An image cropping unit is used to crop the captured image to obtain a hand capture image;
[0050] The data output unit is used to input the hand capture image into the hand prediction network to obtain the pose data corresponding to multiple hand joints output by the hand prediction network.
[0051] The image driving unit is used to drive the virtual image based on the position data of each human joint, the posture data corresponding to each torso joint, and the posture data corresponding to each hand joint.
[0052] A virtual avatar driving device, including a memory and a processor;
[0053] The memory is used to store programs;
[0054] The processor is used to execute the program to implement the various steps of the virtual image driving method described above.
[0055] A readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the virtual avatar driving method described above.
[0056] As can be seen from the above technical solution, the virtual image driving method provided in this application obtains captured images by real-time motion capture of the human body driving the virtual image; acquires a joint prediction network and a limb prediction network; inputs the captured images into the joint prediction network to obtain the position data of multiple human joints output by the joint prediction network and the joint confidence level corresponding to each human joint; it can be seen that the above steps can realize the prediction of the position of each joint of the human body, and the reliability of each position can be determined according to the joint confidence level. The position data of each human joint is input into the limb prediction network to obtain the posture data corresponding to each torso joint output by the limb prediction network. If the position data of each human joint includes the position data of the hand joint and the confidence of the hand joint is greater than a preset threshold, the hand prediction network is obtained. The captured image is cropped to obtain a hand capture image. The hand capture image is input into the hand prediction network to obtain the position data of multiple hand joints and the posture data corresponding to each hand joint output by the hand prediction network. It can be seen that the above process can separate the prediction of human body movements and human hand movements, reduce the complexity of predicted posture data and the difficulty of predicting the body and hand at the same time, realize the lightweighting of the hand prediction network and the limb prediction network, speed up the prediction efficiency of this application, and realize the real-time prediction of posture data. The virtual image is driven according to the position data of each human joint, the posture data corresponding to each torso joint, and the posture data corresponding to each hand joint. It can be seen that the integrated driving of the virtual image can be achieved by fusing the posture data corresponding to the body and the posture data corresponding to the hand. Based on this, this application can simplify the difficulty and complexity of the virtual image driving process, thereby further improving the real-time performance and matching of virtual image driving and human body movements, and achieving accurate synchronization between virtual images and human body movements. Attached Figure Description
[0057] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0058] Figure 1 A flowchart of a virtual avatar driving method provided in an embodiment of this application;
[0059] Figure 2 A schematic diagram of two-dimensional human joints provided in this application embodiment;
[0060] Figure 3 A schematic diagram of three-dimensional human joints provided in this application embodiment;
[0061] Figure 4 A structural block diagram of a virtual avatar driving device provided in an embodiment of this application;
[0062] Figure 5 This is a hardware structure block diagram of a virtual avatar driving device provided in an embodiment of this application. Detailed Implementation
[0063] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0064] The virtual avatar driving method provided in this application can be applied to various system platforms, and the hardware supporting the system platform can be a PC, terminal or other processing device, and the execution subject of the method can be the processor of the system platform.
[0065] Next, combine Figure 1 The virtual avatar driving method of this application is described in detail, including the following steps:
[0066] Step S1: Perform motion capture on the human body driving the virtual image in real time to obtain captured images.
[0067] Specifically, a camera can be used to capture human movements in real time and obtain captured images.
[0068] The captured image can be in RGB format.
[0069] Step S2: Obtain the joint prediction network and limb prediction network.
[0070] Specifically, pre-trained joint prediction networks and pre-trained limb prediction networks can be obtained.
[0071] Step S3: Input the captured image into the joint prediction network to obtain the position data of multiple human joints output by the joint prediction network and the joint confidence level corresponding to each human joint.
[0072] Specifically, the input to the keypoint prediction network is a captured image. A coordinate system can be constructed from the captured image, and the keypoint prediction network outputs the position data of each keypoint corresponding to the captured image and the confidence score of each keypoint.
[0073] The location data can be a coordinate array, and the location data of each human joint can be an array containing the identifier of each human joint and the corresponding coordinates.
[0074] The joint confidence level of each human joint is the maximum thermal value in the thermal map of that human joint.
[0075] Multiple human joints can include multiple torso joints and multiple hand joints, or they can include only hand joints.
[0076] Step S4: Input the position data of each human joint point into the limb prediction network to obtain the posture data corresponding to each body joint point output by the limb prediction network.
[0077] Specifically, images can be captured in real time to obtain multiple consecutive captured images, and the position data of human joints corresponding to each captured image can be obtained by using a joint prediction network.
[0078] The input to the limb prediction network can be the position data of human joints corresponding to multiple consecutive captured images, and the output is the pose data corresponding to each body joint.
[0079] Human joints can refer to all joints in the human body, such as the joints corresponding to the left shoulder, the left wrist, and the right palm.
[0080] See Figure 2 Examples of 18 human joints from multiple human joint points are provided, along with their Chinese abbreviations.
[0081] Body joints can be any joint in the human body except for those in the hands.
[0082] The joints of the hand can include the finger joints of the left and right hands, the left wrist joints, and the right wrist joints.
[0083] The pose data for each body joint can be a rotation matrix or a rotation quaternion for that body joint.
[0084] Step S5: If the position data of each human joint includes the position data of the hand joint and the confidence of the joint corresponding to the hand joint is greater than the preset threshold, obtain the hand prediction network.
[0085] Specifically, it can be determined whether the human joint position data output by the joint prediction network includes hand joint position data, such as the position data of any finger bone joints.
[0086] When confirming the location data containing hand joints, it is determined whether the joint confidence of each hand joint is greater than the threshold. If the joint confidence of all hand joints is greater than the threshold, it is determined that the captured image contains a human hand image, and the virtual avatar's hand needs to be driven.
[0087] This threshold can be set according to actual needs; generally, it can be 0.5.
[0088] A hand prediction network can be used to predict human hand posture data, and based on this, the hands of virtual avatars can be driven.
[0089] Step S6: Crop the captured image to obtain a hand capture image.
[0090] Specifically, the input to the hand prediction network can be a human hand image cropped from a captured image. Therefore, the human hand image in the captured image can be cropped, and the resulting image is the hand capture image.
[0091] Step S7: Input the captured hand image into the hand prediction network to obtain the pose data corresponding to multiple hand joints output by the hand prediction network.
[0092] Specifically, the hand prediction network can output pose data corresponding to multiple hand joints from multiple consecutive captured images obtained by cropping multiple consecutive captured images.
[0093] The hand prediction network can include a hand joint prediction layer and a pose prediction layer. The hand capture image can be input into the hand joint prediction layer to obtain the position data of multiple hand joints output by the hand joint prediction layer. The pose prediction layer can be used to predict the pose data of each hand joint based on the position data of multiple hand joints.
[0094] The hand joint prediction layer can be further divided into a 3D prediction layer and a 2D prediction layer. The 3D prediction layer can be used to predict the 3D position data corresponding to each hand joint based on the input 3D hand capture image; the 2D prediction layer can be used to predict the 2D position data corresponding to each hand joint based on the 2D hand capture image.
[0095] Step S8: Drive the virtual image based on the position data of each human joint, the posture data corresponding to each torso joint, and the posture data corresponding to each hand joint.
[0096] Specifically, the virtual avatar can be driven based on the position and posture data of each human joint to achieve synchronization between the human posture and the virtual avatar.
[0097] As can be seen from the above technical solutions, the embodiments of this application provide a virtual image driving method, which can process the captured images to drive the virtual image. Therefore, this application can be applied to various terminals with cameras capable of capturing human movements, and does not have excessively high requirements for the cameras, nor does it depend on specific motion capture equipment and environments, thus improving the scope of application of this application. Simultaneously, this application utilizes a joint prediction network, a limb prediction network, and captured images to predict the position data of human joints and the posture data of torso joints. Furthermore, the position data of each human joint includes the position data of hand joints, and the confidence level of the joints corresponding to the hand joints is greater than a preset threshold. The hand capture image obtained by cropping the captured image and the hand prediction network are used to predict the posture data of the hand joints. The above process achieves the prediction of position and posture data for each human joint. During posture data prediction, different networks are used for predicting the posture data of the torso joints and the text data of the hand joints, reducing the complexity and difficulty of posture data prediction and further achieving network lightweighting and accuracy. This results in a more accurate match between the virtual image's movements and the human body's movements during subsequent synchronization, and achieves high efficiency in the prediction process. Finally, by using the predicted position and posture data of each human joint to drive the virtual image, it is possible to drive the virtual image based on human movements. Therefore, the virtual image driving method provided in this application not only achieves accurate synchronization between the virtual image and human movements but also possesses reliability, accuracy, efficiency, and wide applicability.
[0098] In some embodiments of this application, after step S4, inputting the position data of each human joint point into the limb prediction network to obtain the posture data corresponding to each torso joint point output by the limb prediction network, if the position data of each human joint point does not include the position data of the hand joint point, the position data of each human joint point and the posture data corresponding to each torso joint point can be directly used to drive the virtual image.
[0099] In some embodiments of this application, if the position data of each human joint includes the position data of the hand joint, but the confidence level of the joint corresponding to the hand joint is not greater than a preset threshold, then the position data of each human joint and the posture data corresponding to each torso joint can be used directly to drive the virtual image.
[0100] In some embodiments of this application, the process of obtaining the joint prediction network in step S2 is described in detail, and the steps are as follows:
[0101] S20. Obtain the initial joint prediction network and multiple human images.
[0102] Specifically, an initial joint prediction network can be constructed. The trained initial joint prediction network is a joint prediction network that is based on the input captured image and outputs the position data of multiple human joints corresponding to the captured image and the joint confidence of each human joint.
[0103] Multiple human images can include RGB images of multiple real human bodies captured by a camera in various poses, or images of a 3D human body model in different poses.
[0104] Each human body image is labeled with the location data of the body's joints.
[0105] A coordinate system can be constructed for each human body image, and the position coordinates of human body joints can be marked on each human body image.
[0106] Human images are not necessarily images that contain all human joints; some human joints may be missing in human images to improve the robustness of the joint prediction network.
[0107] S21. Based on the location data of the human joints marked on each human body image, draw the thermal image of the joints corresponding to the human body image.
[0108] Specifically, for each human body image, a thermal image of the corresponding joint point is drawn for each human body joint point, and each joint point thermal image contains multiple thermal values.
[0109] S22. Use each human body image and its corresponding joint point thermal image as one of the training data for training the initial joint point prediction network to obtain multiple training data.
[0110] Specifically, a human image labeled with the location data of human joints and a thermal image of the joints in the human image are used as one of the training data, and multiple training data are used to train the initial joint prediction network.
[0111] S23. The initial joint prediction network is trained using the multiple training data, and the initial joint prediction network obtained after training is used as the latest joint prediction network.
[0112] Specifically, a human body image is randomly selected and input into the initial joint prediction network to obtain the joint thermal images corresponding to each human body joint predicted by the joint prediction network, as well as the position data and joint confidence of each human body joint predicted by the joint prediction network based on each joint thermal image.
[0113] Based on the joint thermal images predicted by the joint prediction network and the joint thermal images corresponding to the human images input to the initial joint prediction network, the thermal loss value of the initial joint prediction network is calculated. Based on the human joint position data labeled in the human images input to the initial joint prediction network and the position data of each human joint predicted by the joint prediction network, the position loss value is calculated.
[0114] Based on the thermal loss value and the positional loss value, the parameters of the initial joint prediction network are adjusted until the sum of the thermal loss value and the positional loss value of the initial joint prediction network reaches the minimum value that the initial joint prediction network can achieve.
[0115] The final initial joint prediction network is the latest joint prediction network.
[0116] As can be seen from the above technical solution, this embodiment provides an optional method for obtaining a joint prediction network. Using this method, multiple human images labeled with location data and their corresponding joint heatmap images can be used to train an initial joint prediction network to obtain the final joint prediction network. This method can improve the robustness and prediction accuracy of the joint prediction network.
[0117] In some embodiments of this application, the process of obtaining the limb prediction network in step S2 is described in detail, and the steps are as follows:
[0118] S24. Obtain the initial limb prediction network.
[0119] Specifically, an initial limb prediction network can be constructed. The trained initial limb prediction network predicts the posture data corresponding to each body joint based on the input temporal human joint position data.
[0120] The limb prediction network can be used for the SMPLH 3D human body model. The SMPLH 3D human body model is a standard 3D human body model with added hand gestures.
[0121] S25. Set the posture and morphological parameters of the 3D human body model to obtain the position data of multiple body joints and the posture data of multiple body joints corresponding to different parameters.
[0122] Specifically, the different parameters are different attitude parameters and / or morphological parameters.
[0123] By adjusting the morphological parameters of the 3D human body model, 3D human bodies with different bone lengths are obtained. The same posture parameters are set for the 3D human body models with different bone lengths to simulate the position changes of each body joint in the same posture. Based on the position changes of each body joint corresponding to each posture, the posture data of each body joint and the temporal position data of each body joint corresponding to that posture are obtained.
[0124] The location data may include three-dimensional location data, or two-dimensional location data generated by projecting the three-dimensional location data back to the image coordinate system using a camera projection formula.
[0125] The process of repeatedly adjusting morphological and / or posture parameters yields a large amount of temporal position data and posture data of each body joint.
[0126] S26. Use the position data of each body joint corresponding to each parameter and the posture data of each body joint as one of the training data for training the initial limb prediction network to obtain multiple training data.
[0127] Specifically, the posture data of each body joint corresponding to the same posture parameter and the same morphological parameter, as well as the temporal position data of each body joint, constitute a training data set.
[0128] S27. Train the initial limb prediction network using multiple training data, and use the trained initial limb prediction network as the latest limb prediction network.
[0129] Specifically, the posture data of each body joint corresponding to the same posture parameter and the same morphological parameter, as well as the temporal position data of each body joint, are randomly selected and input into the initial limb prediction network to obtain the posture data of each body joint predicted by the initial limb prediction network.
[0130] The loss value of the initial limb prediction network is calculated based on the posture data corresponding to each body joint predicted by the initial limb prediction network and the posture data of each body joint input to the initial limb prediction network.
[0131] Based on this loss value, the parameters of the initial limb prediction network are adjusted until the loss value of the initial limb prediction network reaches the minimum value that the initial limb prediction network can achieve.
[0132] The initial limb prediction network obtained is the latest limb prediction network.
[0133] As can be seen from the above technical solution, this embodiment provides an optional method for obtaining a limb prediction network. This method utilizes a 3D human body model to generate training data, which is then used to train the initial limb prediction network. Therefore, this application can utilize a 3D human body model to generate a large amount of training data that is difficult to obtain from the real world, making the training data for the initial limb prediction network richer and improving the robustness of the limb prediction network.
[0134] In some embodiments of this application, the process of obtaining the hand prediction network in step S5 is described in detail, and the steps are as follows:
[0135] S50. Obtain the initial hand prediction network.
[0136] Specifically, an initial hand prediction network can be constructed. The trained initial hand prediction network is a hand prediction network that outputs the pose data of multiple hand joints corresponding to the input hand capture image.
[0137] S51. The initial hand prediction network is trained using a mixed data method, and the trained initial hand prediction network is used as the latest hand prediction network.
[0138] Specifically, real hand capture images and hand capture images obtained using human models can be used simultaneously as training data to train an initial hand prediction network, and the trained initial hand prediction network can be used as the latest hand prediction network.
[0139] As can be seen from the above technical solution, this embodiment provides an optional method for obtaining a hand prediction network. This method uses training data composed of a mixture of real hand capture images and hand capture images simulated by a human model to train an initial hand prediction network, thereby obtaining a trained hand prediction network. Through the use of real hand capture images, the pose data predicted by the hand prediction network is ensured to conform to natural human perception, improving the prediction reliability of the hand prediction network. Simulated hand capture images enrich the training data, improving the prediction accuracy of the hand prediction network. Therefore, the hand prediction network trained in this application has high prediction accuracy.
[0140] In some embodiments of this application, the process of training the initial hand prediction network using mixed data in step S51, and using the trained initial hand prediction network as the latest hand prediction network, is described in detail below:
[0141] S510. Acquire multiple three-dimensional hand data, wherein each three-dimensional hand data is generated based on multiple two-dimensional images of a real human body in the same pose, and each three-dimensional hand data includes a hand image with two-dimensional position data of multiple hand joints labeled, three-dimensional position data of multiple hand joints, and pose data corresponding to multiple hand joints.
[0142] Specifically, multiple consecutive frames of two-dimensional images of the same real human body from different angles can be used to annotate the two-dimensional position data of each hand joint in each two-dimensional image, and the three-dimensional position data of each hand joint can be determined based on the two-dimensional images from different angles.
[0143] Based on multiple consecutive frames of two-dimensional images and the position data corresponding to each frame of two-dimensional images, pose data of each hand joint and time-seriesd three-dimensional and two-dimensional position data of each hand joint can be generated.
[0144] A hand image labeled with two-dimensional position data of multiple hand joints, three-dimensional position data of multiple hand joints, and posture data corresponding to multiple hand joints can be used to form three-dimensional hand data.
[0145] S511. Acquire multiple two-dimensional hand data, wherein each two-dimensional hand data is generated based on a real two-dimensional image, and each two-dimensional hand data includes a hand image labeled with two-dimensional position data of multiple hand joints and posture data of multiple hand joints.
[0146] Specifically, the two-dimensional position data of hand joints in multiple consecutive frames of real two-dimensional images can be labeled to obtain temporally sequenced two-dimensional position data and posture data of each hand joint.
[0147] Two-dimensional hand data can be composed of two-dimensional position data and posture data of multiple hand joints labeled with two-dimensional position data and multiple hand joints labeled with two-dimensional position data.
[0148] S512. Acquire multiple simulation data. Each simulation data is obtained by setting the posture parameters and morphological parameters of the three-dimensional hand model. Each simulation data includes a hand image with two-dimensional position data of multiple hand joints labeled, three-dimensional position data of multiple hand joints, and posture data corresponding to multiple hand joints.
[0149] Specifically, the three-dimensional hand model can be the MANO human hand model.
[0150] By adjusting the morphological parameters of the 3D hand model, 3D human hands with different bone lengths can be obtained. By setting the same pose parameters for 3D human models with different bone lengths, the 3D hand model can simulate the positional changes of various hand joints in different body postures under the same pose. Based on the positional changes of each hand joint corresponding to each body posture, the pose data of each hand joint and the temporal position data of each hand joint corresponding to that body posture are obtained. The hand image of that body posture under that pose is labeled with the pose data of each hand joint and the temporal position data of each hand joint.
[0151] By repeatedly adjusting the morphological and / or posture parameters, a large amount of time-sequential positional data and posture data of each hand joint can be obtained.
[0152] The simulation data may include time-series three-dimensional and two-dimensional position data of each hand joint with the same posture and morphological parameters, as well as hand images with posture data of each hand joint.
[0153] S513. A training set is formed by combining multiple 3D hand data, multiple 2D hand data, and multiple simulation data.
[0154] Specifically, the training set for the initial hand prediction network can consist of multiple 3D hand data sets, multiple 2D hand data sets, and multiple simulated data sets.
[0155] S514. Using the training set, train the initial hand prediction network.
[0156] Specifically, training data is randomly selected from the training set, which can be three-dimensional hand data, two-dimensional hand data, or simulated data.
[0157] The training data is input into the hand prediction network to obtain the initial position and pose data of each hand joint predicted by the hand prediction network. If the training image is three-dimensional hand data or simulation data, the position data of each hand joint can include the three-dimensional position data and two-dimensional position data of each hand joint. If the training data is two-dimensional hand data, the position data of each joint can be the two-dimensional position data of each hand joint.
[0158] Based on the positional data of each hand joint and the labeled data of the training data, the loss value of the initial hand prediction network is calculated.
[0159] The parameters of the initial hand prediction network are adjusted based on this loss value until the loss value of the initial hand prediction network reaches the minimum value that the initial hand prediction network can achieve.
[0160] The resulting initial hand prediction network is the latest hand prediction network.
[0161] As can be seen from the above technical solution, this embodiment provides an optional method for training a hand prediction network using mixed data. Through the above method, the initial hand prediction network can be trained using real two-dimensional data, real three-dimensional data, and simulated three-dimensional data, which further expands the applicable scenarios of the hand prediction network and improves the practicality of this application.
[0162] In some embodiments of this application, the process of cropping the captured image to obtain a hand capture image in step S6 is described in detail, and the steps are as follows:
[0163] S60. Based on the position data of multiple human joints output by the joint prediction network, determine the rectangular region of the face in the captured image, wherein the rectangular region of the face contains the face image of the human body.
[0164] Specifically, the rectangular region of the face in the captured image can be determined based on the position data of the facial key points in the captured image.
[0165] S61. Enlarge the rectangular area of the face according to a preset ratio to obtain a rectangular frame.
[0166] Specifically, the preset ratio can be set according to actual needs. The ratio can be set to 1.5, and each rectangular side of the face rectangle area can be enlarged according to the ratio to obtain a rectangular frame.
[0167] S62. Based on the position data of the hand joints, the rectangular frame is moved in the captured image to obtain a rectangular area of the hand.
[0168] Specifically, the rectangular frame is moved using the position data of the hand joints. The rectangular area corresponding to the moved rectangular frame is the hand rectangular area, which may contain an image of the human hand.
[0169] When capturing multiple consecutive frames of images, the position data of the hand joints are different in different captured images, and the position of the rectangular area of the hand is also different.
[0170] S63. Crop the rectangular area of the hand in the captured image, and use the cropped rectangular area of the hand as the captured image of the hand.
[0171] Specifically, the rectangular area of the hand containing the human hand image is cropped out of the captured image to obtain the hand capture image.
[0172] As can be seen from the above technical solution, this application can accurately acquire hand capture images by capturing images. By using position data and proportions, the size of the hand capture image can be reduced while ensuring high accuracy in cropping the hand capture image, thereby further reducing the complexity of predicting the pose data of hand joints.
[0173] In some embodiments of this application, the process of driving the virtual image based on the position data of each human joint, the posture data corresponding to each torso joint, and the posture data corresponding to each hand joint is described in detail below:
[0174] S80. Determine the root node position data of the human body image in the captured image, and draw the motion chain corresponding to the human hand with the root node position data as the starting point. The motion chain contains multiple target body joints.
[0175] Specifically, the root node of a human body image is generally located in the abdominal image. A typical human body has a left hand and a right hand. The left hand corresponds to a left-hand kinematic chain, and the right hand corresponds to a right-hand kinematic chain.
[0176] Starting from the root node, the kinematic chain along the torso to the elbow corresponds to the human hand. For example, see... Figure 3 The node corresponding to node 0 is the root node of the human body. The right hand's corresponding right hand kinematic chain is node 0, node 3, node 6, node 9, node 13, node 16 and node 18; the left hand's corresponding left hand kinematic chain is node 0, node 3, node 6, node 9, node 14, node 17 and node 19.
[0177] The target body joints are the various body joints on the kinematic chain corresponding to the human hand.
[0178] S81. Determine the posture data corresponding to the wrist joint from the various posture data, and determine the posture data corresponding to each target body joint from the posture data corresponding to each body joint.
[0179] Specifically, the posture data corresponding to each target body joint and the posture data corresponding to the wrist joint can be determined.
[0180] S82. Update the posture data corresponding to the wrist joint based on the posture data corresponding to each target body joint and the posture data corresponding to the wrist joint.
[0181] Specifically, the rotational postures of the body and wrists can be spliced along the kinetic chain to update the wrist joint posture data. This can be achieved by using a rotational splicing method.
[0182] The rotating splicing method is shown below:
[0183] R H =inv(R0×R3×R6×R9×R 14 ×R 17 ×R 19 )×R h
[0184] Among them, R H For the updated wrist joint pose data, inv indicates converting the matrix to its inverse, and R... h For the wrist joint pose data before the update, R n This refers to the attitude data corresponding to node n.
[0185] S83. Drive the virtual image based on the updated posture data of the wrist joint, the posture data of each hand joint other than the wrist joint, the position data of each human body joint, and the posture data corresponding to each torso joint.
[0186] Specifically, the virtual avatar can be driven based on the position and posture data of the human body's joints to achieve synchronization between human body movements and the virtual avatar.
[0187] As can be seen from the above technical solutions, compared with the above embodiments, this application provides an optional way to drive virtual images. By using the above method, the wrist joint posture data predicted by the hand prediction network and the posture data of the target body joints that affect the wrist joint posture data are used to update the wrist joint posture data, which can further ensure the accuracy and authenticity of the wrist joint posture data.
[0188] Next, we will combine Figure 4 The virtual avatar driving device provided in this application will be described in detail. The virtual avatar driving device described below and the virtual avatar driving method described above can be referred to in correspondence.
[0189] See Figure 4 Therefore, the virtual avatar driving device may include:
[0190] Motion capture unit 1 is used to capture the motion of the human body driving the virtual image in real time and obtain captured images;
[0191] Network acquisition unit 2 is used to acquire the joint prediction network and the limb prediction network;
[0192] Image input unit 3 is used to input the captured image into the joint prediction network to obtain the position data of multiple human joints output by the joint prediction network and the joint confidence level corresponding to each human joint.
[0193] Network utilization unit 4 is used to input the position data of each human joint point into the limb prediction network to obtain the posture data corresponding to each body joint point output by the limb prediction network.
[0194] Threshold comparison unit 5 is used to obtain the hand prediction network if the position data of each human joint includes the position data of the hand joint and the confidence of the joint corresponding to the hand joint is greater than a preset threshold.
[0195] Image cropping unit 6 is used to crop the captured image to obtain a hand capture image;
[0196] The data output unit 7 is used to input the hand capture image into the hand prediction network to obtain the pose data corresponding to multiple hand joints output by the hand prediction network.
[0197] The image driving unit 8 is used to drive the virtual image based on the position data of each human joint, the posture data corresponding to each torso joint, and the posture data corresponding to each hand joint.
[0198] Furthermore, the network acquisition unit may include:
[0199] An image acquisition unit is used to acquire an initial joint prediction network and multiple human images, each of which is labeled with the position data of human joints.
[0200] The thermal image rendering unit is used to render the thermal image of the joint points corresponding to each human body image based on the position data of the human body joint points marked on each human body image;
[0201] A thermal image utilization unit is used to use each human body image and its corresponding joint point thermal image as one of the training data for training the initial joint point prediction network, thereby obtaining multiple training data.
[0202] The training data utilization unit is used to train the initial joint prediction network using the multiple training data, and to use the initial joint prediction network obtained after training as the latest joint prediction network.
[0203] Furthermore, the network acquisition unit may also include:
[0204] The limb prediction network acquisition unit is used to acquire the initial limb prediction network.
[0205] The parameter setting unit is used to set the posture and shape parameters of the three-dimensional human body model, and obtain the position data of multiple body joints and the posture data of multiple body joints corresponding to different parameters.
[0206] The training data acquisition unit is used to take the position data of each body joint corresponding to each parameter and the posture data of each body joint as one of the training data for training the initial limb prediction network, and obtain multiple training data.
[0207] The network training unit is used to train the initial limb prediction network using multiple training data, and to use the trained initial limb prediction network as the latest limb prediction network.
[0208] Furthermore, the threshold comparison unit may include:
[0209] The hand prediction network acquisition unit is used to acquire the initial hand prediction network.
[0210] The hybrid data utilization unit is used to train the initial hand prediction network using hybrid data, and to use the trained initial hand prediction network as the latest hand prediction network.
[0211] Furthermore, the hybrid data utilization unit may include:
[0212] A three-dimensional hand data acquisition unit is used to acquire multiple three-dimensional hand data, wherein each three-dimensional hand data is generated based on multiple two-dimensional images of a real human body in the same pose, and each three-dimensional hand data contains a hand image with two-dimensional position data of multiple hand joints marked, three-dimensional position data of multiple hand joints, and pose data corresponding to multiple hand joints.
[0213] A two-dimensional hand data acquisition unit is used to acquire multiple two-dimensional hand data, wherein each two-dimensional hand data is generated based on a real two-dimensional image, and each two-dimensional hand data includes a hand image with two-dimensional position data and posture data of multiple hand joints marked.
[0214] The simulation data acquisition unit is used to acquire multiple simulation data. Each simulation data is obtained by setting the posture parameters and morphological parameters of the three-dimensional hand model. Each simulation data includes a hand image with two-dimensional position data of multiple hand joints labeled, three-dimensional position data of multiple hand joints, and posture data corresponding to multiple hand joints.
[0215] The training set combination unit is used to combine multiple 3D hand data, multiple 2D hand data, and multiple simulation data into a training set;
[0216] The training set utilization unit is used to train the initial hand prediction network using the training set.
[0217] Furthermore, the image cropping unit may include:
[0218] A face rectangle region determination unit is used to determine a face rectangle region in the captured image based on the position data of multiple human joints output by the joint prediction network, wherein the face rectangle region contains the face image of the human body.
[0219] The rectangular frame acquisition unit is used to enlarge the rectangular area of the face according to a preset ratio to obtain a rectangular frame;
[0220] A rectangular frame moving unit is used to move the rectangular frame in the captured image according to the position data of the hand joints to obtain a rectangular area of the hand;
[0221] A hand rectangular region cropping unit is used to crop a rectangular region of the hand in the captured image, and the cropped rectangular region of the hand is used as the captured hand image.
[0222] Furthermore, the data output unit may include:
[0223] The motion chain drawing unit is used to determine the root node position data of the human body image in the captured image, and draw the motion chain corresponding to the human hand with the root node position data as the starting point. The motion chain contains multiple target body joints.
[0224] The posture data determination unit is used to determine the posture data corresponding to the wrist joint from various posture data, and to determine the posture data corresponding to each target body joint from the posture data corresponding to each body joint.
[0225] The posture data update unit is used to update the posture data corresponding to the wrist joint based on the posture data corresponding to each target body joint and the posture data corresponding to the wrist joint.
[0226] The posture data utilization unit is used to drive the virtual image based on the updated posture data of the wrist joint, the posture data of each hand joint other than the wrist joint, the position data of each human body joint, and the posture data corresponding to each torso joint.
[0227] The virtual avatar driving device provided in this application embodiment can be applied to virtual avatar driving devices, such as PC terminals, mobile terminals, and desktop computers. Optionally, Figure 5 The hardware structure block diagram of the virtual avatar driving device is shown. Figure 5 The hardware structure of a virtual avatar driving device may include: at least one processor 1, at least one communication interface 2, at least one memory 3, and at least one communication bus 4;
[0228] In this embodiment of the application, the number of processor 1, communication interface 2, memory 3, and communication bus 4 is at least one, and processor 1, communication interface 2, and memory 3 communicate with each other through communication bus 4;
[0229] Processor 1 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.
[0230] Memory 3 may include high-speed RAM, and may also include non-volatile memory, such as at least one disk storage device;
[0231] The memory stores a program, which the processor can call. The program is used for:
[0232] Real-time motion capture of the human body driving the virtual avatar is performed to obtain captured images;
[0233] Obtain the joint prediction network and limb prediction network;
[0234] The captured image is input into the joint prediction network to obtain the position data of multiple human joints output by the joint prediction network and the joint confidence level corresponding to each human joint.
[0235] The position data of each human joint is input into the limb prediction network to obtain the posture data corresponding to each body joint output by the limb prediction network.
[0236] If the location data of each human joint includes the location data of the hand joint and the confidence of the joint corresponding to the hand joint is greater than a preset threshold, obtain the hand prediction network.
[0237] The captured image is cropped to obtain a hand capture image;
[0238] The hand capture image is input into the hand prediction network to obtain the pose data corresponding to multiple hand joints output by the hand prediction network;
[0239] The virtual avatar is driven based on the position data of each human body joint, the posture data corresponding to each torso joint, and the posture data corresponding to each hand joint.
[0240] Optionally, the refined and extended functions of the program can be found in the description above.
[0241] This application embodiment also provides a readable storage medium that can store a program suitable for execution by a processor, the program being used for:
[0242] Real-time motion capture of the human body driving the virtual avatar is performed to obtain captured images;
[0243] Obtain the joint prediction network and limb prediction network;
[0244] The captured image is input into the joint prediction network to obtain the position data of multiple human joints output by the joint prediction network and the joint confidence level corresponding to each human joint.
[0245] The position data of each human joint is input into the limb prediction network to obtain the posture data corresponding to each body joint output by the limb prediction network.
[0246] If the location data of each human joint includes the location data of the hand joint and the confidence of the joint corresponding to the hand joint is greater than a preset threshold, obtain the hand prediction network.
[0247] The captured image is cropped to obtain a hand capture image;
[0248] The hand capture image is input into the hand prediction network to obtain the pose data corresponding to multiple hand joints output by the hand prediction network;
[0249] The virtual avatar is driven based on the position data of each human body joint, the posture data corresponding to each torso joint, and the posture data corresponding to each hand joint.
[0250] Optionally, the refined and extended functions of the program can be referred to the above description.
[0251] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0252] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0253] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. The various embodiments of this application can be combined with each other. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A virtual avatar driving method, characterized in that, include: Real-time motion capture of the human body driving the virtual avatar is performed to obtain captured images; Obtain the joint prediction network and limb prediction network; The captured image is input into the joint prediction network to obtain the position data of multiple human joints output by the joint prediction network and the joint confidence level corresponding to each human joint. The position data of each human joint is input into the limb prediction network to obtain the posture data corresponding to each body joint output by the limb prediction network. If the location data of each human joint includes the location data of the hand joint and the confidence of the joint corresponding to the hand joint is greater than a preset threshold, obtain the hand prediction network. The captured image is cropped to obtain a hand capture image; The hand capture image is input into the hand prediction network to obtain the pose data corresponding to multiple hand joints output by the hand prediction network; The virtual avatar is driven based on the position data of each human body joint, the posture data corresponding to each torso joint, and the posture data corresponding to each hand joint.
2. The virtual avatar driving method according to claim 1, characterized in that, The acquisition of the key point prediction network includes: An initial joint prediction network and multiple human images are obtained, with the location data of human joints marked on each human image. Based on the location data of the human joints marked on each human body image, draw the thermal image of the joints corresponding to the human body image; Each human body image and its corresponding joint point thermal image are used as one of the training data for training the initial joint point prediction network, resulting in multiple training data. The initial joint prediction network is trained using the multiple training data, and the initial joint prediction network obtained after training is used as the latest joint prediction network.
3. The virtual avatar driving method according to claim 1, characterized in that, The limb prediction network includes: Obtain the initial limb prediction network; Set the pose and shape parameters of the 3D human body model to obtain the position data of multiple body joints and the pose data of multiple body joints corresponding to different parameters. The position data of each body joint corresponding to each parameter and the posture data of each body joint are used as one of the training data for training the initial limb prediction network, resulting in multiple training data. The initial limb prediction network is trained using multiple training data sets, and the trained initial limb prediction network is used as the latest limb prediction network.
4. The virtual avatar driving method according to claim 1, characterized in that, The method for obtaining the hand prediction network includes: Obtain the initial hand prediction network; The initial hand prediction network is trained using a mixed data approach, and the trained initial hand prediction network is used as the latest hand prediction network.
5. The virtual avatar driving method according to claim 4, characterized in that, The initial hand prediction network is trained using a mixed data approach, including: Multiple three-dimensional hand data are acquired, wherein each three-dimensional hand data is generated based on multiple two-dimensional images of a real human body in the same pose, and each three-dimensional hand data contains a hand image with two-dimensional position data of multiple hand joints labeled, three-dimensional position data of multiple hand joints, and pose data corresponding to multiple hand joints. Multiple two-dimensional hand data are acquired, wherein each two-dimensional hand data is generated based on a real two-dimensional image, and each two-dimensional hand data includes a hand image labeled with two-dimensional position data of multiple hand joints and posture data of multiple hand joints; Multiple simulation data are acquired. Each simulation data is obtained by setting the posture parameters and morphological parameters of a three-dimensional hand model. Each simulation data includes a hand image with two-dimensional position data of multiple hand joints labeled, three-dimensional position data of multiple hand joints, and posture data corresponding to multiple hand joints. The training set is composed of the multiple three-dimensional hand data, the multiple two-dimensional hand data, and the multiple simulation data; The initial hand prediction network is trained using the training set.
6. The virtual avatar driving method according to any one of claims 1-5, characterized in that, The captured image is cropped to obtain a hand capture image, including: Based on the position data of multiple human joints output by the joint prediction network, a rectangular region of the face in the captured image is determined, and the rectangular region of the face contains the face image of the human body. The rectangular area of the face is enlarged according to a preset ratio to obtain a rectangular frame; Based on the position data of the hand joints, the rectangular frame is moved in the captured image to obtain a rectangular area of the hand; A rectangular area of the hand is cropped from the captured image, and the cropped rectangular area of the hand is used as the captured image of the hand.
7. The virtual avatar driving method according to any one of claims 1-5, characterized in that, The process of driving the virtual avatar based on the position data of each human body joint, the posture data corresponding to each torso joint, and the posture data corresponding to each hand joint includes: Determine the root node position data of the human body image in the captured image, and draw the motion chain corresponding to the human hand with the root node position data as the starting point. The motion chain contains multiple target body joints. The posture data corresponding to the wrist joint is determined from the posture data of each posture data, and the posture data corresponding to each target body joint is determined from the posture data corresponding to each body joint. The posture data corresponding to the wrist joint is updated based on the posture data corresponding to each target body joint and the posture data corresponding to the wrist joint. The virtual avatar is driven based on the updated posture data of the wrist joint, the posture data of each hand joint other than the wrist joint, the position data of each human body joint, and the posture data corresponding to each torso joint.
8. A virtual avatar driving device, characterized in that, include: A motion capture unit is used to capture the motion of the human body driving the virtual avatar in real time to obtain captured images; The network acquisition unit is used to acquire the joint prediction network and the limb prediction network. An image input unit is used to input the captured image into the joint prediction network to obtain the position data of multiple human joints output by the joint prediction network and the joint confidence level corresponding to each human joint. The network utilization unit is used to input the position data of each human joint point into the limb prediction network to obtain the posture data corresponding to each body joint point output by the limb prediction network. The threshold comparison unit is used to obtain the hand prediction network if the position data of each human joint includes the position data of the hand joint and the confidence of the joint corresponding to the hand joint is greater than a preset threshold. An image cropping unit is used to crop the captured image to obtain a hand capture image; The data output unit is used to input the hand capture image into the hand prediction network to obtain the pose data corresponding to multiple hand joints output by the hand prediction network. The image driving unit is used to drive the virtual image based on the position data of each human joint, the posture data corresponding to each torso joint, and the posture data corresponding to each hand joint.
9. A virtual avatar driving device, characterized in that, Including memory and processor; The memory is used to store programs; The processor is configured to execute the program to implement each step of the virtual image driving method as described in any one of claims 1-7.
10. A readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the virtual image driving method as described in any one of claims 1-7.
Citation Information
Patent Citations
Virtual object driving method and device, electronic equipment and readable memory
CN111694429A
Human body posture estimation method and device, electronic equipment and storage medium
CN115346239A