Video action extraction method, storage medium and electronic device

CN122513612APending Publication Date: 2026-08-04NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NETEASE (HANGZHOU) NETWORK CO LTD
Filing Date
2026-06-24
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

[0003]然而,在对现有技术的研究和实践过程中发现,现有的视频动作提取方法的动作捕捉精度较低,使得视频动作提取效率较低

Benefits of technology

[0010]This application embodiment identifies the detection area information, body posture features, and acquisition device parameter information of the target object in the target video; extracts body movements from the target video based on the detection area information, body posture features, and acquisition device parameter information to obtain the body movement information of the target object in the target video; extracts hand movements from the target video based on the body movement information to obtain the hand movement information of the target object and the rotation information of the target joints associated with the hand; updates the body movement information based on the rotation information of the target joints; optimizes the physical rationality of the updated body movement information; and generates full-body movement information of the target object in the target video based on the optimized body movement information and hand movement information. Therefore, by using the target object's detection area information, body posture features, and acquisition device parameter information, the body motion information of the target object in the target video can be predicted. Then, by combining the body motion information, the hand motion information and the rotation information of the target joints associated with the hand can be accurately extracted from the target video. The body motion information is updated based on the rotation information of the target joints, and the updated body motion information is optimized for physical rationality. Based on the optimized body motion information and hand motion information, accurate and physically consistent full-body motion information can be extracted from the target video, improving the accuracy of video motion capture and further improving the efficiency of video motion extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122513612A_ABST
    Figure CN122513612A_ABST
Patent Text Reader

Abstract

This application discloses a video motion extraction method, storage medium, and electronic device. The method involves identifying the detection region information, body posture features, and initial acquisition device parameters of a target object in a target video. Based on these parameters, body motion is extracted to obtain the target object's body motion information. Hand motion is extracted from the target video, yielding hand motion information and rotation information of the target joints associated with the hand. The body motion information is then updated based on the joint rotation information. The updated body motion information is then physically optimized. Finally, based on the optimized body motion and hand motion information, the full-body motion information of the target object and the corresponding camera pose are obtained. This improves the accuracy and efficiency of motion extraction from videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, specifically to a video motion extraction method, apparatus, storage medium, and electronic device. Background Technology

[0002] Video motion capture, as a crucial technology in modern digital content creation, has played a key role in the gaming industry. Compared to traditional hand-drawn animation and complex light-capture animation, video motion capture technology requires only a few cameras to capture realistic human movement, balancing animation production efficiency and quality.

[0003] However, research and practice on existing technologies have revealed that existing video motion extraction methods have low motion capture accuracy, resulting in low video motion extraction efficiency. Summary of the Invention

[0004] This application provides a video motion extraction method, apparatus, storage medium, and electronic device, which can improve the accuracy of video motion capture and further improve the efficiency of video motion extraction.

[0005] This application provides a video action extraction method, including: In the target video, the detection area information, body posture features, and acquisition device parameter information corresponding to the target object in the target video are identified; Based on the detection area information, the body posture features, and the acquisition device parameter information, body motion is extracted from the target video to obtain the body motion information of the target object in the target video. Based on the body motion information, hand motion is extracted from the target video to obtain the hand motion information of the target object and the rotation information of the target joint associated with the hand; The body motion information is updated based on the rotation information of the target joint; The updated body motion information is then optimized for physical plausibility. Based on the optimized body motion information and hand motion information, full-body motion information of the target object in the target video is generated.

[0006] Accordingly, embodiments of this application provide a video motion extraction device, including: The recognition unit is used to identify the detection area information, body posture features, and acquisition device parameter information corresponding to the target object in the target video. The first extraction unit is used to extract body movements from the target video based on the detection area information, the body posture features and the acquisition device parameter information, to obtain the body movement information of the target object in the target video. The second extraction unit is used to extract hand movements from the target video based on the body movement information, so as to obtain the hand movement information of the target object and the rotation information of the target joint associated with the hand. An update unit is used to update the body motion information based on the rotation information of the target joint; The optimization unit is used to perform physical rationality optimization on the updated body motion information; The generation unit is used to generate full-body motion information of the target object in the target video based on the optimized body motion information and the hand motion information.

[0007] Furthermore, embodiments of this application also provide a computer-readable storage medium storing a computer program adapted for loading by a processor to execute steps in any of the video motion extraction methods provided in embodiments of this application.

[0008] Furthermore, this application also provides an electronic device, including a processor and a memory, wherein the memory stores an application program, and the processor is used to run the application program in the memory to implement the video motion extraction method provided in this application.

[0009] This application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. When the processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, causing the electronic device to perform the steps in the video motion extraction method provided in this application.

[0010] This application embodiment identifies the detection area information, body posture features, and acquisition device parameter information of the target object in the target video; extracts body movements from the target video based on the detection area information, body posture features, and acquisition device parameter information to obtain the body movement information of the target object in the target video; extracts hand movements from the target video based on the body movement information to obtain the hand movement information of the target object and the rotation information of the target joints associated with the hand; updates the body movement information based on the rotation information of the target joints; optimizes the physical rationality of the updated body movement information; and generates full-body movement information of the target object in the target video based on the optimized body movement information and hand movement information. Therefore, by using the target object's detection area information, body posture features, and acquisition device parameter information, the body motion information of the target object in the target video can be predicted. Then, by combining the body motion information, the hand motion information and the rotation information of the target joints associated with the hand can be accurately extracted from the target video. The body motion information is updated based on the rotation information of the target joints, and the updated body motion information is optimized for physical rationality. Based on the optimized body motion information and hand motion information, accurate and physically consistent full-body motion information can be extracted from the target video, improving the accuracy of video motion capture and further improving the efficiency of video motion extraction. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a schematic diagram illustrating an implementation scenario of a video action extraction method provided in this application embodiment; Figure 2 This is a flowchart illustrating a video action extraction method provided in an embodiment of this application; Figure 3a This is a schematic diagram of body motion extraction in a video motion extraction method provided in an embodiment of this application; Figure 3b This is a schematic diagram of hand motion extraction in a video motion extraction method provided in this application embodiment; Figure 3c This is a schematic diagram of another hand motion extraction method provided in this application embodiment for video motion extraction; Figure 3dThis is a schematic diagram of the body motion post-processing optimization process of a video motion extraction method provided in an embodiment of this application; Figure 3e This is a schematic diagram of the center of gravity optimization of a video action extraction method provided in an embodiment of this application; Figure 3f This is a schematic diagram of footstep optimization for a video motion extraction method provided in an embodiment of this application; Figure 4 This is a schematic diagram illustrating the specific process of a video motion extraction method provided in an embodiment of this application; Figure 5 This is a schematic diagram of the video motion extraction device provided in the embodiments of this application; Figure 6 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0014] This application provides a video motion extraction method, apparatus, storage medium, and electronic device. The video motion extraction apparatus can be integrated into an electronic device, which may be a server or a terminal, etc.

[0015] The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal can include, but is not limited to, mobile phones, computers, smart voice interaction devices, smart home appliances, vehicle terminals, and aircraft. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited herein.

[0016] Please see Figure 1 Taking the integration of video motion extraction devices into electronic devices as an example, Figure 1This is a schematic diagram illustrating an implementation scenario of the video motion extraction method provided in this application. The electronic device can identify the detection area information, body posture features, and acquisition device parameter information corresponding to the target object in the target video. Based on the detection area information, body posture features, and acquisition device parameter information, it extracts body motion from the target video to obtain the target object's body motion information in the target video. Based on the body motion information, it extracts hand motion from the target video to obtain the target object's hand motion information and the rotation information of the target joints associated with the hand. It updates the body motion information based on the rotation information of the target joints. It optimizes the updated body motion information for physical plausibility. Based on the optimized body motion information and hand motion information, it generates the target object's full-body motion information in the target video.

[0017] It should be noted that, Figure 1 The illustrated scenario of the video action extraction method is merely an example. The implementation environment of the video action extraction method described in this application is for the purpose of more clearly illustrating the technical solution of this application and does not constitute a limitation on the technical solution provided in this application. Those skilled in the art will understand that with the evolution of data processing and the emergence of new business scenarios, the technical solution provided in this application is also applicable to similar technical problems.

[0018] The solutions provided in this application are specifically illustrated through the following embodiments. It should be noted that the order of description of the following embodiments is not intended to limit the preferred order of the embodiments.

[0019] This embodiment will be described from the perspective of a video motion extraction device, which can be integrated into an electronic device, such as a terminal and / or a server, and this application does not impose any limitations on it.

[0020] Please see Figure 2 , Figure 2 This is a flowchart illustrating the video motion extraction method provided in this application embodiment. The video motion extraction method includes: In step 101, the detection area information, body posture features, and acquisition device parameter information corresponding to the target object in the target video are identified in the target video.

[0021] The target video can be a video from which motion extraction is to be performed. For example, the target video can be a dynamic perspective video, i.e., a video with camera movement effects. Camera movement (i.e., moving shots) refers to a shooting method that creates dynamic video footage by changing the camera's position, direction, or speed. When the target video has camera movement effects, the pose of the camera capturing the target video changes. The target object can be an object with motion in the target video, such as a person, animal, or swaying plant. The detection region information can be information indicating the location of the target object in the video frame image of the target video. For example, the detection region information can be an object detection box. The object detection box can be a detection box indicating the location of the target object in the video frame image of the target video. The body pose feature can be information describing the body pose of the target object. For example, the body pose feature can be the body key points of the target object. The body key points can be information used to indicate the body pose of the target object, such as the target object's nose, left eye, right eye, left shoulder, right shoulder, elbow, and ankle. The acquisition device parameter information can be the parameter information of the acquisition device for the target video. The acquisition device can be a camera, and the parameter information can be camera parameters. The camera can be the one capturing the target video. When the target video is a dynamic viewpoint video, the camera's pose changes. The acquisition device parameter information can be relevant camera parameters, such as the camera's angular velocity, camera intrinsic parameters, and camera pose. This acquisition device parameter information can also be the camera's initial acquisition device parameter information, for example, initial camera parameters estimated from the target video.

[0022] There are several ways to identify the detection area information, body posture features, and acquisition device parameter information of the target object in the target video. For example, object recognition can be performed on the video frame images in the target video to obtain the detection area information and body posture features of the target object in the target video. Based on the feature points sampled in the area outside the detection area information in the video frame image, the acquisition device parameter information of the target video can be calculated.

[0023] The feature point can be a sampling point in the video frame image used to calculate the parameter information of the acquisition device.

[0024] There are several ways to calculate the acquisition device parameters corresponding to the target video based on feature points sampled from regions outside the detection region information in the video frame image. For example, these acquisition device parameters can include camera pose, angular velocity, and camera intrinsic parameters. Visual odometry can be used to estimate the camera pose and calculate the angular velocity. Then, feature points are sampled in the video frame image, and it is determined whether the sampled feature points are within the detection region information. If the sampled feature point is within the detection region information, it is discarded; if it is not, the camera's intrinsic parameters are estimated based on that feature point. Since the area where the target object is located in the video frame image is often dynamically changing, feature point sampling can be performed in areas outside the target object's location to accurately obtain the camera intrinsic parameters.

[0025] In step 102, body motion is extracted from the target video based on the detection area information, body posture features, and acquisition device parameter information to obtain the body motion information of the target object in the target video.

[0026] The body motion information can be information indicating the body movements of a target object in a target video. For example, the body motion information can include rotation information of the target object's bones or joints, as well as displacement information of the root bones. The rotation information can describe the rotation angle of the bones or joints. The displacement information can describe the displacement of the root bones.

[0027] Optionally, the body motion information may include the target object's body motion in the world coordinate system, the body motion in the camera's camera coordinate system, and the ground contact probability. The body motion may include information such as the rotation angles of the target object's bones or joints and the displacement of the root bones. The ground contact probability may be information indicating the likelihood of at least some joints in the body motion contacting the ground.

[0028] Optionally, each video frame in the target video can correspond to a detection area information, body posture features, and acquisition device parameter information. Based on the detection area information, body posture features, and acquisition device parameter information corresponding to each video frame, body motion information corresponding to each video frame can be obtained.

[0029] There are several ways to extract body movements from a target video based on detection area information, body posture features, and acquisition device parameter information to obtain the body movement information of the target object in the target video. For example, feature extraction can be performed on the detection area information, body posture features, and acquisition device parameter information to obtain the corresponding feature information. Then, the feature information can be input into a trained neural network model to output the body movement information of the target object in the target video.

[0030] In one embodiment, taking a person as an example, please refer to... Figure 3a , Figure 3a This is a schematic diagram illustrating body motion extraction in a video motion extraction method provided in this application embodiment. The body motion extraction method provided in this application embodiment includes human detection and tracking, an input preparation stage, and body motion prediction. In the human detection and tracking stage, an efficient strategy of tracking as the primary method and detection as a secondary method can be adopted. For example, the object detection box tracked in the previous video frame can be used for detection in the current video frame. Only when no human body is detected will a re-detection mechanism be initiated. The input preparation stage aims to provide the necessary input for the neural network in the next stage. First, the body key points of the target object can be extracted using the Vision Transformer (ViT) network structure. To facilitate the subsequent acquisition of the cropping box of the hand region, fingertip key points can be additionally annotated for training. To improve dance movement performance, large-scale dance movement data can be collected, and pelvic twisting effects can be specifically optimized. Second, visual odometry can be used to estimate the camera pose and calculate the angular velocity. For example, Deep Patch Visual Odometry (DPVO) can be used to set a mask for the dynamic region where the human body is located based on the body detection box, ensuring that most of the sampled feature points come from the static background. At the same time, the camera intrinsic parameters can be estimated based on the image pixel information of the video frame image.

[0031] Optionally, human visual features can be extracted from video frame images of the target video as auxiliary input for subsequent body motion prediction to improve performance in certain truncated scenes. If this step is not performed, an all-zero feature vector is used instead. This visual feature extraction can be implemented using a backbone network for single-frame human pose estimation.

[0032] In the body motion prediction stage, the body detection boxes, body keypoints, camera angular velocity, camera intrinsic parameters, and optional human visual features predicted in the previous stage can be used as inputs. These are converted into feature vectors by their respective encoders. These feature vectors are then concatenated and input into the neural network model (also called the body model) provided in this embodiment. The model is encoded using the encoder structure (Transformer Encoder) and decoded using the decoder, outputting the target object's body motion in the world coordinate system, the body motion in the camera coordinate system, and the ground contact probability. The neural network model used in this embodiment can be based on a Transformer structure and can incorporate relative position encoding. For example, Rotary Position Embedding (RoPE) can be used to enhance temporal modeling capabilities. Model training can utilize diverse action sequence data collected and organized independently. First, training can be performed on the full dataset to obtain strong generalization ability. Then, fine-tuning can be done on a carefully selected dance dataset to enhance the representation of action details such as pelvic twisting, thereby improving the accuracy of video motion extraction.

[0033] In step 103, hand movements are extracted from the target video based on body motion information to obtain the hand movement information of the target object and the rotation information of the target joints associated with the hand.

[0034] The hand movement information can be information describing the hand movements of the target object, such as the rotation values ​​of the joints of the target object's hand. The target joint can be a joint connected to the hand, such as an arm joint.

[0035] There are several ways to extract hand movements from a target video based on body motion information to obtain the hand movement information of the target object and the rotation information of the target joints associated with the hand. For example, the hand region of the target object and the orientation of the target object's palm can be determined in the video frame image of the target video based on body posture features; the hand key points of the target object can be predicted based on the hand region and the orientation of the palm; and the hand movement information of the target object and the rotation information of the target joints associated with the hand can be identified based on the hand key points, the hand region and body motion information.

[0036] The hand region can be the area that includes the target object's hand. The palm orientation can be information describing the orientation of the hand. The hand key points can be information describing the hand posture, and a digital model of the hand can be constructed by capturing the positions of the wrist, knuckles, and fingertips. For example, the hand key points can include key points such as the wrist joint, finger base, knuckles, and fingertips.

[0037] There are several ways to determine the hand region of a target object in the video frame image of a target video based on body posture features. For example, the size of the target part of the target object can be determined based on body posture features; the size of the target part of the hand can be determined based on the size ratio between the target part and the hand, as well as the size of the part; or the hand region including the hand of the target object can be determined in the video frame image of the target video based on the size of the target part.

[0038] The target area can be a region used to determine the size of the hand area, such as the forearm, neck, or head. The size of the target area can be its dimensions, such as its length. The proportional relationship can be the ratio between the target area and the hand; for example, assuming the target area is the neck, and the hand's length is 1.5 times the neck's length, the proportional relationship could describe that the hand's length is 1.5 times the neck's length. The target area size can also be the hand's dimensions, such as its length or width.

[0039] In one embodiment, there may be at least two target parts. To avoid errors in the determined target part size due to partial folding or obstruction, the part sizes of at least two target parts can be compared, and the target part size of the hand of the target object can be determined based on the target part with the larger part size.

[0040] Optionally, there are several ways to identify the hand movement information of the target object and the rotation information of the target joints associated with the hand based on hand key points, hand regions, and body movement information. For example, the position information of bones in a preset part of the target object can be extracted from the body movement information, and the position information of the bones can be normalized based on the position of the root bone corresponding to the bone; the hand key points can be normalized based on the hand region; and the hand movement information of the target object and the rotation information of the target joints associated with the hand can be predicted based on the normalized bone position information and hand key points.

[0041] The preset part can be a part related to the hand, such as the upper body or arm of the target object. The location information can be information indicating the position of the bones.

[0042] In one embodiment, please refer to Figure 3b , Figure 3bThis is a schematic diagram of hand motion extraction in a video motion extraction method provided in this application embodiment. The hand motion extraction method for dynamic viewpoints provided in this application embodiment includes two stages: hand keypoint extraction and hand motion prediction, and both stages incorporate the body information of the target object. In the hand keypoint extraction stage, the goal is to accurately locate finger keypoints from the input video frame image. First, the wrist and fingertip positions of the target object can be predicted based on the body model to assist in cropping the hand image. To avoid incomplete hand images due to fingertip position prediction errors, the cropping frame size of the hand area can be adjusted proportionally based on the forearm length and neck length to ensure the entire hand is included within the hand area. Second, to prevent large wrist rotation or twisting, the palm orientation (i.e., the direction from the wrist to the fingertips) can be input as an additional condition into the finger model to provide approximate palm information, thereby improving the stability and accuracy of keypoint prediction. In this way, hand images cropped from the hand region and the palm orientation can be combined and input into the hand model. Encoding is performed using the ViT Encoder and MLPencoder (Multilayer Perceptron Encoder) in the hand model, and finally the finger key point positions and confidence scores are decoded using the Decoder in the hand model.

[0043] Optionally, when training the hand model, to enrich the hand keypoint data, a large number of internet images can be collected, and pseudo-labels can be generated using high-precision models. For example, models such as Sapiens or the Hand Mesh Reconstructor (HaMeR) can be used. Sapiens can provide unstructured, purely 2D keypoint outputs, with a wide range of applications; HaMeR, on the other hand, can output keypoints with hand structure information, making it more robust. The complementary advantages of both can effectively improve the accuracy of pseudo-labels. Figure 3b In the middle, the rightmost figure can show that when a person makes a fist with their arm outstretched, the distance from the wrist to the fingertips alone cannot obtain a sufficiently large clipping frame. However, the embodiments of this application, based on the size ratio between the target part and the hand, can automatically adjust the clipping frame of the hand area during inference, thereby improving the accuracy and stability of hand motion extraction.

[0044] Optional, please refer to Figure 3c , Figure 3cThis is another hand motion extraction diagram provided in this application embodiment of a video motion extraction method. The hand motion prediction stage aims to accurately predict hand motion and update arm rotation by combining finger keypoints with the estimated body motion. Specifically, for dynamic camera environments, the three-dimensional position of the upper body skeleton (i.e., the preset part) can be extracted from the body motion in the camera coordinate system and normalized by translation with the position of the root bone as a reference. Subsequently, the finger keypoints can be normalized based on the cropping box of the hand region and concatenated with the upper body skeleton data to construct the input feature vector. Among them, finger keypoints with a confidence level lower than the preset confidence threshold will be set to zero. The feature sequence composed of the feature vectors corresponding to each video frame image in the target video is combined with the relative position encoding and input into the Transformer neural network to output the rotation sequence of the finger and arm joints. Here, the re-estimated arm joint rotation will cover the prediction result of the body model on the body motion, realizing the optimization of arm motion (mainly rotation around the bone axis) through finger information, which is equivalent to updating the body estimation using finger information, further improving the accuracy of video motion capture.

[0045] During the training of the hand model, the human body can be projected by randomly sampling the pose of a static camera, and random perturbations can be applied to the 2D projection positions of finger keypoints. The perturbation amplitude is set according to the accuracy of the actual 2D keypoint estimator. In addition, some hand keypoints can be randomly set to invisible. The upper body skeletal data can be kept without noise to ensure the stability of the skeletal information.

[0046] In step 104, the body motion information is updated based on the rotation information of the target joint.

[0047] Specifically, the rotation information of the target joint in the body motion information can be updated based on the rotation information of the target joint to obtain the updated body motion information.

[0048] Therefore, by combining body motion information to predict hand movements and generating hand motion information, and then optimizing the body motion information based on the hand motion prediction results, it is possible to accurately extract the hand and body movements of the target object in the target video, thereby improving the accuracy of motion extraction from the video.

[0049] In step 105, the updated body motion information is optimized for physical plausibility.

[0050] In order to extract full-body movements from the video that are more in line with the laws of physics, the extracted body movement information can be optimized for physical rationality.

[0051] Optionally, based on the detection area information, body posture features, and acquisition device parameter information, the ground contact probability of multiple candidate ground-contacting joints corresponding to body movement information can also be predicted.

[0052] The ground contact probability can be information indicating the likelihood of a candidate ground contact joint touching the ground. The candidate ground contact joint can be any joint that is likely to touch the ground; for example, it could include six candidate ground contact joints such as the ankle, toes, and hand. Optionally, each video frame corresponds to a body motion information set, and each body motion information set corresponds to a set of ground contact probabilities. This set of ground contact probabilities includes the probability of multiple candidate ground contact joints of the target object touching the ground in the current video frame.

[0053] There are several ways to optimize the physical rationality of the updated body motion information. For example, the target ground-touching joint can be determined from the candidate ground-touching joints based on the ground-touching probability, the updated body motion information can be optimized based on the target ground-touching joint, and the full-body motion information of the target object in the target video can be generated based on the optimized body motion information and hand motion information.

[0054] The target ground-contact joint can be a candidate ground-contact joint that has made contact with the ground. Ground contact refers to the joint touching the ground.

[0055] There are several ways to determine the target contact joint among the candidate contact joints based on the contact probability. For example, the contact probability can be corrected based on the contact correlation parameters of the candidate contact joints to obtain the target contact probability. The contact correlation parameters include at least one of vertical height and velocity. The target contact joint among the candidate contact joints is then determined based on the target contact probability.

[0056] The ground contact correlation parameter can be a parameter related to the ground contact situation, such as vertical height, velocity, acceleration, etc. The target ground contact probability can be a corrected ground contact probability.

[0057] Among them, the ground contact probability is corrected based on the ground contact correlation parameters of the candidate ground contact joints. There are several ways to obtain the target ground contact probability. For example, since the height of the ground contact joint is low enough and the speed is small enough, the ground contact probability of the candidate ground contact joint with a vertical height not less than a preset height threshold and / or a speed not less than a preset speed threshold can be set to 0 to obtain the target ground contact probability.

[0058] After correcting the contact probability based on the contact correlation parameters of the candidate contact joints, the target contact joint can be determined from the candidate contact joints according to the target contact probability. There are several ways to determine the target contact joint from the candidate contact joints according to the target contact probability. For example, candidate contact joints with a target contact probability greater than a preset probability threshold can be determined as the target contact joints.

[0059] The preset probability threshold can be a value such as 0.5 or 0.6, and the specific value can be set according to the actual situation. This application embodiment does not limit it here.

[0060] After determining the target ground-touching joint among the candidate ground-touching joints based on the ground-touching probability, the updated body motion information can be optimized based on the target ground-touching joint. There are several ways to optimize the updated body motion information based on the target ground-touching joint. For example, the target ground-touching joint can be moved to the ground position, and the positions of other bones or joints in the body motion information can be adjusted synchronously based on the position of the moved target ground-touching joint.

[0061] Optionally, there are several other ways to optimize the updated body motion information based on the target ground-touching joint. For example, if the target object is in a ground-touching state, the center of gravity position of the target object is corrected based on at least one first loss function. If the target object is not in a ground-touching state, the center of gravity position of the target object is corrected based on a second loss function, and the body motion information is optimized based on the corrected center of gravity position.

[0062] The first loss function can be used to constrain the motion of the target's ground-touching joints and the change in the center of gravity of the target object, while the second loss function can be used to constrain the trajectory of the center of gravity of the target object.

[0063] For example, the first loss function may include a loss function for constraining the velocity of the target ground-touching joint, a loss function for constraining the vertical height of the target ground-touching joint, a loss function for constraining the smoothness of the center of gravity change, and a loss function for constraining the distance between adjacent center of gravity positions, etc.

[0064] There are several ways to determine the second loss function. For example, the center of gravity trajectory of the target object can be parabolically fitted based on the center of gravity position of the target object in the body motion information to obtain the center of gravity trajectory curve; the second loss function can be determined based on the distance between the center of gravity of the target object and the center of gravity trajectory curve.

[0065] The center of gravity trajectory curve can be a curve obtained by parabolic fitting based on the center of gravity position of the target object in the body motion information corresponding to multiple video frames.

[0066] When the target object is not in contact with the ground, i.e., when the target object is in a suspended state, the trajectory curve of the target object's center of gravity can be considered to be approximately a parabola. Therefore, a parabolic fit can be performed on the trajectory of the target object's center of gravity. Then, the distance from the target object's center of gravity to the trajectory curve can be used as a loss function to optimize the center of gravity to approximate the parabola, thereby improving the accuracy of body motion extraction.

[0067] There are several ways to correct the center of gravity position of the target object based on at least one first loss function. For example, a third loss function can be determined based on the distance between the target trigger joint and the center of gravity. The center of gravity position and the position of the target ground-touching joint can be corrected based on the third loss function and the second loss function. The step of optimizing the body motion information based on the corrected center of gravity position can include: optimizing the body motion information based on the corrected center of gravity position and the position of the target ground-touching joint.

[0068] The third loss function can be used to constrain the distance between the target's ground-touching joint and its center of gravity.

[0069] In one embodiment, to extract more physically accurate full-body movements from the video, the body movements can be optimized based on the ground contact probability. For example, please refer to... Figure 3d , Figure 3d This is a schematic diagram of the body motion post-processing optimization process of a video motion extraction method provided in this application embodiment. This application embodiment provides a motion and shot optimization method for dynamic perspectives, which can include motion post-processing and camera motion post-processing. For physics-oriented character motion post-processing, in order to obtain human motion that better conforms to physical laws, the optimization scheme provided in this application embodiment can be adopted. Through a multi-stage, multi-objective design, the stability and effectiveness of the optimization are ensured. First, the body motion information can be corrected for floating and its trajectory smoothed. This step aims to eliminate the floating phenomenon of the human body by calculating the overall vertical offset of the motion and placing the lowest point of the human body on the ground. Simultaneously, the center of gravity trajectory is smoothed to effectively suppress jitter. Thus, the center of gravity T of the body motion information can be initially optimized.

[0070] Next, contact continuity correction can be performed. This step aims to filter out unreasonable contact estimates based on the velocity, height, and continuity assumptions of the contact joint. For example, the floating-point contact probability predicted by the neural network can be transformed into a binary variable C of 0 / 1 by setting a preset probability threshold. Then, the contact probability can be further corrected based on the velocity and height of the target contact joint to ensure that the velocity of the contact joint is sufficiently small and the height above the ground is sufficiently low, thereby updating the contact result. The specific optimization formula can be expressed as:

[0071] Where j can represent the index of the target ground contact joint, P j,y It can represent the vertical height of the target's ground-contact joint, & can represent logical AND, and | can represent logical OR. It can represent the velocity of the target's ground-contact joint. and The threshold values ​​represent vertical height and velocity; the superscripts 1 and 2 can indicate two sets of thresholds. For example, this... It can be 0.4. It can be 0.1. It can be 0.05. The threshold can be 0.01, in meters and meters per second. The first group of thresholds is more lenient and is mainly used to process results that have been predicted as target contact joints. The second group of thresholds is more stringent and is used to filter out possible contact points from video frames that have never been predicted as target contact joints.

[0072] Optionally, the continuity of ground contact status can be ensured by filtering out sporadic ground contact estimates. For example, a candidate ground contact joint can be identified as the target ground contact joint only if it is continuously in contact with the ground for a certain number of frames (e.g., 5 frames). Then, it can be verified whether the target ground contact joint is the joint with the lowest height in the body movement, and results that do not meet this assumption can be eliminated. In this way, a more accurate ground contact estimate can be obtained, providing a reliable basis for subsequent optimization of the center of gravity and limbs.

[0073] Optionally, physical center of gravity optimization can be performed on body movements. This step aims to optimize the character's center of gravity curve (i.e., root bone displacement) based on the ground contact assumption and the non-ground contact parabolic assumption, which can be divided into two cases. When the target object is in contact with the ground, the center of gravity T is used as the optimization variable, and the following first loss function L is established. cog :

[0074] Where j can represent the index of the target ground-contact joint, and T0 represents the initial value of the center of gravity. This indicates the new position of the target's ground-touching joint after the center of gravity has shifted. The subscripts x, y, and z can represent the corresponding coordinate axes. This represents the appropriate height for the corresponding joint to touch the ground; for example, it could be set to 0.2 cm for the toes. The entire optimization can be achieved using gradient descent.

[0075] Among them, L contact It can be used to constrain the speed of a target joint that is in contact with the ground. When a joint is determined to be in contact with the ground, its movement speed should be close to 0. C jThis can be expressed as the probability of touching the ground, or as a confidence weight. The larger this value is, the stronger the constraint, as long as it's certain the joint is on the ground. L ground It can be used to constrain the vertical height of the target's ground-contacting joint. When the joint contacts the ground, its vertical height (y-axis coordinate) should be equal to a preset reasonable height, which can prevent the foot from sinking into the ground or floating in mid-air. L smooth This can be expressed as the acceleration of the center of gravity should be close to 0, which can be used for smoothing constraints. L data It can be used to constrain the optimized centroid position T to not deviate too far from the originally captured centroid position T0. || can be represented as a norm, used to calculate the magnitude of a vector.

[0076] Optionally, a loss function L for balancing constraints can be enabled. balance This constraint is used to prevent the center of gravity from deviating too far from the line connecting the feet when the foot touches the ground. This constraint helps maintain balance during smaller movements.

[0077] When the target object is not in contact with the ground, i.e., in a suspended state, the center of gravity curve can be considered to approximate a parabola. Therefore, we can first perform parabolic fitting on the center of gravity trajectory, and then use the distance from the center of gravity to the center of gravity trajectory curve as a second loss function to optimize the center of gravity to make it approximate a parabola.

[0078] For example, please refer to Figure 3e , Figure 3e This is a schematic diagram of the center of gravity optimization of a video motion extraction method provided in this application embodiment. It can show the center of gravity trajectory of the target object before and after center of gravity optimization when the target object jumps from the spot. It can be seen that after center of gravity optimization, the center of gravity trajectory of the target object is closer to a parabola.

[0079] Optionally, distance-aware motion optimization can be performed on the body motion information. For example, based on the previous center of gravity optimization, limb movements in the body motion information can be further optimized. Specifically, the positions of the center of gravity and the ground-contact joints can be optimized, and the loss function L(T, P) is represented as follows:

[0080] Among them, L offset This can be represented as a third loss function, which can be used to constrain the distance from the center of gravity to the target contact joint to remain constant. cog This can be represented as the first loss function. For those contained in L cog Among the other loss functions, only the smoothing loss L... smoothThe position of the target ground-touching joint will be additionally considered, while other losses remain unchanged. The center of gravity and the target ground-touching joint position can be jointly optimized using gradient descent. Finally, based on the optimized center of gravity and target ground-touching joint positions, the rotation angle θ of the relevant limbs can be calculated using whole-body inverse kinematics, and the body motion information can be updated.

[0081] For example, please refer to Figure 3f , Figure 3f This is a schematic diagram of footwork optimization in a video motion extraction method provided in this application embodiment. The diagram compares the foot trajectories before and after center of gravity optimization. The red curve represents the foot trajectory before optimization, which is relatively disordered, while the green curve represents the foot trajectory after optimization, which is simpler. This shows that the embodiment of this application can achieve a more stable foot motion extraction effect.

[0082] Optionally, considering that real-world videos often feature half-body shots or characters with small displacements (such as popular short videos), applying excessive optimization in these situations, while improving the stability of the ground-touching joints and center of gravity, might cause unnecessary body shaking and affect the visual effect. Therefore, the weight parameters and iteration count can be adaptively adjusted based on the character's displacement in the world coordinate system and its distance from the camera. For example, the smaller the displacement and the closer the distance, the lower the corresponding optimization intensity. This differentiated correction strategy effectively ensures the stability and performance of the video motion extraction algorithm in such cases.

[0083] In step 106, based on the optimized body motion information and hand motion information, full-body motion information of the target object in the target video is generated.

[0084] The full-body motion information can be information describing the full-body motion of the target object, such as the target object's body movements and hand movements.

[0085] In one embodiment, video frames in the target video can be identified to obtain the detection area information, body posture features, and acquisition device parameter information corresponding to each video frame of the target object. For each video frame, body motion extraction can be performed based on the detection area information, body posture features, and acquisition device parameter information to obtain the body motion information of the target object in the current video frame. Then, hand motion extraction can be performed based on the body motion information to obtain the hand motion information and rotation information of the target joints associated with the hand in the current video frame. Next, the body motion information can be updated based on the rotation information of the target joints, and the updated body motion information can be physically optimized. Then, based on the optimized body motion information and hand motion information, the full-body motion information of the target object in the current video frame can be obtained. Thus, based on the full-body motion information corresponding to each video frame in the target video, the full-body motion sequence of the target object captured in the target video, i.e., skeletal animation, can be obtained.

[0086] In one embodiment, in order to restore the camera movement effect in the target video and improve the stability and accuracy of video motion extraction, the camera pose of the video frames in the target video can be extracted.

[0087] For example, the camera pose of the camera corresponding to the target video can be determined based on body motion information, and the camera pose is used to reconstruct the camera movement effect in the target video based on the full-body motion information.

[0088] The camera pose can be information describing the camera's pose, including information such as the camera's orientation and displacement.

[0089] The body and hand motion information from multiple video frames corresponding to the target video constitutes a motion sequence. These are combined to obtain a full-body motion sequence, i.e., human animation, which can be used to drive virtual game characters. The camera pose sequence corresponds one-to-one with body or hand motions. This human animation can be projected onto the screen using a camera model to recreate the camera movement effects of the input target video. This can be used by players to replicate shots or to prevent developers from manually drawing keyframes.

[0090] There are several ways to determine the camera pose of the camera corresponding to the target video based on body motion information. For example, based on body motion information, the first orientation and first displacement of the target object in the world coordinate system, and the second orientation and second displacement in the camera's camera coordinate system can be determined; based on the first orientation and second orientation, the target orientation of the camera in the world coordinate system can be calculated; based on the first orientation, second orientation, first displacement and second displacement, the target displacement of the camera in the world coordinate system can be calculated; and based on the target orientation and target displacement, the camera pose of the camera can be obtained.

[0091] In one embodiment, this application provides an adaptive camera motion post-processing method, which can indirectly obtain the camera pose by utilizing the orientation and displacement of the human body in different reference coordinate systems. The formula for this method can be as follows:

[0092] In this formula, the superscript denotes the reference coordinate system, where w represents the world coordinate system and c represents the camera coordinate system. The superscript T denotes matrix transpose, and the subscript denotes the object, where h represents the human body and c represents the camera. R and T can represent rotation (or orientation) and translation (or displacement), respectively. The above formula can be expressed as follows: through the first orientation of the human body in the world coordinate system... and the second orientation in the camera coordinate system It can calculate the target orientation of the camera in the world coordinate system. Furthermore, this is combined with the first displacement of the human body in the world coordinate system. and the second displacement in the camera coordinate system It can calculate the target displacement of the camera in the world coordinate system. Thus, the camera pose is obtained based on the target orientation and target displacement.

[0093] Optionally, there are several other ways to determine the camera pose of the target video based on body motion information. For example, a 3D skeletal model of the target object can be built based on the body motion information and key points corresponding to each video frame, and a pose estimation algorithm can be used to detect 2D key points (such as shoulders, elbows, knees, etc.) corresponding to the target object in each video frame. Then, these 2D key points are associated in time series using optical flow or feature matching to form a trajectory. Next, a perspective n-point algorithm can be used to match the detected 2D key points with corresponding points in the 3D skeletal model. Since the target object is in motion, this matching is dynamic. By minimizing the reprojection error (i.e., the difference between the position of the 3D point projected onto the 2D plane and the actual position of the detected 2D point), the camera pose, such as the camera rotation matrix and translation vector, relative to the previous frame can be solved.

[0094] To improve the accuracy of camera pose determination, kinematic constraints can be incorporated. For example, for the human body, stride length is finite, and arm swing has a specific frequency. These constraints can be used as part of the loss function to help filter out erroneous pose estimates, making the results smoother and more consistent with physical laws.

[0095] Since the preceding process uses neural networks to predict body movements in both the world coordinate system and the camera coordinate system, the initial camera pose can be derived using the formulas described above. This pose estimation method naturally combines the acquisition device parameter information directly obtained from visual odometry with human motion to obtain a camera pose consistent with the camera movement of the target video, thus improving the efficiency of camera pose acquisition.

[0096] Optionally, there are several other ways to determine the camera pose of the target video based on body motion information. For example, the type of camera motion amplitude can be determined based on the acquisition device parameter information; and the camera pose can be determined based on the type of motion amplitude and body motion information.

[0097] The motion amplitude type can be information describing the motion amplitude of the camera.

[0098] There are several ways to determine the camera pose based on the type of motion amplitude and body motion information. For example, if the type of motion amplitude is stationary, the camera pose is determined based on the body motion information corresponding to the first video frame in the target video; if the type of motion amplitude is small-amplitude motion, target video frames are sampled in the target video at preset frame intervals, and the camera pose is determined based on the body motion information corresponding to the target video frames; if the type of motion amplitude is large-amplitude motion, the camera pose is determined based on the body motion information corresponding to each video frame in the target video.

[0099] Therefore, when the camera is stationary, the camera pose calculated based on the body motion information of the first frame can be determined as the camera pose in the target video. When the camera is moving slightly, to reduce computational resource consumption, samples can be taken intermittently in the target video, and the camera pose calculated based on the body motion information of the sampled target video frames can be determined as the camera pose in the target video. When the camera is moving significantly, to improve the accuracy of the camera pose, the camera pose calculated based on the body motion information of each video frame in the target video can be determined as the camera pose in that video frame, thus obtaining the camera pose in each video frame.

[0100] To preserve large camera movements while reducing unnecessary small-amplitude shakes, this application employs a hierarchical optimization of camera motion. Specifically, based on the camera pose, two criteria are set to determine the magnitude of camera motion, categorizing initial camera motion into three types: stationary, small-amplitude, and large-amplitude. For stationary cameras, the camera pose of the first frame can be directly used as the camera pose for the entire video. For cameras with small-amplitude motion, sampling is performed at intervals of a certain number of frames, and then the camera pose corresponding to the sampled target video frames is segmented and smoothly interpolated. For cameras with large-amplitude motion, the camera pose is calculated for each video frame. For example, a stationary camera is defined as one where the variance of the orientation change angle in the acquisition device parameter information is less than 0.0005 degrees and the movement range is less than 2 meters; a small-amplitude motion camera is defined as one where the angle variance is between 0.0005 degrees and 0.01 degrees and the movement range is less than 2.2 meters; all others can be considered large-amplitude motion cameras. Thus, the hierarchical optimization scheme provided by this application helps reduce unnecessary camera shakes and computational resource consumption, improving the efficiency of camera pose calculation.

[0101] In one embodiment, please refer to Figure 4 , Figure 4 This is a schematic diagram illustrating the specific process of a video motion extraction method provided in this application embodiment. This application embodiment offers a method for extracting full-body character movements and camera motion from dynamic viewpoint video. The process includes: extracting body movements from the dynamic viewpoint video; fusing body information to extract hand movements from the dynamic viewpoint video and updating arm rotation; physics-guided multi-stage character motion post-processing; and calculating and adaptively post-processing the camera pose. The framework design of this method achieves hierarchical modeling of the body, hands, and camera, and closely links each module, improving the overall accuracy and robustness of video motion extraction.

[0102] Therefore, users only need to input any captured video, and based on the embodiments of this application, 15 seconds of human body movements and camera motion information can be extracted within half a minute. Specifically, the embodiments of this application can predict accurate full-body movements. On the one hand, it can obtain stable body movements under dynamic perspectives. On the other hand, by introducing body information in two stages—hand key point estimation and hand movement prediction—it takes into account the flexibility and smoothness of the wrist joint and effectively reduces the impact of dynamic perspectives on hand movement estimation, significantly improving the expressiveness of hand movements. In addition, the embodiments of this application do not use a scheme that predicts the body and hands together, which can avoid interference with body movement prediction when there is a lot of hand noise. Through physical-guided character movement post-processing, the naturalness of the center of gravity and feet can be enhanced. On the one hand, a multi-stage optimization design is adopted to minimize foot slippage and reduce the unnatural performance of feet easily sticking to the ground. On the other hand, physical correction is introduced in the center of gravity optimization stage to ensure that the movement conforms to the physical trajectory even when in the air. Furthermore, the embodiments of this application infer camera motion by human body movements, which can naturally combine the camera motion information directly obtained by visual odometry with human body movements to obtain camera pose consistent with the camera movement effect of the input video, and adopt a hierarchical optimization strategy to improve the estimation accuracy of different types of camera motion.

[0103] As can be seen from the above, the embodiments of this application identify the detection area information, body posture features, and acquisition device parameter information of the target object in the target video; extract body movements from the target video based on the detection area information, body posture features, and acquisition device parameter information to obtain the body movement information of the target object in the target video; extract hand movements from the target video based on the body movement information to obtain the hand movement information of the target object and the rotation information of the target joints associated with the hand; update the body movement information based on the rotation information of the target joints; optimize the physical rationality of the updated body movement information; and generate the full-body movement information of the target object in the target video based on the optimized body movement information and hand movement information. Therefore, by using the target object's detection area information, body posture features, and acquisition device parameter information, the body motion information of the target object in the target video can be predicted. Then, by combining the body motion information, the hand motion information and the rotation information of the target joints associated with the hand can be accurately extracted from the target video. The body motion information is updated based on the rotation information of the target joints, and the updated body motion information is optimized for physical rationality. Based on the optimized body motion information and hand motion information, accurate and physically consistent full-body motion information can be extracted from the target video, improving the accuracy of video motion capture and further improving the efficiency of video motion extraction.

[0104] To better implement the above methods, embodiments of the present invention also provide a video motion extraction device, which can be integrated into an electronic device, such as a server.

[0105] For example, such as Figure 5 The diagram shown is a structural schematic of a video motion extraction device provided in an embodiment of this application. The video motion extraction device may include a recognition unit 201, a first extraction unit 202, a second extraction unit 203, an update unit 204, an optimization unit 205, and a generation unit 206, as follows: The recognition unit 201 is used to identify the detection area information, body posture features and acquisition device parameter information corresponding to the target object in the target video. The first extraction unit 202 is used to extract body movements from the target video based on detection area information, body posture features and acquisition device parameter information, to obtain the body movement information of the target object in the target video. The second extraction unit 203 is used to extract hand movements from the target video based on body movement information, so as to obtain the hand movement information of the target object and the rotation information of the target joint associated with the hand. The update unit 204 is used to update the body motion information based on the rotation information of the target joint; Optimization unit 205 is used to optimize the physical rationality of the updated body motion information; The generation unit 206 is used to generate full-body motion information of the target object in the target video based on the optimized body motion information and hand motion information.

[0106] In some embodiments, the second extraction unit is configured to: Based on body posture features, the hand region of the target object and the orientation of the target object's palm are determined in the video frame images of the target video; Based on the hand area and palm orientation, predict the key points of the target object's hand; Based on key hand points, hand regions, and body movement information, the system identifies the hand movement information of the target object and the rotation information of the target joints associated with the hand.

[0107] In some embodiments, the above-mentioned identification of the target object's hand movement information and the rotation information of the target joints associated with the hand based on hand key points, hand regions, and body movement information is specifically used for: Extract the position information of bones in preset parts of the target object from the body motion information, and normalize the position information of the bones based on the position of the root bones corresponding to the bones. Normalization of key hand points is performed based on the hand region; Based on the normalized bone position information and hand key points, predict the hand movement information of the target object and the rotation information of the target joints associated with the hand.

[0108] In some embodiments, the above-mentioned method of determining the hand region of the target object in the video frame image of the target video based on body posture features is specifically used for: Based on body posture characteristics, determine the size of the target part in the target object; Based on the size ratio between the target part and the hand, and the size of the part, determine the size of the target part of the hand of the target object; Based on the size of the target part, the hand region including the hand of the target object is determined in the video frame image of the target video.

[0109] In some embodiments, the video motion extraction device is further configured to: Based on the detection area information, body posture features, and data acquisition device parameter information, the ground contact probability of multiple candidate ground contact joints corresponding to body movement information is predicted. Optimization unit, used for: The target ground contact joint is determined from the candidate ground contact joints based on the ground contact probability. The updated body motion information is optimized based on the target ground-contact joint; Based on the optimized body and hand motion information, full-body motion information of the target object in the target video is generated.

[0110] In some embodiments, the above-described determination of the target ground contact joint among candidate ground contact joints based on the ground contact probability is specifically used for: Based on the ground contact correlation parameters of the candidate ground contact joints, the ground contact probability is corrected to obtain the target ground contact probability. The ground contact correlation parameters include at least one of vertical height and velocity. The target contact joint is determined from the candidate contact joints based on the target contact probability.

[0111] In one embodiment, the above-mentioned optimization processing of the updated body motion information based on the target ground-contact joint is specifically used for: If the target object is in a ground-touching state, the center of gravity position of the target object is corrected based on at least one first loss function. The first loss function is used to constrain the motion of the target ground-touching joint and to constrain the change in the center of gravity position of the target object. If the target object is not in a ground-touching state, the center of gravity position of the target object is corrected based on the second loss function, which is used to constrain the center of gravity trajectory of the target object. The body motion information is optimized based on the corrected center of gravity position.

[0112] In some embodiments, the video motion extraction device is further configured to: Based on the center of gravity position of the target object in the body motion information, the trajectory of the center of gravity of the target object is parabolic and the center of gravity trajectory curve is obtained. The second loss function is determined based on the distance between the target object's center of gravity and its trajectory curve.

[0113] In some embodiments, the video motion extraction device is further configured to: Based on body motion information, the camera pose of the corresponding camera in the target video is determined. The camera pose is used to reconstruct the camera movement effect in the target video based on the full-body motion information.

[0114] In some embodiments, the method of determining the camera pose of the camera corresponding to the target video based on body motion information is specifically used for: Determine the type of camera motion amplitude based on the parameter information collected from the equipment; The camera pose is determined based on the type of motion amplitude and body movement information.

[0115] In some embodiments, determining the camera pose based on the type of motion amplitude and body movement information is specifically used for: If the motion amplitude type is stationary, the camera pose is determined based on the body motion information corresponding to the first video frame in the target video. If the motion type is small motion, the target video frame is sampled in the target video based on the preset frame interval, and the camera pose is determined based on the body motion information corresponding to the target video frame. If the motion type is large-amplitude motion, the camera pose is determined based on the body motion information corresponding to each video frame in the target video.

[0116] In practice, each of the above units can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units, please refer to the previous method embodiments, which will not be repeated here.

[0117] As can be seen from the above, in this embodiment of the application, the identification unit 201 identifies the detection area information, body posture features, and acquisition device parameter information of the target object in the target video; the first extraction unit 202 extracts body movements from the target video based on the detection area information, body posture features, and acquisition device parameter information to obtain the body movement information of the target object in the target video; the second extraction unit 203 extracts hand movements from the target video based on the body movement information to obtain the hand movement information of the target object and the rotation information of the target joint associated with the hand; the update unit 204 updates the body movement information based on the rotation information of the target joint; the optimization unit 205 optimizes the physical rationality of the updated body movement information; and the generation unit 206 generates the full-body movement information of the target object in the target video based on the optimized body movement information and hand movement information. Therefore, by using the target object's detection area information, body posture features, and acquisition device parameter information, the body motion information of the target object in the target video can be predicted. Then, by combining the body motion information, the hand motion information and the rotation information of the target joints associated with the hand can be accurately extracted from the target video. The body motion information is updated based on the rotation information of the target joints, and the updated body motion information is optimized for physical rationality. Based on the optimized body motion information and hand motion information, accurate and physically consistent full-body motion information can be extracted from the target video, improving the accuracy of video motion capture and further improving the efficiency of video motion extraction.

[0118] This application also provides an electronic device, such as... Figure 6 The diagram shows a schematic representation of the structure of an electronic device according to an embodiment of this application. This electronic device can be a terminal or a server. Specifically: The electronic device 300 includes a processor 301 with one or more processing cores, a memory 302 with one or more computer-readable storage media, and a computer program stored in the memory 302 and executable on the processor. The processor 301 and the memory 302 are electrically connected. Those skilled in the art will understand that the electronic device structure shown in the figures does not constitute a limitation on the electronic device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0119] The processor 301 is the control center of the electronic device 300. It connects various parts of the electronic device 300 through various interfaces and lines. By running or loading software programs and / or modules stored in the memory 302, and calling data stored in the memory 302, it performs various functions of the electronic device 300 and processes data, thereby monitoring the electronic device 300 as a whole.

[0120] In this embodiment, the processor 301 in the electronic device 300 loads the instructions corresponding to the processes of one or more applications into the memory 302 according to the following steps, and the processor 301 runs the applications stored in the memory 302 to realize various functions: The system identifies the detection area information, body posture features, and acquisition device parameter information of the target object in the target video. Based on the detection area information, body posture features and acquisition device parameter information, body motion is extracted from the target video to obtain the body motion information of the target object in the target video. Hand movements are extracted from the target video based on body motion information to obtain the hand movement information of the target object and the rotation information of the target joints associated with the hand. Update body motion information based on the rotational information of the target joint; Optimize the physical plausibility of the updated body movement information; Based on the optimized body and hand motion information, full-body motion information of the target object in the target video is generated.

[0121] This solution identifies the detection area information, body posture features, and acquisition device parameters of the target object in the target video. Based on the detection area information, body posture features, and acquisition device parameters, it extracts body movements from the target video to obtain the target object's body movement information. Based on the body movement information, it extracts hand movements from the target video to obtain the target object's hand movement information and the rotation information of the target joints associated with the hand. It updates the body movement information based on the rotation information of the target joints. It optimizes the updated body movement information for physical plausibility. Based on the optimized body movement information and hand movement information, it generates the target object's full-body movement information in the target video. Therefore, by using the target object's detection area information, body posture features, and acquisition device parameter information, the body motion information of the target object in the target video can be predicted. Then, by combining the body motion information, the hand motion information and the rotation information of the target joints associated with the hand can be accurately extracted from the target video. The body motion information is updated based on the rotation information of the target joints, and the updated body motion information is optimized for physical rationality. Based on the optimized body motion information and hand motion information, accurate and physically consistent full-body motion information can be extracted from the target video, improving the accuracy of video motion capture and further improving the efficiency of video motion extraction.

[0122] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0123] Optional, such as Figure 6 As shown, the electronic device 300 also includes: a touch display screen 303, a radio frequency circuit 304, an audio circuit 305, an input unit 306, and a power supply 307. The processor 301 is electrically connected to the touch display screen 303, the radio frequency circuit 304, the audio circuit 305, the input unit 306, and the power supply 307. Those skilled in the art will understand that... Figure 6 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0124] The touch display screen 303 can be used to display a graphical user interface (GUI) and receive operation commands generated by the user interacting with the GUI. The touch display screen 303 may include a display panel and a touch panel. The display panel can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the electronic device. These graphical user interfaces can be composed of graphics, text, icons, video, and any combination thereof. Optionally, the display panel can be configured using a liquid crystal display (LCD), organic light-emitting diode (OLED), or other similar technologies. The touch panel can be used to collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel), generate corresponding operation commands, and execute the corresponding program according to the operation commands. Optionally, the touch panel may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch location and the signal generated by the touch operation, transmitting the signal to the touch controller. The touch controller receives touch information from the touch detection device, converts it into touch point coordinates, and sends it to the processor 301. It can also receive and execute commands from the processor 301. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it transmits the information to the processor 301 to determine the type of touch event. Subsequently, the processor 301 provides corresponding visual output on the display panel based on the type of touch event. In this embodiment, the touch panel and the display panel can be integrated into the touch display screen 303 to achieve input and output functions. However, in some embodiments, the touch panel and the touch display screen 303 can be implemented as two independent components to achieve input and output functions. That is, the touch display screen 303 can also be used as part of the input unit 306 to achieve input functions.

[0125] The radio frequency circuit 304 can be used to transmit and receive radio frequency signals to establish wireless communication with network devices or other electronic devices, and to transmit and receive signals with network devices or other electronic devices.

[0126] Audio circuitry 305 can be used to provide an audio interface between a user and an electronic device via a speaker and a microphone. Audio circuitry 305 converts received audio data into electrical signals, transmits them to the speaker, and the speaker converts them into sound signals for output. Conversely, the microphone converts collected sound signals into electrical signals, which are then received by audio circuitry 305, converted back into audio data, and then processed by processor 301 before being transmitted via radio frequency circuitry 304 to, for example, another electronic device, or output to memory 302 for further processing. Audio circuitry 305 may also include an earphone jack to facilitate communication between peripheral headphones and electronic devices.

[0127] The input unit 306 can be used to receive input numbers, characters, or user characteristic information (such as fingerprints, iris, facial information, etc.), and to generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control.

[0128] Power supply 307 is used to supply power to various components of electronic device 300. Optionally, power supply 307 can be logically connected to processor 301 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. Power supply 307 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0129] although Figure 6 As not shown in the diagram, the electronic device 300 may also include a camera, sensor, wireless fidelity module, Bluetooth module, etc., which will not be described in detail here.

[0130] In the above embodiments, the descriptions of each embodiment have different focuses. Parts not described in detail in a particular embodiment can be found in the relevant descriptions of other embodiments. It should be noted that the electronic device provided in this application embodiment and the video motion extraction method described in the above embodiments belong to the same concept, and its specific implementation process is detailed in the above method embodiments, and will not be repeated here.

[0131] As can be seen from the above, the electronic device provided in this application embodiment can identify the detection area information, body posture features, and acquisition device parameter information corresponding to the target object in the target video; extract body movements from the target video based on the detection area information, body posture features, and acquisition device parameter information to obtain the body movement information of the target object in the target video; extract hand movements from the target video based on the body movement information to obtain the hand movement information of the target object and the rotation information of the target joint associated with the hand; update the body movement information based on the rotation information of the target joint; optimize the physical rationality of the updated body movement information; and generate the full-body movement information of the target object in the target video based on the optimized body movement information and hand movement information. Therefore, by using the target object's detection area information, body posture features, and acquisition device parameter information, the body motion information of the target object in the target video can be predicted. Then, by combining the body motion information, the hand motion information and the rotation information of the target joints associated with the hand can be accurately extracted from the target video. The body motion information is updated based on the rotation information of the target joints, and the updated body motion information is optimized for physical rationality. Based on the optimized body motion information and hand motion information, accurate and physically consistent full-body motion information can be extracted from the target video, improving the accuracy of video motion capture and further improving the efficiency of video motion extraction.

[0132] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by a computer program, or by a computer program controlling related hardware. The computer program can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0133] Therefore, embodiments of this application provide a computer-readable storage medium storing a computer program that can be loaded by a processor to execute the steps of any of the video motion extraction methods provided in embodiments of this application. For example, the computer program can execute the following steps: The system identifies the detection area information, body posture features, and acquisition device parameter information of the target object in the target video. Based on the detection area information, body posture features and acquisition device parameter information, body motion is extracted from the target video to obtain the body motion information of the target object in the target video. Hand movements are extracted from the target video based on body motion information to obtain the hand movement information of the target object and the rotation information of the target joints associated with the hand. Update body motion information based on the rotational information of the target joint; Optimize the physical plausibility of the updated body movement information; Based on the optimized body and hand motion information, full-body motion information of the target object in the target video is generated.

[0134] This solution identifies the detection area information, body posture features, and acquisition device parameters of the target object in the target video. Based on the detection area information, body posture features, and acquisition device parameters, it extracts body movements from the target video to obtain the target object's body movement information. Based on the body movement information, it extracts hand movements from the target video to obtain the target object's hand movement information and the rotation information of the target joints associated with the hand. It updates the body movement information based on the rotation information of the target joints. It optimizes the updated body movement information for physical plausibility. Based on the optimized body movement information and hand movement information, it generates the target object's full-body movement information in the target video. Therefore, by using the target object's detection area information, body posture features, and acquisition device parameter information, the body motion information of the target object in the target video can be predicted. Then, by combining the body motion information, the hand motion information and the rotation information of the target joints associated with the hand can be accurately extracted from the target video. The body motion information is updated based on the rotation information of the target joints, and the updated body motion information is optimized for physical rationality. Based on the optimized body motion information and hand motion information, accurate and physically consistent full-body motion information can be extracted from the target video, improving the accuracy of video motion capture and further improving the efficiency of video motion extraction.

[0135] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0136] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0137] Since the computer program stored in the computer-readable storage medium can execute the steps of any of the video action extraction methods provided in the embodiments of this application, the beneficial effects that any of the video action extraction methods provided in the embodiments of this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.

[0138] According to one aspect of this application, a computer program product is provided, comprising a computer program stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, causing the electronic device to perform the methods provided in the various optional implementations of the above embodiments.

[0139] The foregoing has provided a detailed description of a video motion extraction method, apparatus, storage medium, and electronic device provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A video motion extraction method, characterized in that, include: In the target video, the detection area information, body posture features, and acquisition device parameter information corresponding to the target object in the target video are identified; Based on the detection area information, the body posture features, and the acquisition device parameter information, body motion is extracted from the target video to obtain the body motion information of the target object in the target video. Based on the body motion information, hand motion is extracted from the target video to obtain the hand motion information of the target object and the rotation information of the target joint associated with the hand; The body motion information is updated based on the rotation information of the target joint; The updated body motion information is then optimized for physical plausibility. Based on the optimized body motion information and hand motion information, full-body motion information of the target object in the target video is generated.

2. The video motion extraction method as described in claim 1, characterized in that, The step of extracting hand movements from the target video based on the body movement information to obtain the hand movement information of the target object and the rotation information of the target joints associated with the hand includes: Based on the body posture features, the hand region of the target object and the orientation of the target object's palm are determined in the video frame image of the target video; Based on the hand region and the palm orientation, predict the key points of the target object's hand; Based on the key hand points, the hand region, and the body movement information, the hand movement information of the target object and the rotation information of the target joints associated with the hand are identified.

3. The video motion extraction method as described in claim 2, characterized in that, The step of identifying the hand movement information of the target object and the rotation information of the target joints associated with the hand based on the hand key points, the hand area, and the body movement information includes: The position information of bones in a preset part of the target object is extracted from the body movement information, and the position information of the bones is normalized based on the position of the root bone corresponding to the bone. The key points of the hand are normalized based on the hand region; Based on the normalized bone position information and the hand key points, predict the hand movement information of the target object and the rotation information of the target joints associated with the hand.

4. The video motion extraction method as described in claim 2, characterized in that, The step of determining the hand region of the target object in the video frame image of the target video based on the body posture features includes: Based on the body posture characteristics, the size of the target part in the target object is determined; Based on the size ratio between the target part and the hand, and the size of the part, the size of the target part of the hand of the target object is determined; Based on the size of the target part, a hand region including the hand of the target object is determined in the video frame image of the target video.

5. The video motion extraction method as described in claim 1, characterized in that, Also includes: Based on the detection area information, the body posture features, and the acquisition device parameter information, the ground contact probability of multiple candidate ground-contact joints corresponding to the body movement information is predicted. The physical plausibility optimization of the updated body movement information includes: The target ground contact joint among the candidate ground contact joints is determined based on the ground contact probability. The updated body motion information is optimized based on the target ground-contact joint.

6. The video motion extraction method as described in claim 5, characterized in that, The step of determining the target contact joint among the candidate contact joints based on the contact probability includes: Based on the ground contact correlation parameters of the candidate ground contact joints, the ground contact probability is corrected to obtain the target ground contact probability. The ground contact correlation parameters include at least one of vertical height and velocity. The target contact joint is determined from the candidate contact joints based on the target contact probability.

7. The video motion extraction method as described in claim 5, characterized in that, The optimization process based on the target ground-contact joint to update the body motion information includes: If the target object is in a ground-touching state, the center of gravity position of the target object is corrected based on at least one first loss function. The first loss function is used to constrain the motion of the target ground-touching joint and to constrain the change in the center of gravity position of the target object. If the target object is not in a ground-touching state, the center of gravity position of the target object is corrected based on the second loss function, which is used to constrain the center of gravity trajectory of the target object; The body motion information is optimized based on the corrected center of gravity position.

8. The video motion extraction method as described in claim 7, characterized in that, Before correcting the centroid position of the target object based on the second loss function, the method further includes: Based on the center of gravity position of the target object in the body motion information, a parabolic fitting is performed on the center of gravity trajectory of the target object to obtain the center of gravity trajectory curve. The second loss function is determined based on the distance between the centroid of the target object and the centroid trajectory curve.

9. The video motion extraction method according to any one of claims 1 to 8, characterized in that, The method further includes: Based on the body motion information, the camera pose of the camera corresponding to the target video is determined, and the camera pose is used to reconstruct the camera movement effect in the target video based on the full-body motion information.

10. The video motion extraction method as described in claim 9, characterized in that, Determining the camera pose of the camera corresponding to the target video based on the body motion information includes: The motion amplitude type of the camera is determined based on the parameter information of the acquisition device; The camera pose is determined based on the motion amplitude type and the body movement information.

11. The video motion extraction method as described in claim 10, characterized in that, Determining the camera pose based on the motion amplitude type and the body movement information includes: If the motion amplitude type is static, the camera pose is determined based on the body motion information corresponding to the first video frame in the target video. If the motion amplitude type is a small motion type, target video frames are sampled in the target video based on a preset frame interval, and the camera pose of the camera is determined based on the body motion information corresponding to the target video frames. If the motion amplitude type is a large motion type, the camera pose of the camera is determined based on the body motion information corresponding to each video frame in the target video.

12. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program that, when executed by the processor, causes the processor to perform the steps of any of the methods described in claims 1 to 11.

13. A computer-readable storage medium, characterized in that, It includes a computer program that, when run on an electronic device, causes the electronic device to perform the steps of any of the methods described in claims 1 to 11.