Video-based digital human motion capture and motion generation method and device, and medium
Through the video-based digital human motion capture method, human body and hand detection algorithms are used to combine SMPL and SMPL-X models for 3D reconstruction to generate skeletal motion files, and the problems of high-cost and repetitive work in the existing technology are solved through cross-skeletal motion migration technology, achieving low-cost and natural and smooth motion migration.
Patent Information
- Application Number
- CN202510594526.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-08-08
AI Technical Summary
The existing digital human motion capture technology is expensive, and bone adjustments have caused animators to refine animations, increase workload and unnecessary repetitive work.
Through a video-based digital human motion capture method, image frames are extracted using human body and hand detection algorithms, combined with SMPL and SMPL-X models for 3D human body reconstruction, generate skeletal motion files, and migrate the action files to the target human model through cross-skeletal motion migration technology, and optimize the motion fluency using reverse kinematics and baking techniques.
It reduces the cost of digital human motion capture, solves the problem of re-animation caused by bone adjustment, and realizes natural and smooth migration of cross-skeleton movement.
Smart Images

Figure CN120451350A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of machine vision technology, and in particular to a method, device, and medium for capturing and generating digital human motion based on video. Background Art
[0002] Motion capture is often required when creating digital human movements. Currently, the most widely used methods are inertial and optical. However, both require expensive equipment and locations, making the costs prohibitive. Furthermore, the captured data typically requires refined processing by professional animators. However, if the skeleton needs to be readjusted during digital human creation, such as adjusting the position or length of bones, all the animators' previous fine-tuning efforts will be lost, forcing them to re-fine the digital human's movements based on the new skeleton. This undoubtedly adds unnecessary and repetitive work to the animators, and any change to the skeleton affects all movements, significantly increasing the animator's workload.
[0003] Therefore, the present invention proposes a video-based digital human motion capture and motion generation solution to address the problems existing in the prior art. Summary of the Invention
[0004] This application proposes a video-based digital human motion capture and motion generation solution, which not only reduces the cost of digital human body motion capture, but also solves the problem of needing to re-animate due to skeleton adjustment.
[0005] In a first aspect, the present application provides a method for capturing digital human motion and generating motion based on video, comprising:
[0006] Get a video containing character actions;
[0007] Extracting a body region image frame and a hand region image frame corresponding to the person based on the video;
[0008] Reconstructing a human body based on the body region image frame and the hand region image frame to obtain a 3D human body model;
[0009] Human skeleton information is acquired based on the 3D human body model, and a skeleton motion file is generated based on the human skeleton information.
[0010] In some embodiments, extracting the body region image and the hand region image corresponding to the person based on the video includes:
[0011] Based on a human body detection algorithm, detecting and determining the position of a human body from each frame of the video;
[0012] Based on a hand key point detection algorithm, a hand position is selected at the position of the human body;
[0013] determining a hand region frame and a body region frame based on the position of the hand and the position of the human body;
[0014] A body region image frame and a hand region image frame corresponding to the person are obtained based on the body region frame and the hand region frame selected in each frame image.
[0015] In some embodiments, the above-mentioned reconstructing the human body based on the body region image frame and the hand region image frame to obtain a 3D human body model includes:
[0016] Predicting shape parameters and posture parameters of the body based on the body region image frame and the first neural network model, and inputting the shape parameters and posture parameters of the body into the SMPL model to obtain a 3D body model;
[0017] Predicting shape parameters and posture parameters of the hand based on the hand region image frame and the second neural network model, and inputting the shape parameters and posture parameters of the hand into the SMPL-X model to obtain a 3D finger model;
[0018] The 3D body model and the 3D finger model are spliced together to obtain the 3D human body model.
[0019] Furthermore, the joints of the 3D body model include wrist joints but do not include palm joints, and the joints of the 3D finger model include wrist joints and palm joints. Accordingly, the above-mentioned splicing of the 3D body model and the 3D finger model to obtain the 3D human body model includes: according to the corresponding relationship between the wrist joints of the 3D body model and the wrist joints of the 3D finger model, connecting the palm joints of the 3D finger model to the wrist joints of the 3D body model.
[0020] In some embodiments, the steps of obtaining human skeleton information based on the 3D human body model and generating a skeleton motion file based on the human skeleton information include:
[0021] Determining a human skeleton containing semantics based on a pairwise connection relationship of 3D key points of the 3D human body model, and obtaining a topological structure of the skeleton based on the human skeleton and its semantics;
[0022] Calculating the rotation matrix of each human bone according to the posture parameters of the 3D human body model to obtain the motion value of the bone;
[0023] The skeleton motion file is formed according to the topological structure of the skeleton and the motion value.
[0024] In some embodiments, the method further includes: migrating the skeletal motion file to a target human body model, wherein the skeletal motion file is used to drive the target human body model to perform actions consistent with actions of the character in the video.
[0025] Furthermore, migrating the skeleton motion file to the target human body model includes:
[0026] Using the topological structure and the motion value of the skeleton in the skeleton motion file as data to be migrated;
[0027] Matching the skeleton in the data to be migrated with the skeleton of the target human body model based on the semantics of the skeleton to obtain a mapping relationship;
[0028] Based on the mapping relationship, the posture of the corresponding skeleton of the target human body model is adjusted according to the motion value of the skeleton in the data to be migrated.
[0029] Furthermore, after migrating the skeleton motion file to the target human body model, the method further includes:
[0030] a step of correcting the data transferred to the target human body model, and / or a step of baking the data transferred to the target human body model.
[0031] In a second aspect, the present application provides a computer device comprising a processor and a storage device, wherein the storage device is suitable for storing multiple program codes, and the program codes are suitable for being loaded and run by the processor to execute the method described in any one of the technical solutions of the above-mentioned video-based digital human motion capture and motion generation method.
[0032] In a third aspect, the present application provides a computer-readable storage medium storing a plurality of program codes, wherein the program codes are suitable for being loaded and run by a processor to execute the method described in any one of the technical solutions of the above-mentioned video-based digital human motion capture and motion generation method.
[0033] One or more of the above-mentioned technical solutions of the present application have at least one or more of the following beneficial effects: the present application proposes human body reconstruction and motion capture based on video streams, which reduces the cost of capturing digital human body motions; the present application solves the problem of needing to re-animate due to skeletal adjustment. The skeletal motion file obtained based on the method of the present application can be migrated across skeletal motions to different target human models, that is, it can drive the motion of digital humans with different skeletons. By adding inverse kinematics and baking, the motion migrated to the new skeleton has no local sliding, and the resulting motion is more natural and smooth. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] The disclosure of this application will be more easily understood with reference to the accompanying drawings. Those skilled in the art will readily appreciate that these drawings are for illustrative purposes only and are not intended to limit the scope of protection of this application. Furthermore, similar numbers in the figures represent similar components, where:
[0035] Figure 1 This is a schematic diagram of the implementation process of a method for video-based digital human motion capture and motion generation according to an embodiment of the present application;
[0036] Figure 2 Schematic diagram of the effect of the original skeleton and the target skeleton before motion migration in an embodiment of the present application;
[0037] Figure 3 is based on Figure 2 Schematic diagram of the effect of the skeleton after movement migration;
[0038] Figure 4 This is a schematic block diagram of a video-based digital human motion capture and motion generation system according to an embodiment of the present application. DETAILED DESCRIPTION
[0039] Some embodiments of the present application are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present application and are not intended to limit the scope of protection of the present application.
[0040] In the description of this application, "module" and "processor" may include hardware, software, or a combination of the two. A module may include hardware circuits, various suitable sensors, communication ports, and memory, and may also include software components, such as program code, or a combination of software and hardware. The processor may be a central processing unit, a microprocessor, an image processor, a digital signal processor, or any other suitable processor. The processor has data and / or signal processing functions. The processor may be implemented in software, hardware, or a combination of the two. The term "A and / or B" represents all possible combinations of A and B, such as only A, only B, or A and B. The term "at least one A or B" or "at least one of A and B" has a similar meaning to "A and / or B" and may include only A, only B, or A and B. The singular terms "one" and "the" may also include plural forms.
[0041] First, the nouns involved in this application are explained.
[0042] SMPL model: It is a parametric 3D human body model that represents the human body as a low-dimensional parameter space and describes the human body through posture parameters and shape parameters. This model can capture a wide range of morphological changes of the human body and express it in an efficient mathematical form.
[0043] SMPL-X model: It is an extension of the SMPL model, which enhances the modeling capabilities of men, women, and children and introduces additional parameters to cover facial and hand poses.
[0044] Super-resolution: It is an image processing technology that aims to improve the resolution of the original image. It is mainly used to reconstruct the corresponding high-resolution image from the observed low-resolution image.
[0045] bvh: A file format that records bone topology and motion data.
[0046] Inverse Kinematics (IK): A commonly used algorithm in 3D skeletal animation, IK can deduce the rotation of other joints (joints connected to the free end joint) given the position of the free end joint. It is often used to derive the rotation of knee and elbow joints. Inverse kinematics, also known as inverse kinematics, is a technique that infers the position and transformation of a parent bone from the position and transformation of its child bones.
[0047] The preferred embodiments of the present application are described below in conjunction with the drawings in the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application and are not used to limit the present application. In addition, the embodiments and features in the embodiments of the present application can be combined with each other if there is no conflict.
[0048] The digital human motion capture and motion generation method of the present application captures the motions of characters in a video based on a deep learning model, determines the skeletal topology and motion values of the character in each frame of the video, generates a skeletal motion file corresponding to the video, and enables the target human body model to execute the skeletal motion file and display motions corresponding to the motions of the character in the video through cross-skeletal motion migration.
[0049] like Figure 1 As shown in FIG, the main steps of the method for capturing and generating digital human motion based on video in one embodiment of the present application are as follows: Figure 1 The process shown mainly includes the following steps: step S11 to step S14:
[0050] Step S11: obtaining a video containing character actions;
[0051] In this embodiment, the video can be understood as a plurality of frames of images containing continuous character movements.
[0052] Step S12: extracting body region image frames and hand region image frames corresponding to the person based on the video;
[0053] In this embodiment, the existing human body detection algorithm and finger key point detection algorithm can be used to detect each frame image in the video to extract the body area frame and the hand area frame, wherein the image within the body area frame on each frame image is the body area image frame, and the image within the hand area frame on each frame image is the hand area image frame.
[0054] Step S13: reconstructing the human body based on the body region image frame and the hand region image frame to obtain a 3D human body model;
[0055] In this embodiment, a 3D body model can be obtained by performing 3D reconstruction of the human body region based on the SMPL model and the body region image frames, and a 3D finger model can be obtained by performing 3D reconstruction of the hand region based on the SMPL-X model and the hand region image frames. A complete 3D human body model can be obtained by splicing the 3D body model and the 3D finger model.
[0056] Step S14: acquiring human skeleton information based on the 3D human body model, and generating a skeleton motion file based on the human skeleton information;
[0057] In this embodiment, the topological structure of the skeleton can be determined based on the 3D key points of the 3D human body model, and the rotation matrix of each bone in the topological structure can be calculated based on the posture parameters of the 3D human body model. The skeleton action file can be formed according to the topological structure of the skeleton and the rotation matrix of each bone, wherein the skeleton action file is a bvh file.
[0058] In some embodiments, step S15 may be further included after step S14.
[0059] Step S15: Migrating the skeleton motion file to the target human body model.
[0060] The skeleton action file is used to drive the target human body model to perform actions consistent with the actions of the character in the video.
[0061] In this embodiment, the designed cross-skeleton motion migration method can match bones with the same semantics, such as arm to arm, thigh to thigh, etc., interactively adjust the postures of different bones, and use the IK algorithm to suppress the sliding of the feet and adjust the movement of the arms, so that the movement migrated to the new skeleton is more natural; the movement after baking the migration makes the movement migrated to the new skeleton smoother and more fluent.
[0062] In a specific implementation of the above-mentioned step S11, the step of video preprocessing may be further included after obtaining the video containing the action of the person. Specifically, the step may be to use the basic white balance algorithm and wide dynamic algorithm in digital image processing so that the image in the video will not have obvious color cast or unclear mixture of dark and bright areas after processing.
[0063] In a specific implementation of the above step S12, extracting the body region image frame and the hand region image frame corresponding to the person based on the video may specifically include the following steps 121 to 124:
[0064] Step 121: Based on a human body detection algorithm, detecting and determining the position of the human body from each frame of the video;
[0065] Specifically, the human body detection algorithm can be a method of human body key point detection or posture estimation in the existing technology. Human body key point detection refers to detecting the key point positions of the human body, such as the head, shoulders, wrists, knees, etc. from an image or video. It is usually trained and inferred using a deep learning model, and the position of each key point is predicted by regression or classification, and then the position of the human body is determined based on the position of the key points.
[0066] Step 122: Based on a hand key point detection algorithm, select a hand position in the position of the human body;
[0067] Specifically, the hand key point detection algorithm may be a hand recognition algorithm based on key point detection in the prior art, which can detect the position of the hand from an image or video.
[0068] Generally speaking, models that can be applied to human keypoint detection tasks can also be applied to hand keypoint detection tasks. In this embodiment, the hand keypoint detection algorithm and the human body detection algorithm can be implemented based on the same deep learning model, but they differ in the number of keypoints and whether the objects within them need to be framed. For example, the human body generally has 17 keypoints, while the hand has 21 keypoints. Based on the detected positions of the hand keypoints, a hand region frame can be obtained, while based on the detected positions of the human body keypoints, a body region frame can be obtained.
[0069] Step 123: determining a hand region frame and a body region frame based on the position of the hand and the position of the human body;
[0070] Step 124: Obtain a body region image frame and a hand region image frame corresponding to the person based on the body region frame and the hand region frame selected in each frame of image.
[0071] In a specific implementation of the above step S13, the step of reconstructing the human body based on the body region image frame and the hand region image frame to obtain a 3D human body model specifically includes the following steps 131 to 133:
[0072] Step 131: Predicting shape parameters and posture parameters of the body based on the body region image frame and the first neural network model, and inputting the shape parameters and posture parameters of the body into the SMPL model to obtain a 3D body model;
[0073] Step 132: Predicting shape parameters and posture parameters of the hand based on the hand region image frame and the second neural network model, and inputting the shape parameters and posture parameters of the hand into the SMPL-X model to obtain a 3D finger model;
[0074] Step 133: splicing the 3D body model and the 3D finger model to obtain the 3D human body model.
[0075] It is understandable that the parameters of the spliced 3D human body model include shape parameters and posture parameters. More specifically, the shape parameters and posture parameters of the 3D human body model can be used to reflect the shape of the body and hands, that is, the shape of the human body and the posture of all key points of the human body.
[0076] The architecture of the first neural network model or the second neural network model in this embodiment includes an encoder and a regressor. In the above steps 131 and 132, the neural network-based encoder can encode the input data to obtain a high-dimensional feature vector, and the neural network-based regressor can regress the high-dimensional feature vector into smpl parameters or smpl-x parameters. It should be understood that smpl parameters refer to parameters of the SMPL model, which may include shape parameters (or body parameters) and posture parameters, and smpl-x parameters refer to parameters of the SMPL-X model, which may include shape parameters, posture parameters, and expression parameters. In the embodiment of the present application, the SMPL model is mainly used for 3D reconstruction of the human body. The smpl parameters predicted and output by the first neural network model are specifically the shape parameters and posture parameters of the body. The SMPL-X model is mainly used for 3D reconstruction of the hand. The smpl-x parameters predicted and output by the second neural network model are specifically the shape parameters and posture parameters of the hand.
[0077] Furthermore, the above step 132 can also be specifically as follows: first, the hand area image frame is optimized using a super-resolution algorithm to improve the clarity of the fingers, and then the super-resolved image is encoded through a Transformer, encoded into a high-dimensional vector and then the smplx parameter is regressed.
[0078] In this embodiment, the joint points of the 3D body model reconstructed based on the SMPL model include wrist joint points but do not include palm joint points, and the joint points of the 3D finger model reconstructed based on the SMPL-X model include wrist joint points and palm joint points; the splicing of the 3D body model and the 3D finger model to obtain the 3D human body model in the above step 133 may specifically include: according to the correspondence between the wrist joint points of the 3D body model and the wrist joint points of the 3D finger model, connecting the palm joint points of the 3D finger model to the wrist joint points of the 3D body model.
[0079] In a specific implementation of the above step S14, obtaining human skeleton information based on the 3D human body model and generating a skeleton motion file based on the human skeleton information specifically include the following steps 141 to 143:
[0080] Step 141: determining a human skeleton containing semantics according to a pairwise connection relationship of 3D key points of the 3D human body model, and obtaining a topological structure of the skeleton based on the human skeleton and its semantics;
[0081] For example, the 3D key points of the 3D human body model include the cervical vertebrae, lumbar vertebrae, elbows, etc. The 3D key points connected in pairs can obtain semantic bone information, such as the pelvic 3D key point connected to the knee 3D key point to obtain the thigh bone, the knee 3D key point connected to the ankle 3D key point to obtain the calf bone, etc.
[0082] Step 142: Calculating the rotation matrix of each human bone according to the posture parameters of the 3D human body model to obtain the motion value of the bone;
[0083] For example, the calculated rotation matrix may be recorded as the movement value of the bone.
[0084] It is understandable that during the reconstruction of the 3D human body model, there may be jumps and instabilities between the previous and next frames after the video is frame-shifted. Therefore, in this step, the motion values can also be smoothed before forming a skeleton motion file. Specifically, a smoothing algorithm (such as SmoothNet) can be used to smooth the motion values.
[0085] Step 143: forming the skeleton motion file according to the skeleton topology structure and the motion value.
[0086] For example, the topological structure and the motion value of the skeleton may be saved as a bvh file.
[0087] In a specific implementation of the above step S15, migrating the skeletal motion file to the target human body model may specifically include the following steps 151 to 153:
[0088] Step 151: using the topological structure and the motion value of the skeleton in the skeleton motion file as data to be migrated;
[0089] Step 152: matching the skeleton in the data to be migrated with the skeleton of the target human body model based on the semantics of the skeleton to obtain a mapping relationship;
[0090] For example, the bones in the data to be migrated are recorded as original bones (source), and the bones of the target human body model are recorded as target bones (target). This step specifically matches bones with the same semantics, such as the calf bones of the source match the calf bones of the target, the lumbar vertebrae bones of the source match the lumbar vertebrae bones of the target, and so on.
[0091] Step 153: Based on the mapping relationship, the posture of the corresponding skeleton of the target human body model is adjusted according to the motion value of the skeleton in the data to be migrated.
[0092] For example, the initial pose of the original skeleton is T-pose, and the initial pose of the target skeleton is A-pose. This step specifically adjusts all joints of the target skeleton so that the pose of the target skeleton is also T-pose, and the poses of the two skeletons are basically close.
[0093] For example, in the above steps 151 to 153, after the data is imported into the software, the effect diagrams presented by the original skeleton and the target skeleton are as follows: Figure 2 As shown, the effect diagram presented after the original skeleton and the target skeleton are matched and the posture is adjusted is as follows Figure 3 As shown, Figure 2 and Figure 3 The bone on the left is the original bone, and the bone on the right is the target bone.
[0094] Furthermore, after step 153, the above step may further include the step of correcting the data transferred to the target human body model and / or the step of baking the data transferred to the target human body model. For example, the correction may be to suppress foot sliding and hand misalignment using the IK (Inverse Kinematics) method, and the baking may be to perform interpolation smoothing on the transferred data, using linear interpolation or F-Curves interpolation to smooth the motion data.
[0095] The video-based digital human motion capture and motion generation method proposed in this application can reconstruct a 3D human body based on video images. By converting the 3D key points of the human body into bvh-formatted data that drives the 3D digital human's motion and smoothing the data, the video motion can be effectively transferred to the target 3D character, achieving the purpose of real-time control of the digital human. The method provided in this application achieves low-cost motion capture and can drive humanoid digital humans. It can also maintain real-time drive even after replacing the digital human with different skeletons.
[0096] See attached Figure 4 , Figure 4 2 is a main structural block diagram of a video-based digital human motion capture and motion generation system according to an embodiment of the present application. The system mainly includes a data acquisition module 201, a data processing module 202, a 3D reconstruction module 203 and a motion generation module 204, wherein:
[0097] The data acquisition module 201 is used to acquire a video containing character movements;
[0098] The data processing module 202 is configured to extract body region image frames and hand region image frames corresponding to the person based on the video acquired by the data acquisition module 201;
[0099] The 3D reconstruction module 203 is configured to reconstruct a human body based on the body region image frames and the hand region image frames extracted by the data processing module 202 to obtain a 3D human body model;
[0100] The motion generation module 204 is configured to obtain human skeleton information based on the 3D human body model reconstructed by the 3D reconstruction module 203 and generate a skeleton motion file based on the human skeleton information.
[0101] Furthermore, the system of the embodiment of the present application may also include a motion redirection module 205 for migrating the skeletal motion file generated by the motion generation module 204 to a target human body model, wherein the skeletal motion file is used to drive the target human body model to perform actions consistent with the character actions in the video.
[0102] For ease of explanation, the above introduction to the video-based digital human motion capture and motion generation system only shows the part related to the embodiment of the present application. For specific technical details not disclosed, please refer to the method part of the embodiment of the present application.
[0103] It should be understood that since the configuration of each module is merely for the purpose of illustrating the functional units of this application, the physical devices corresponding to these modules may be the processor itself, or a portion of the software in the processor, a portion of the hardware, or a combination of software and hardware. Therefore, the number of modules in the figure is merely illustrative.
[0104] Those skilled in the art will appreciate that the various modules in the system can be adaptively split or merged. Such splitting or merging of specific modules will not cause the technical solution to deviate from the principles of this application. Therefore, the technical solutions after splitting or merging will fall within the scope of protection of this application.
[0105] It will be understood by those skilled in the art that all or part of the processes in the method for implementing the above embodiment of the present application can also be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of the above-mentioned various method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable storage medium may include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electric carrier signal, telecommunication signal and software distribution medium, etc. that can carry the computer program code.
[0106] Furthermore, the present application also provides a computer device.
[0107] In an embodiment of a computer device according to the present application, the computer device primarily includes a processor and a storage device. The storage device can be configured to store a program for executing the video-based digital human motion capture and motion generation method of the above-described method embodiment, and the processor can be configured to execute the program in the storage device, including but not limited to a program for executing the video-based digital human motion capture and motion generation method of the above-described method embodiment. For ease of illustration, only the portions relevant to the embodiments of the present application are shown. For specific technical details not disclosed, please refer to the method section of the embodiments of the present application.
[0108] In the embodiment of the present application, the computer device may be a control device device formed by various electronic devices. In some possible implementations, the computer device may include multiple storage devices and multiple processors. The program for executing the video-based digital human motion capture and motion generation method of the above-mentioned method embodiment can be divided into multiple subroutines, and each subroutine can be loaded and run by the processor to execute different steps of the video-based digital human motion capture and motion generation method of the above-mentioned method embodiment. Specifically, each subroutine can be stored in different storage devices respectively, and each processor can be configured to execute the programs in one or more storage devices to jointly implement the video-based digital human motion capture and motion generation method of the above-mentioned method embodiment, that is, each processor executes different steps of the video-based digital human motion capture and motion generation method of the above-mentioned method embodiment respectively to jointly implement the video-based digital human motion capture and motion generation method of the above-mentioned method embodiment.
[0109] The aforementioned multiple processors may be processors deployed on the same device. For example, the aforementioned computer device may be a high-performance device composed of multiple processors, and the aforementioned multiple processors may be processors configured on the high-performance device. Furthermore, the aforementioned multiple processors may also be processors deployed on different devices. For example, the aforementioned computer device may be a server cluster, and the aforementioned multiple processors may be processors on different servers in the server cluster.
[0110] Furthermore, the present application also provides a computer-readable storage medium.
[0111] In a computer-readable storage medium embodiment according to the present application, the computer-readable storage medium can be configured to store a program for executing the video-based digital human motion capture and motion generation method of the above-mentioned method embodiment. The program can be loaded and executed by a processor to implement the above-mentioned video-based digital human motion capture and motion generation method. For ease of explanation, only the parts related to the embodiment of the present application are shown. For specific technical details not disclosed, please refer to the method section of the embodiment of the present application. The computer-readable storage medium can be a storage device formed by various electronic devices. Optionally, the computer-readable storage medium in the embodiment of the present application is a non-transitory computer-readable storage medium.
[0112] Thus far, the technical solutions of the present application have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is readily understood by those skilled in the art that the scope of protection of the present application is obviously not limited to these specific embodiments. Without departing from the principles of the present application, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present application.
Claims
1. A method for capturing and generating digital human motion based on video, characterized in that: The method comprises: Get a video containing character actions; Extracting a body region image frame and a hand region image frame corresponding to the person based on the video; Reconstructing a human body based on the body region image frame and the hand region image frame to obtain a 3D human body model; Human skeleton information is acquired based on the 3D human body model, and a skeleton motion file is generated based on the human skeleton information.
2. The method according to claim 1, characterized in that The extracting of the body region image frame and the hand region image frame corresponding to the character based on the video comprises: Based on a human body detection algorithm, detecting and determining the position of a human body from each frame of the video; Based on a hand key point detection algorithm, a hand position is selected at the position of the human body; determining a hand region frame and a body region frame based on the position of the hand and the position of the human body; A body region image frame and a hand region image frame corresponding to the person are obtained based on the body region frame and the hand region frame selected in each frame image.
3. The method according to claim 1, characterized in that The step of reconstructing a human body based on the body region image frame and the hand region image frame to obtain a 3D human body model comprises: Predicting shape parameters and posture parameters of the body based on the body region image frame and the first neural network model, and inputting the shape parameters and posture parameters of the body into the SMPL model to obtain a 3D body model; Predicting shape parameters and posture parameters of the hand based on the hand region image frame and the second neural network model, and inputting the shape parameters and posture parameters of the hand into the SMPL-X model to obtain a 3D finger model; The 3D body model and the 3D finger model are spliced together to obtain the 3D human body model.
4. The method according to claim 3, characterized in that The joints of the 3D body model include wrist joints but do not include palm joints, and the joints of the 3D finger model include wrist joints and palm joints. The step of splicing the 3D body model and the 3D finger model to obtain the 3D human body model includes: According to the correspondence between the wrist joint points of the 3D body model and the wrist joint points of the 3D finger model, the palm joint points of the 3D finger model are connected to the wrist joint points of the 3D body model.
5. The method according to claim 1, wherein The acquiring of human skeleton information based on the 3D human body model and generating a skeleton motion file based on the human skeleton information comprises: Determining a human skeleton containing semantics based on a pairwise connection relationship of 3D key points of the 3D human body model, and obtaining a topological structure of the skeleton based on the human skeleton and its semantics; Calculating the rotation matrix of each human bone according to the posture parameters of the 3D human body model to obtain the motion value of the bone; The skeleton motion file is formed according to the topological structure of the skeleton and the motion value.
6. The method according to claim 1, characterized in that The method further includes: migrating the skeleton motion file to a target human body model, wherein the skeleton motion file is used to drive the target human body model to perform actions consistent with actions of the character in the video.
7. The method according to claim 6, characterized in that Migrating the skeleton motion file to the target human body model includes: Using the topological structure and the motion value of the skeleton in the skeleton motion file as data to be migrated; Matching the skeleton in the data to be migrated with the skeleton of the target human body model based on the semantics of the skeleton to obtain a mapping relationship; Based on the mapping relationship, the posture of the corresponding skeleton of the target human body model is adjusted according to the motion value of the skeleton in the data to be migrated.
8. The method according to claim 7, characterized in that After migrating the skeleton motion file to the target human body model, the method further includes: a step of correcting the data transferred to the target human body model, and / or a step of baking the data transferred to the target human body model.
9. A computer device comprising a processor and a storage device, wherein the storage device is suitable for storing a plurality of program codes, wherein: The program code is suitable for being loaded and run by the processor to execute the video-based digital human motion capture and motion generation method according to any one of claims 1 to 8.
10. A computer-readable storage medium storing a plurality of program codes, characterized in that: The program code is suitable for being loaded and run by a processor to execute the video-based digital human motion capture and motion generation method according to any one of claims 1 to 8.
Citation Information
Cited By
Demonstration method and system of humanoid robot based on vision
CN120862643A