Method, apparatus, electronic device, and computer program for generating virtual object animations
The method improves motion capture efficiency and accuracy by using monocular video to correct posture changes in virtual object animations, addressing the limitations of existing technologies.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2024-01-09
- Publication Date
- 2026-04-27
AI Technical Summary
Existing motion capture technologies, such as optical and inertial methods, are costly and require dedicated locations and equipment, while end-to-end video methods suffer from insufficient training data leading to inaccurate motion retargeting and low generalization ability.
A method and apparatus for generating virtual object animations that utilize monocular video to determine initial three-dimensional and two-dimensional pose information, combined with foot contact information, to correct posture changes and retarget motion to virtual objects, improving efficiency and continuity.
This approach enhances motion capture accuracy and efficiency by correcting posture information using monocular video, ensuring continuity and rationality in virtual object animations without the need for additional equipment or locations.
Smart Images

Figure 2026513435000001_ABST
Abstract
Description
Technical Field
[0001] (Cross - reference to related applications) This application is filed based on a Chinese patent application with the application number 202310244258.5, which was filed with the Chinese Patent Office on March 06, 2023. This application claims the priority of the Chinese patent application, and all the contents of the Chinese patent application are incorporated herein by reference.
[0002] This application relates to image processing technology, and particularly to a method, apparatus, electronic device, computer - readable storage medium, and computer program product for generating virtual object animations.
Background Art
[0003] In related technologies, motion retargeting is widely applied in animation and film production, and motion capture is one of the key technologies in motion retargeting. Motion capture is the process of recording the movement of an object in three-dimensional space and simulating its trajectory in a digital model. For example, it involves detecting and recording the movement trajectory of an actor's limbs in three-dimensional space, capturing the actor's posture and movements, and converting the captured posture and movements into digitized abstract movements. This allows a virtual model within a software application to perform the same movements as the actor, generating an animation sequence. In recent years, motion capture technology has been widely used in many fields, including virtual reality, 3D games, and human bioengineering. Motion capture in related technologies can be implemented using optical motion capture, inertial motion capture, and end-to-end video motion capture methods. Optical and inertial motion capture methods are costly because they require the use of external equipment and locations for motion capture. On the other hand, end-to-end video motion capture methods suffer from limitations due to insufficient training data, resulting in low generalization ability and poor motion capture effectiveness. This leads to problems such as a lack of motion continuity or inaccurate motion execution when retargeting captured motion to virtual objects. [Overview of the Initiative]
[0004] Embodiments of the present invention provide a method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating virtual object animations that can reduce the cost of motion capture while ensuring the accuracy of motion obtained by motion capture.
[0005] The technical solution of the embodiment of the present application is implemented as follows.
[0006] Embodiments of the present application provide a method for generating virtual object animations, the method being: The process involves acquiring a video to be processed and determining the initial three-dimensional pose information, two-dimensional pose information, and foot contact information of the target object in the video to be processed. The steps include determining target pose change information between two adjacent video frames in the video to be processed, based on the initial three-dimensional pose information of the target object, The steps include determining the target object's awaiting correction posture information based on the target object's initial three-dimensional posture information and posture change information, The steps include: performing a correction process on the target object's awaiting correction posture information based on the target object's two-dimensional posture information and foot contact information to obtain the corrected posture information of the target object; The process includes the steps of motion retargeting the modified pose information to a virtual object and generating a virtual object animation video corresponding to the video to be processed.
[0007] Embodiments of the present application provide a virtual object animation generation device, the device is A first acquisition module is configured to acquire a video to be processed and to determine the initial three-dimensional pose information, two-dimensional pose information, and foot contact information of the target object in the video to be processed. A first decision module is configured to determine the pose change information between adjacent video frames based on the initial three-dimensional pose information of the target object, A second determination module is configured to determine the target object's awaiting correction orientation information based on the target object's initial three-dimensional orientation information and orientation change information between adjacent video frames. A first processing module is configured to perform a correction process on the target object's awaiting correction posture information based on the target object's two-dimensional posture information and foot contact information, and to obtain the corrected posture information of the target object. The system includes a motion retargeting module configured to motion retarget the modified posture information to a virtual object and generate a virtual object animation video corresponding to the video to be processed.
[0008] Embodiments of the present application provide an electronic device, said electronic device, Memory configured to store computer executable instructions, The present invention comprises a processor configured to realize the virtual object animation generation method provided in the embodiment of the present invention by executing computer executable instructions stored in the memory.
[0009] Embodiments of the present invention provide a computer-readable storage medium in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, a virtual object animation generation method provided in embodiments of the present invention is realized.
[0010] Embodiments of the present application provide a computer program product including a computer program or computer executable instructions, which, when executed by a processor, implements a virtual object animation generation method provided in embodiments of the present application. [Effects of the Invention]
[0011] The embodiments of this application have the following beneficial effects.
[0012] After acquiring the video to be processed, first, the initial three-dimensional posture information, two-dimensional posture information, and foot contact information of the target object in the video to be processed are determined. Next, based on the initial three-dimensional posture information of the target object, posture change information between adjacent video frames is determined. This posture change information is used as a motion precedent and combined with the initial three-dimensional posture information of each video frame to determine the posture information awaiting correction. Subsequently, based on the two-dimensional posture information and foot contact information of the target object, a correction process is performed on the posture information awaiting correction of the target object to obtain the corrected posture information of the target object. Since the posture information awaiting correction of the target object is determined based on the posture change information as a motion precedent and the initial three-dimensional posture information of each video frame, correcting the posture information awaiting correction is essentially equivalent to correcting the posture change information. This not only improves the efficiency of motion capture by improving correction efficiency, but also ensures the continuity and rationality of the motion of the virtual object animation video obtained by retargeting the corrected posture information to the virtual object through motion retargeting. [Brief explanation of the drawing]
[0013] [Figure 1] This is a schematic diagram of the architectural structure of the motion capture system 100 according to an embodiment of the present invention. [Figure 2] This is a schematic diagram of the structure of server 400 according to an embodiment of the present invention. [Figure 3A] This is a flowchart of a virtual object animation generation method according to an embodiment of the present invention. [Figure 3B] This is an implementation flowchart for determining the initial three-dimensional orientation information, two-dimensional orientation information, and foot contact information of a target object according to an embodiment of the present invention. [Figure 3C] This is a flowchart for determining posture change information according to an embodiment of the present invention. [Figure 3D] This is an implementation flowchart for training a motion priori model according to an embodiment of the present invention. [Figure 4A] A realization flowchart for determining posture information awaiting correction according to an embodiment of the present application. [Figure 4B] A realization flowchart for performing correction processing on posture information awaiting correction according to an embodiment of the present application. [Figure 4C] A realization flowchart for determining reprojection error according to an embodiment of the present application. [Figure 5A] A realization flowchart for determining prior information according to an embodiment of the present application. [Figure 5B] A realization flowchart for determining regularization term error according to an embodiment of the present application. [Figure 5C] Another flowchart of a method for generating virtual object animation according to an embodiment of the present application. [Figure 6] A schematic diagram of a joint tree according to an embodiment of the present application. [Figure 7A] A human body mesh obtained using a base template constructed with an SMPL model according to an embodiment of the present application. [Figure 7B] A human body mesh obtained using a base template constructed with an SMPL model according to an embodiment of the present application. [Figure 7C] A schematic diagram of a human body model obtained after adjusting shape parameters according to an embodiment of the present application. [Figure 8A] Yet another flowchart of a method for generating virtual object animation according to an embodiment of the present application. [Figure 8B] An overall processing flowchart of a method for generating virtual object animation according to an embodiment of the present application. [Figure 9] A schematic diagram of the network structure of a 3D pose estimation module according to an embodiment of the present application. [Figure 10A] A schematic diagram of the network structure of a 2D pose estimation module according to an embodiment of the present application. [Figure 10B] A schematic diagram of the network structure of a foot contact prediction module according to an embodiment of the present application. [Figure 10C]This is a schematic diagram of a variational autoencoder according to an embodiment of the present invention. [Figure 10D] This is a schematic diagram of the network structure of the motion prior model according to the embodiment of the present invention. [Figure 11] This is a schematic diagram of the application interface in the virtual engine of the virtual object animation generation method according to an embodiment of the present invention. [Modes for carrying out the invention]
[0014] To further clarify the purpose, technical solutions, and advantages of this application, the application is described in more detail below with reference to the drawings. The embodiments described are not limiting to this application, and all other embodiments that can be obtained without creative effort by those skilled in the art are included within the scope of this application.
[0015] In the following description, the term “several embodiments” refers to a subset of all possible embodiments, but understandably, “several embodiments” can be the same subset or different subsets of all possible embodiments, and can be combined with each other as long as they do not contradict each other.
[0016] In the following description, terms such as “first / second / third” are used merely to distinguish similar objects and do not represent a specific order of objects. Since “first / second / third” can, where possible, change a specific order or sequence, it should be understood that the embodiments of the present application described herein may be performed in an order other than those illustrated or described herein.
[0017] In embodiments of the present application, the terms “module” or “unit” refer to a computer program or part of a computer program having a predetermined function, which works in conjunction with other relevant parts to achieve a predetermined purpose, and which can be fully or partially implemented by software, hardware (e.g., processing circuits and memory), or a combination thereof. Similarly, one or more modules or units can be implemented using one processor (or more processors or memory). Furthermore, each module or unit may be part of an overall module or unit that includes the function of that module or unit.
[0018] Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as those commonly understood by those skilled in the art. The terms used in this application are solely for illustrative purposes of the embodiments of this application and are not intended to limit it.
[0019] Before describing the embodiments of this application in detail, the nouns and terms used in the embodiments of this application will be explained below.
[0020] 1) Motion capture, also known as motion capture, is the process of recording how an object moves in three-dimensional space and simulating its motion trajectory in a digital model.
[0021] 2) SMPL (Skinned Multi-Person Linear Model) is a nude, vertex-based three-dimensional model of the human body that can accurately represent different shapes and poses of the human body.
[0022] Motion capture is one of the key techniques in virtual object animation generation methods. To better understand the virtual object animation generation method provided in the embodiment of this application, we will first describe the implementation processes and drawbacks of optical motion capture, inertial motion capture, and end-to-end video motion capture methods provided in related technologies.
[0023] Optical motion capture systems are based on the principles of computer vision, using multiple high-speed cameras to continuously identify target feature points from different angles and simultaneously performing motion capture in combination with skeletal calculation algorithms. Theoretically, for any point in space, if that point can be observed simultaneously by two or more cameras, its 3D position in space at that moment can be determined. When cameras continuously capture images at a high frame rate, the motion trajectory of the point can be obtained from the image sequence. This technology requires a dedicated location, demands no obvious interference from the surroundings, and necessitates a professional performer wearing the optical motion capture device to capture movement. While the capture results of optical motion capture technology are highly accurate, it has requirements for location, equipment, and personnel, resulting in high usage costs.
[0024] Inertial motion capture technology requires the attachment of integrated inertial sensor devices, such as accelerometers, gyroscopes, and magnetometers, to key joint points on the human body. These sensor devices capture motion data of a target object, including information such as the posture and orientation of body parts, and transmit this data to a data processing device via a data transmission device. After data correction and processing, a three-dimensional model is finally constructed, and this model is made to move in accordance with the actual and natural movements of the human body. However, this technology requires a professional performer to wear specific equipment, and due to accuracy issues with the inertial devices, the capture effect is inferior to optical motion capture technology, and it cannot solve the problem of keeping the soles of the virtual character's feet firmly on the ground.
[0025] End-to-end video motion capture methods require acquiring 3D human body data in real-world environments, which is highly challenging. Implementing this technology requires not only cameras but also additional equipment, making it extremely difficult to directly output 3D position data as network input. To address the data shortage problem, this approach involves training by transforming the background or changing the clothing of the human body, but the capture effect is not ideal.
[0026] Based on this, embodiments of the present application provide a method, apparatus, device, and computer-readable storage medium for generating virtual object animations, which are based on end-to-end prediction and incorporate prior information on human body movements, allowing for further iterative optimization to adapt accuracy to practical application needs. At the same time, compared to optical motion capture and inertial motion capture, this solution does not require additional locations or equipment and only requires monocular video, thus expanding the application scenarios of motion capture. The following describes exemplary applications of the computer device provided in embodiments of the present application. The device provided in embodiments of the present application may be implemented as various user terminals such as laptop computers, tablet computers, desktop computers, set-top boxes, and mobile devices (e.g., mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable game consoles, etc.), or as a server. The following describes exemplary applications when the device is implemented as a server.
[0027] Referring to Figure 1, which is a schematic diagram of the architecture of a motion capture system 100 provided in an embodiment of the present invention, in which terminals (exemplarily shown as terminals 200-1 and terminals 200-2) are connected to a server 400 via a network 300, the network 300 may be a wide area network, a local network, or a combination of both.
[0028] In the motion capture system, terminal 200 is used to collect at least video and transmit the collected video to server 400 via network 300, where server 400 either performs motion capture processing or stores it in database 500. Terminal 200 can collect video using an image acquisition device within the terminal.
[0029] Server 400 acquires the video to be processed, which may be a video acquired from terminal 200 or a pre-stored video acquired from database 500. Subsequently, Server 400 determines the initial three-dimensional pose information, two-dimensional pose information, and foot contact information of the target object in the video to be processed. Next, based on the initial three-dimensional pose information of the target object, it determines the pose information of the target object awaiting correction based on the initial three-dimensional pose information and pose change information of the target object. Based on the two-dimensional pose information and foot contact information of the target object, it performs a correction process on the pose information awaiting correction of the target object to obtain the corrected pose information of the target object. Finally, by retargeting the corrected pose information to a virtual object using motion retargeting, it generates a virtual object animation video corresponding to the video to be processed. The target object's awaiting correction posture information is determined based on posture change information as prior to motion. Therefore, correcting the awaiting correction posture information is essentially equivalent to correcting the posture change information. This not only improves the efficiency of motion capture by improving correction efficiency, but also ensures the continuity and coherence of the virtual object animation video obtained by retargeting the corrected posture information to the virtual object through motion retargeting. Subsequently, the server 400 transmits the virtual object animation video to the terminal 200, enabling the terminal 200 to play the virtual object animation video. In the virtual object animation video, the virtual object performs actions corresponding to the corrected posture information.
[0030] In some embodiments, the server 400 may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDNs, and big data and artificial intelligence platforms. The terminal 200 may be, but is not limited to, a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, or in-vehicle terminal. The terminal and server may be connected directly or indirectly by wired or wireless communication, and the embodiments of this application are not limited to these.
[0031] Referring to Figure 2, which is a schematic diagram showing the configuration of a server 400 provided in an embodiment of the present application, the server 400 shown in Figure 2 includes at least one processor 410, memory 450, at least one network interface 420, and a user interface 430. Each component in the server 400 is coupled via a bus system 440. Understandably, the bus system 440 is used to enable connection communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, various buses are represented as the bus system 440 in Figure 2.
[0032] The processor 410 may be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), a programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, where the general-purpose processor may be a microprocessor or any conventional processor.
[0033] The user interface 430 includes one or more output devices 431 for displaying media content, each output device 431 including one or more speakers and / or one or more visual displays. The user interface 430 further includes one or more input devices 432, each input device 432 including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touchscreen display, camera, and other input buttons and controls.
[0034] The memory 450 may be removable, non-removable, or a combination of both. Exemplary hardware devices include solid-state memory, hard drives, optical drives, and the like. The memory 450 includes one or more storage devices located physically separate from the processor 410.
[0035] The memory 450 may include volatile memory or non-volatile memory, or both volatile and non-volatile memory. Non-volatile memory may be read-only memory (ROM), and volatile memory may be random-access memory (RAM). The memory 450 described in the embodiments of this application is intended to include any suitable type of memory.
[0036] In some embodiments, memory 450 can store data to support various operations, as illustrated below, and examples of this data include programs, modules, and data structures, or subsets or supersets thereof.
[0037] Operating System 451 is configured to perform various basic tasks and handle hardware-based tasks, including system programs for handling various basic system services such as the frame layer, core library layer, and drive layer, and for performing hardware-related tasks.
[0038] The network communication module 452 is for accessing other computing devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including Bluetooth, Wireless Compatibility Certification (WiFi), and Universal Serial Bus (USB).
[0039] The display module 453 is used to enable the display of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with the user interface 430.
[0040] The input processing module 454 is for detecting one or more user inputs or interactions from one or more input devices 432 and for translating the detected inputs or interactions.
[0041] In some embodiments, the apparatus provided in the embodiments of the present invention may be implemented in software form. Figure 2 shows a virtual object animation generation device 455 stored in memory 450. The virtual object animation generation device 455 may be software in the form of a program or plug-in, and comprises a first acquisition module 4551, a first decision module 4552, a second decision module 4553, a first processing module 4554, and a motion retargeting module 4555. Since these modules are logical, they can be arbitrarily combined or further divided depending on the function to be implemented. The function of each module will be described below.
[0042] In some other embodiments, the apparatus provided in the embodiments of the present application can be implemented in hardware form, for example, the apparatus provided in the embodiments of the present application may be a hardware decoding processor, which is programmed to perform the virtual object animation generation method provided in the embodiments of the present application, and for example, the hardware decoding processor may employ one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic elements.
[0043] The virtual object animation generation method provided in the embodiment of this application will be described with reference to the exemplary application and implementation of the server provided in the embodiment of this application.
[0044] To better understand the virtual object animation generation method provided in the embodiments of this application, we will first explain artificial intelligence, the various fields of artificial intelligence, and the application fields related to the virtual object animation generation method provided in the embodiments of this application.
[0045] Artificial intelligence (AI) is the theory, methods, techniques, and applied systems that use digital computers or machines controlled by digital computers to simulate, extend, and augment human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve the best results. In other words, AI is a comprehensive field of computer science that attempts to understand the nature of intelligence and create new intelligent machines that can react in a similar way to human intelligence. AI studies the design principles and implementation methods of various intelligent machines to enable them to perceive, reason, and make decisions.
[0046] Artificial intelligence technology is a comprehensive field encompassing a wide range of disciplines, including both hardware and software technologies. Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. Artificial intelligence software technologies primarily include key areas such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. The virtual object animation generation method provided in the embodiments of this application is primarily related to the field of machine learning and is described below.
[0047] Machine learning (ML) is a multidisciplinary field encompassing various disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. Machine learning specializes in the study of how computers can simulate or replicate human learning behavior to acquire new knowledge and skills, reorganize existing knowledge structures, and continuously improve performance. Machine learning is the core of artificial intelligence, a fundamental means of giving computers intelligence, and is applied to various fields of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, trust networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.
[0048] With the ongoing research and advancement of artificial intelligence technology, it is being studied and applied in various fields such as smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, drones, robotics, smart healthcare, and smart customer service. As the technology develops, it is expected that artificial intelligence technology will be applied in even more fields and will become increasingly important.
[0049] The following describes a method for generating virtual object animations. As mentioned above, the electronic device that implements the virtual object animation generation method according to the embodiment of this application may be a server. Therefore, the entity that executes each step will not be described repeatedly below.
[0050] Referring to Figure 3A, which is a flowchart of the virtual object animation generation method provided in an embodiment of the present application, the steps shown in Figure 3A will be described with reference to them.
[0051] In step 101, the video to be processed is acquired, and the initial three-dimensional pose information, two-dimensional pose information, and foot contact information of the target object in the video to be processed are determined.
[0052] Here, the video to be processed is a monocular video, i.e., a video collected by a single image acquisition device. The target object in the video to be processed is not stationary but performs a series of actions. The target object may be a natural person, a digital human, or a virtual character. The video to be processed may be collected using the image acquisition device of a terminal, downloaded from a network, or transmitted to a server from another device.
[0053] In some embodiments, referring to Figure 3B, the "determining the initial three-dimensional pose information, two-dimensional pose information, and foot contact information of the target object in the video to be processed" in step 101 of Figure 3A can be achieved by the following steps 1011 to 1014, which will be explained below with reference to Figure 3B.
[0054] In step 1011, a trained three-dimensional pose estimation model, a trained two-dimensional pose estimation model, and a trained foot contact detection model are obtained.
[0055] In some embodiments, a predetermined initial three-dimensional pose estimation model, an initial two-dimensional pose estimation model, and an initial foot contact detection model are trained using a first training video to obtain a trained three-dimensional pose estimation model, a trained two-dimensional pose estimation model, and a trained foot contact detection model. During the training process, first, a predetermined initial three-dimensional pose estimation model, an initial two-dimensional pose estimation model, and an initial foot contact detection model are obtained, and then a first training video containing multiple first training video frames is obtained, and the three-dimensional pose information, two-dimensional pose information, and foot contact information of the target object in each first training video frame are obtained, respectively, the three-dimensional pose information of the target object in each first training video frame is determined as the three-dimensional annotation information of each first training video frame, and the two-dimensional pose information of the target object in each first training video frame is determined as the two-dimensional annotation information of each first training video frame, and the target object in each first training video frame The foot contact information of the object is determined as the ground contact annotation information for each first training video frame. Then, using each first training video frame and the corresponding three-dimensional annotation information, an initial three-dimensional posture estimation model is trained to obtain a trained three-dimensional posture estimation model. Using each first training video frame and the corresponding two-dimensional annotation information, an initial two-dimensional posture estimation model is trained to obtain a trained two-dimensional posture estimation model. Using each first training video frame and the corresponding ground contact annotation information, an initial foot contact detection model is trained to obtain a trained foot contact detection model.
[0056] In step 1012, the trained three-dimensional pose estimation model is used to perform pose estimation on the video to be processed, and initial three-dimensional pose information of the target object in the video to be processed is obtained.
[0057] In some embodiments, first, a constructed joint point tree and the total number of video frames N in the video to be processed are obtained. The joint point tree is a tree-like structure constructed based on a predefined number of joint points and their mutual influences. Then, using a trained three-dimensional pose estimation model, pose estimation is performed on the target object in the i-th video frame out of the N video frames, and initial three-dimensional pose information of the target object in that i-th video frame is obtained. This initial three-dimensional pose information includes rotation information for each joint point of the target object in the i-th video frame relative to its parent node, as well as orientation information, rotation information, and offset information for the entire skeleton, where i = 1, 2, ..., N.
[0058] In step 1013, the trained two-dimensional pose estimation model is used to perform pose estimation on the video to be processed, and two-dimensional pose information of the target object is obtained.
[0059] In some embodiments, the i-th video frame out of N video frames is input to a trained two-dimensional pose estimation model, and pose estimation is performed on the target object in the i-th video frame to obtain initial two-dimensional pose information for the target object in the i-th video frame. This initial two-dimensional pose information includes the position coordinates of each joint point of the target object in at least the i-th video frame, where i = 1, 2, ..., N. Assuming that the target object is defined to have 24 joint points, the position coordinates of the j-th joint point are (m j ,n j ) can be expressed as follows, and the initial two-dimensional pose information of the target object in the i-th video frame can be represented as a single 24*2-dimensional feature matrix.
[0060] In some embodiments, the two-dimensional orientation information of the target object further includes confidence information for a given joint point, which is used to represent the reliability of the position coordinates of each joint point, with a higher confidence information for a given joint point indicating greater reliability of its position coordinates.
[0061] In step 1014, the trained foot contact detection model is used to perform prediction processing on the two-dimensional posture information of the target object to obtain the foot contact information of the target object.
[0062] The two-dimensional pose information of the target object in the i-th video frame obtained in step 1013 is input into the trained foot contact detection model to determine whether the target object's feet are in contact with the ground in the i-th video frame, thereby obtaining the target object's foot contact information for the i-th video frame. This foot contact information can be represented as a four-dimensional vector (rf,rb,lf,lb), where rf indicates whether the right toe is in contact with the ground, rb indicates whether the right heel is in contact with the ground, lf indicates whether the left toe is in contact with the ground, and lb indicates whether the left heel is in contact with the ground. For example, if the right toe is in contact with the ground, the value of rf is 1, and if the right toe is not in contact with the ground, the value of rf is 0. Assuming that the target object's right toe is in contact with the ground, the right heel is not, the left toe is in contact with the ground, and the left heel is in contact with the ground, the target object's foot contact information is (1,0,1,1).
[0063] In step 102, the orientation change information between two adjacent video frames in the video to be processed is determined based on the initial three-dimensional orientation information of the target object.
[0064] In some embodiments, the initial three-dimensional pose information may include initial pose parameters and initial shape parameters. Assuming the current video frame is the i-th video frame, in the implementation of step 102, prediction processing is performed on the initial pose parameters of the i-th video frame based on the motion priori model to obtain pose change information of the target object from the (i-1)th video frame to the i-th video frame. Note that the pose change information of the target object from the (i-1)th video frame to the i-th video frame, determined based on the motion priori model, is an estimate and differs from the pose change information of the target object from the (i-1)th video frame to the i-th video frame, which is determined based on the initial pose parameters of the target object in the (i-1)th video frame and the initial pose parameters of the target object in the i-th video frame. Here, i = 1, 2, ..., N-1, where N is the total number of video frames in the video to be processed.
[0065] In some embodiments, referring to Figure 3C, step 102 in Figure 3A can be achieved by the following steps 1021 to 1022, which will be explained below with reference to Figure 3C.
[0066] In step 1021, a trained motor a priori model is obtained.
[0067] In some embodiments, as shown in Figure 3D, a trained motor prior model can be obtained by steps 211 to 215 below, prior to step 1021, which are described below with reference to Figure 3D.
[0068] In step 211, obtain the training video and the designated exercise trial model.
[0069] Here, the training video includes multiple training video frames, in which the target object performs a series of actions. The motion priori model may be a neural network model or a deep learning model. The motion priori model is used to predict information about the pose change of the target object in two adjacent video frames.
[0070] In step 212, the k-th training video frame and the (k+1)-th training video frame are determined to be the k-th training data pair.
[0071] In some embodiments, first, M training video frames are obtained from the training video. Then, two adjacent training video frames are used as one training data pair, meaning that M-1 training data pairs can be constructed from the M training video frames, and the k-th training data pair consists of the k-th training video frame and the (k+1)-th training video frame. Here, k = 1, 2, ..., M-1, and M is the total number of video frames in the training video.
[0072] In step 213, encoding is performed on the two training video frames in the k-th training data pair to obtain the encoded results of the two training video frames in the k-th training data pair in the latent space.
[0073] Here, first, the training pose parameters of two training video frames in the k-th training data pair are obtained, and then the training pose parameters of the two training video frames are encoded to obtain the encoded result of the two training video frames in the k-th training data pair in the latent space. The encoder used to encode the training pose parameters may be a variational autoencoder, which is an encoder obtained by pre-training using a large dataset of 3D human poses. The 3D human poses included in the dataset used for training are poses that the human body can perform, and by encoding the training pose parameters of the k-th and (k+1)-th training video frames in the k-th training data pair in the latent space, the appearance of poses that the human body cannot perform can be prevented.
[0074] In step 214, predictive change information for the two training video frames in the k-th training data pair is determined based on the encoding results of the two training video frames in the k-th training data pair in the latent space.
[0075] In some embodiments, coding difference information is determined between the coding result of the k-th training video frame in the latent space and the coding result of the (k+1)-th training video frame in the latent space. Then, a decoder is used to decode the coding difference information to obtain predicted change information for the two training video frames in the k-th training data pair.
[0076] In step 215, the parameters of the predetermined motor a priori model are adjusted based on the predicted change information of two training video frames within each training data pair to obtain a trained motor a priori model.
[0077] In some embodiments, reference change information for the k-th and k+1-th training video frames is determined based on the training pose parameters of the k-th and k+1-th training video frames. Then, the difference in change between the reference change information and the predicted change information for each training data pair is determined. This difference in change is backpropagated to the motion priori model, and the parameters of the motion priori model are adjusted using a gradient descent algorithm to obtain a trained motion priori model.
[0078] In some embodiments, the motion priori model may include an encoder, a decoder, and a posture change determination module, where the parameters of the posture change determination module can be determined by the mean change and standard deviation of change determined from multiple coded difference information. The posture change determination module of the trained motion priori model may be a module that follows a normal distribution.
[0079] In steps 211 to 215 above, first, multiple training data pairs are constructed from two adjacent training video frames within a training video. Then, an encoding process is performed on the two training video frames contained in each training data pair using a variational autoencoder. This ensures that the poses corresponding to the obtained encoding results are poses that can be performed by the human body. Furthermore, based on the encoding results, predictive change information for the two training video frames within each training data pair is determined, and based on the multiple pose change information, the motor a priori model parameters are adjusted to obtain a trained motor a priori model.
[0080] In step 1022, the trained motion priori model is used to perform a prediction process on the initial posture parameters of the target object and obtain posture change information between adjacent video frames.
[0081] In some embodiments, first, an encoder within a trained motion priori model is used to encode the initial pose parameters in each video frame of the target object, obtaining the encoded result corresponding to each video frame of the target object. Next, a pose change determination module within the trained motion priori model is used to perform a prediction process on the encoded result, obtaining each prediction result. Finally, a decoder within the trained motion priori model is used to decode the prediction result, obtaining pose change information between adjacent video frames.
[0082] The dimensions of the posture change information and the posture parameters are the same; for example, they may be 24*3.
[0083] Steps 1021 to 1022 above allow us to ensure the continuity of motion by using a trained motion priori model to determine information about changes in posture between adjacent video frames of the target object.
[0084] In step 103, the target object's awaiting correction orientation information is determined based on the target object's initial three-dimensional orientation information and orientation change information.
[0085] In some embodiments, referring to Figure 4A, step 103 in Figure 3A can be achieved by the following steps 1031 to 1035, which will be explained below with reference to Figure 4A.
[0086] In step 1031, the video to be processed is divided into multiple video segments.
[0087] In some embodiments, the total number of video frames of the video to be processed is obtained first, then the number of video frames to be included in each video segment is determined based on the total number of video frames, and the video to be processed can be divided into multiple video segments using the number of video frames. The total number of video frames can be an integer multiple of the number of video frames, thereby ensuring that the number of video frames included in each video segment is the same. For example, if the total number of video frames is 500 and the number of video frames is 50, the video to be processed can be divided into 10 video segments, each video segment containing 50 video frames.
[0088] In some embodiments, the video to be processed can be directly divided into multiple video segments according to a predetermined number of video frames. For example, the video to be processed can be divided with 45 video frames, in which case only the last video segment may contain fewer than 45 video frames, while all other video segments contain 45 video frames.
[0089] In step 1032, if it is determined that the i-th video frame is the first frame in the video segment, the initial three-dimensional pose information of the target object is determined as the pose information of the target object awaiting correction.
[0090] Here, i = 1, 2, ..., N, where N is the total number of video frames in the video to be processed, and the i-th video frame is the current video frame in which the target object is located.
[0091] In the embodiments of this invention, the initial three-dimensional pose information of the target object in the first video frame is used as a reference, and therefore the initial three-dimensional pose information of the target object in the first video frame is determined as the target object's pose information awaiting correction.
[0092] In step 1033, if it is determined that the i-th video frame is not the first frame in the video segment, the modified pose information of the target object in the (i-1)-th video frame is obtained.
[0093] Here, the corrected posture information includes the corrected posture parameters and the corrected shape parameters.
[0094] In step 1034, the target object's awaited pose parameter is determined based on the pose change information of the target object from the (i-1)th video frame to the i-th video frame and the modified pose parameter of the target object corresponding to the (i-1)th video frame.
[0095] In some embodiments, the pose change information of the target object from the (i-1)th video frame to the i-th video frame, and the cumulative sum of the modified pose parameters of the target object corresponding to the (i-1)th video frame, are determined as the target object's pose parameters awaiting modification.
[0096] In step 1035, the initial shape parameters corresponding to the target object and the target object's awaiting modification orientation parameters are determined as the target object's awaiting modification orientation information.
[0097] Since a person's height and body shape do not usually change significantly by performing different actions, in the embodiment of the present invention, the initial shape parameter corresponding to the i-th video frame of the target object and the awaiting posture parameter in the i-th video frame of the target object are determined as the awaiting posture information in the i-th frame of the target object.
[0098] In steps 1031 to 1035 described above, first, the video to be processed is divided into multiple video segments, and the first video frame of each video segment is used as reference information, thereby reducing the cumulative amount of error. Furthermore, when determining the pose information awaiting correction, the initial shape parameters of the target object are determined as the shape parameters awaiting correction, thereby reducing the amount of data processing. Moreover, since the pose parameters awaiting correction of the target object in video frames other than the first frame in a video segment are obtained by accumulating the corrected pose parameters and pose change information of the previous frame, it can be understood that only the pose change information is corrected during correction, which not only improves correction efficiency but also guarantees the continuity and naturalness of the motion.
[0099] In step 104, based on the two-dimensional posture information and foot contact information of the target object, a correction process is performed on the posture information of the target object awaiting correction, and the corrected posture information of the target object is obtained.
[0100] In some embodiments, referring to Figure 4B, step 104 in Figure 3A can be achieved by the following steps 1041 to 1045, which will be explained below with reference to Figure 4B.
[0101] In step 1041, the reprojection error of the i-th video frame is determined based on the target object's awaited correction pose information, two-dimensional pose information, and camera parameters.
[0102] Here, the awaiting correction pose information includes awaiting correction pose parameters and awaiting correction shape parameters, and the two-dimensional pose information includes reference coordinate information for each joint point of the target object and the confidence level of each joint point. Camera parameters include camera external parameters and camera internal parameters.
[0103] In some embodiments, referring to Figure 4C, step 1041 in Figure 4B can be achieved by the following steps 411 to 414, which will be explained below with reference to Figure 4C.
[0104] In step 411, the camera coordinate information of each joint point of the target object is determined based on the target object's awaiting correction pose information and two-dimensional pose information.
[0105] In some embodiments, the camera coordinate information of each joint point of the target object can be determined using equation (1-1).
[0106] TIFF2026513435000002.tif14170 Here, TIFF2026513435000003.tif5170 contains basic human body posture mesh information. TIFF2026513435000004.tif5170 is the target object's pending pose parameter in the i-th video frame, TIFF2026513435000005.tif5170 is the shape parameter of the target object awaiting modification in the i-th video frame. TIFF2026513435000006.tif5170 is a linear transformation matrix corresponding to the shape parameter, TIFF2026513435000007.tif6170 is a linear transformation matrix corresponding to the attitude parameters, TIFF2026513435000008.tif5170 is two-dimensional pose information.
[0107] The camera coordinate information for each joint point obtained by equation (1-1) is the three-dimensional coordinate of each joint point in the camera coordinate system.
[0108] In step 412, the image coordinate information of each joint point is determined based on the camera parameters and the camera coordinate information of each joint point of the target object.
[0109] In some embodiments, the image coordinate information of each joint point can be determined by equation (1-2).
[0110] TIFF2026513435000009.tif16170 Here, [x j ,y j ] is the image coordinate information of the j-th joint point, and [X j ,Y j ,Z j ] is the camera coordinate information of the j-th joint point, and f is the focal length of the camera.
[0111] In step 413, the coordinate error of each joint point is determined based on the image coordinate information of each joint point and the reference coordinate information of each joint point.
[0112] In some embodiments, the reference coordinate information of each joint point can be obtained from two-dimensional pose information, the distance between the image coordinate information and the reference coordinate information of each joint point is determined, and this distance is determined as the coordinate error of the key point.
[0113] In step 414, the reprojection error of the target object is determined based on the coordinate error of each joint point and the reliability of each joint point.
[0114] In some embodiments, the reprojection error of the i-th video frame can be determined by equation (1-3).
[0115] TIFF2026513435000010.tif15170 Here, TIFF2026513435000011.tif5170 is the confidence level corresponding to the j-th articular point. TIFF2026513435000012.tif6170 is the coordinate error of the j-th joint point. TIFF2026513435000013.tif8170 represents the Geman-McClure penalty function, where σ is set to 100.
[0116] In steps 411 to 414 above, a penalty function is introduced when determining the reprojection error in the i-th video frame of the target object. This penalty function is primarily used to punish outliers, thereby making the iterative optimization more stable.
[0117] In step 1042, prior information of the target object is determined based on the target object's awaiting correction posture information.
[0118] Here, prior information includes posture prior information and shape prior information, and posture information awaiting modification includes posture parameters awaiting modification and shape parameters awaiting modification.
[0119] In some embodiments, referring to Figure 5A, step 1042 in Figure 4B can be achieved by the following steps 421 to 424, which will be explained below with reference to Figure 5A.
[0120] In step 421, the shape prior information of the target object is determined based on the shape parameters of the target object awaiting modification.
[0121] In some embodiments, the target object's awaited shape parameter refers to the target object's awaited shape parameter in the i-th video frame, and the target object's prior shape information refers to the target object's prior shape information in the i-th video frame. The awaited shape parameter is represented as a vector, which may be, for example, a 1*10 vector, and in this step, the norm of the vector representing the awaited shape parameter can be determined as the prior shape information.
[0122] In step 422, encoding is performed on the target object's awaiting pose parameters to obtain the encoded result of the awaiting pose parameters in latent space.
[0123] In some embodiments, a variational autoencoder can be used to encode the awaiting pose parameters in the i-th video frame of the target object, thereby obtaining the encoded result of the awaiting pose parameters in latent space. By encoding the awaiting pose parameters using a variational autoencoder, it is possible to ensure that the pose corresponding to the encoded result is a feasible movement of the human body.
[0124] In step 423, the attitude prior information of the target object is determined based on the encoding result of the awaited attitude parameters in latent space.
[0125] The encoding result can be represented as a vector, which may be, for example, a 32-dimensional vector. In some embodiments, the target object's pose prior information refers to the pose prior information of the target object in the i-th video frame, and the norm of the vector representing the encoding result is determined as the target object's pose prior information.
[0126] In step 424, the prior information of the target object is determined based on the prior information of the shape of the target object and the prior information of the posture of the target object.
[0127] In some embodiments, a first weight corresponding to shape priori information and a second weight corresponding to posture priori information are obtained. Then, the first and second weights are used to perform weighted addition on the shape priori information and posture priori information to obtain priori information for the target object.
[0128] In steps 421 to 424 described above, shape prior information is determined based on the shape parameters of the target object awaiting modification, encoding is performed on the posture parameters awaiting modification, and posture prior information is determined based on the encoding results. This ensures that the posture prior information can represent the feasible movements of the human body, and the rationality of the inter-frame changes in human posture can be ensured by shape prior and posture prior.
[0129] In step 1043, the normalization term error of the target object is determined based on the foot contact information and corrected posture information of the target object in the (i-1)th video frame and the foot contact information and awaiting correction posture information of the target object in the ith video frame.
[0130] In some embodiments, referring to Figure 5B, step 1043 in Figure 4B can be achieved by the following steps 431 to 434, which will be explained below with reference to Figure 5B.
[0131] In step 431, the positional difference values of each joint point of the target object and the length difference values of each skeleton of the target object are determined based on the modified pose information of the target object in the (i-1)th video frame and the pending pose information of the target object in the i-th video frame.
[0132] In some embodiments, the (i-1)th position coordinate of each joint point is obtained from the modified pose information of the target object in the (i-1)th video frame, and the i-th position coordinate of each joint point is obtained from the awaited pose information of the target object in the i-th video frame. Then, the position difference value of each joint point is determined based on the (i-1)th position coordinate and the i-th position coordinate of each joint point. Then, the (i-1)th length value of each skeleton is determined based on the (i-1)th position coordinates of the two joint points on each skeleton, and the i-th length value of each skeleton is determined based on the i-th position coordinates of the two joint points on each skeleton, and the length difference value of each skeleton is determined.
[0133] In step 432, the foot velocity error of the target object is determined based on the foot contact information in the (i-1)th video frame of the target object and the foot contact information in the ith video frame of the target object.
[0134] In some embodiments, first, the first velocity difference between the left toe of the target object in the (i-1)th video frame and the left toe of the target object in the i-th video frame is determined, then the second velocity difference between the left heel of the target object in the (i-1)th video frame and the left heel of the target object in the i-th video frame is determined, and then the third velocity difference between the right toe of the target object in the (i-1)th video frame and the right toe of the target object in the i-th video frame is determined, The fourth velocity difference value is determined between the right heel of the target object in the (i-1)th video frame and the right heel of the target object in the ith video frame. Next, the first velocity difference value is multiplied by the foot contact information of the left toe to obtain the first product. The second velocity difference value is multiplied by the foot contact information of the left heel to obtain the second product. The third velocity difference value is multiplied by the foot contact information of the right toe to obtain the third product. The fourth velocity difference value is multiplied by the foot contact information of the right heel to obtain the fourth product. Finally, the sum of the first, second, third, and fourth products is determined as the foot velocity error.
[0135] In step 433, the difference in height between the foot joint of the target object and the ground is determined.
[0136] In some embodiments, the i-th position coordinate of the foot joint of the target object is determined, and then the elevation difference between the foot joint and the ground is determined based on the i-th position coordinate.
[0137] In step 434, the normalized term error of the target object is determined based on the positional difference values of each joint point of the target object, the length difference values of each skeleton of the target object, the foot velocity error of the target object, and the height difference between the leg joint points of the target object and the ground.
[0138] In some embodiments, a third weight corresponding to the joint position difference, a fourth weight corresponding to the skeletal length difference, a fifth weight corresponding to the foot velocity error, and a sixth weight corresponding to the altitude difference are obtained. Then, based on the third, fourth, fifth, and sixth weights, the joint position difference, skeletal length difference, foot velocity error, and altitude difference are weighted and summed to obtain the normalized term error of the i-th video frame.
[0139] In the embodiments of this application, weighting each difference or prior information using different weights is performed in order to map the different differences or prior information into a unified numerical range, thereby balancing the limiting effects of the different errors or prior information.
[0140] In steps 431 to 434 above, the positional difference value of the joint points restricts the position of each joint point in adjacent frames to be relatively close, the skeletal length difference value restricts the length of the skeleton in different frames to be the same, the leg velocity restricts the foot movement velocity to be 0 when both feet of two adjacent frames are in contact with the ground, and the height difference between the foot joint point and the ground restricts the height of the foot joint point from the ground to be below a certain threshold when the foot is in contact with the ground in a certain frame. For example, this threshold can be set to 2 centimeters, thereby ensuring the rationality of the human posture through the normalization term error.
[0141] In step 1044, the total error of the target object is determined based on the reprojection error value of the target object, prior information, and normalization term error.
[0142] In some embodiments, the sum of the reprojection error value, prior information, and normalization term error is determined as the total error in the i-th video frame of the target object.
[0143] In step 1045, the target object's awaited posture information is iteratively optimized using gradient descent until the total error of the target object reaches its minimum value, thereby obtaining the corrected posture information of the target object.
[0144] Gradient descent is a first-order optimization algorithm, also known as the steepest descent method. To find the local minimum of a function at a given point, it is necessary to perform an iterative search in the reverse direction of the gradient (or approximate gradient) at the current point on the function for a specified number of steps. When the total error of the target object reaches its minimum, the corrected pose information of the target object is obtained.
[0145] In step 105, the modified pose information is motion retargeted to the virtual object, and a virtual object animation video corresponding to the video to be processed is generated.
[0146] In some embodiments, referring to Figure 5C, step 105 can be achieved by the following steps 1051 to 1052, which are described in detail below.
[0147] In step 1051, the limb mesh information of the virtual object to be driven is obtained.
[0148] In some embodiments, the virtual object being driven may be a virtual object in a game, a virtual object in an animation, or a virtual object in virtual reality. The limb mesh information of the virtual object includes the coordinate information of each mesh point of the virtual object.
[0149] In step 1052, the virtual object is driven to perform an action corresponding to the modified posture information by fitting the modified posture information to the virtual object through motion retargeting, based on the limb mesh information of the virtual object.
[0150] In some embodiments, the coordinate information of each joint point corresponding to the target object is determined based on the limb mesh information of the virtual object to be driven. Next, the motion retargeting system matches each joint point of the target object with each joint point of the virtual object to obtain matching information. Based on this matching information, the virtual object is driven to perform an action corresponding to the modified posture information by fitting the modified posture information of the target object to the virtual object through motion retargeting.
[0151] Steps 1051 to 1052 above allow the captured pose information to be fitted to any virtual object using motion retargeting technology, thereby improving the utilization rate of the modified pose information.
[0152] According to the virtual object animation generation method provided in the embodiment of the present application, after acquiring a video to be processed, first, the initial three-dimensional posture information, two-dimensional posture information, and foot contact information of the target object in the video to be processed are determined. Next, based on the initial three-dimensional posture information of the target object, posture change information between adjacent video frames is determined. Using this posture change information as a motion precedent, it is combined with the initial three-dimensional posture information of each video frame to determine posture information awaiting correction. Subsequently, based on the two-dimensional posture information and foot contact information of the target object, a correction process is performed on the posture information awaiting correction of the target object to obtain the corrected posture information of the target object. Since the posture information awaiting correction of a video frame is determined based on the posture change information as a motion precedent, correcting the posture information awaiting correction is essentially equivalent to correcting the posture change information. This not only improves the efficiency of motion capture by improving correction efficiency, but also ensures the continuity and rationality of the operation of the virtual object animation video obtained by retargeting the corrected posture information to the virtual object through motion retargeting.
[0153] The following describes exemplary applications of the embodiments of this application in actual application scenarios.
[0154] The virtual object animation generation method provided in the embodiment of this application can be applied to application scenes such as driving virtual digital humans and animation production.
[0155] Before describing the virtual object animation generation method provided in the embodiment of this application, we will first explain the SMPL (Skinned Multi-Person Linear) model and the operation of the SMPL model according to the embodiment of this application.
[0156] The SMPL model can define the shape of the human body, such as obesity level, height, and posture during human movement. To define human movement, it is necessary to parameterize each movable joint point of the human body, and changing the parameter of a joint point will change the posture of the human body accordingly.
[0157] The SMPL model is a statistical model, an average human body mesh calculated using a large amount of scanned human body shapes, and the model describes the human body using two types of statistical parameters.
[0158] 1. Shape Parameters: A set of shape parameters contains 10-dimensional numerical values and is used to describe the shape of a person. The value of each dimension can be interpreted as a specific indicator of human body shape (e.g., height or degree of obesity).
[0159] 2. Pose parameters: A set of pose parameters has a 24x3 dimensional numerical value and is used to describe the movement posture of the human body at a given time. In the 24x3, "24" represents the 24 defined joint points of the human body, and "3" refers to the axis-angle representation of the rotation angle of a node relative to its parent node. For these 24 nodes, in the embodiment of this application, a set of joint point trees is defined as shown in Figure 6.
[0160] The SMPL model's operation process can be divided into three steps.
[0161] 1. Shape-based blend shapes.
[0162] At this stage, we have a base template (also called a statistical mean template) as shown in Figure 7A. Using TIFF2026513435000014.tif5170 as the basic pose for the entire human body, this base template is obtained statistically, with N=6890 vertices representing the entire mesh, each vertex having three-dimensional spatial coordinates (x,y,z). Then, using the shape parameter β, the offset amount between the human body pose and this basic pose is described, and by superimposing this offset amount onto the basic pose, the final expected human body pose is formed (see Figure 7B). This process is a linear process, where, TIFF2026513435000015.tif7170 is a matrix multiplication process of a linear matrix with respect to the shape parameter β. The resulting human body mesh pose is called a rest pose (also known as a T-pose) because it does not take into account the influence of the pose parameter.
[0163] 2. Pose Blend Shapes After specifying the shape of the human body mesh based on the designated shape parameter β, a human body mesh with a specific degree of obesity and height can be obtained, as shown in Figure 7B. By superimposing the posture parameter θ onto the human body mesh in a static posture, a human body mesh with a specific degree of obesity, height, and posture can be obtained, as shown in Figure 7C. Certain movements can affect specific shape changes in localized parts of the human body; for example, a person's abdomen may not be visible when standing, but may protrude when sitting. Therefore, the posture parameter θ also affects the mesh shape in a static posture to some extent.
[0164] 3. Skinning Up to this point, calculations have all been performed on a mesh in a static position. However, when the skeletal joints of the human body move, the "skin," which consists of endpoints (vertices), changes in accordance with the movement of the skeletal joints. This process is called skinning. The skinning process can be viewed as a weighted linear combination in which skin nodes are generated in response to changes in skeletal joints. Simply put, the closer an endpoint is to a skeletal joint, the more strongly it is affected by changes such as rotation / translation of that skeletal joint.
[0165] According to the above analysis, to obtain a mesh model of the human body, it is sufficient to obtain the shape parameter β and posture parameter θ of the SMPL-parameterized human body model. Therefore, the entire algorithmic flow of the embodiment of this application can be said to be the problem of estimating these two parameters.
[0166] The virtual object animation generation method provided in the embodiment of the present invention can be realized by steps 801 to 804 shown in Figure 8A, and the overall processing flow of the virtual object animation generation method is as shown in Figure 8B, which will be explained below with reference to Figures 8A and 8B.
[0167] In step 801, the video stream is obtained.
[0168] The video stream may be a recording of motion clips made using a regular camera, or it may be related video containing human motion downloaded directly from a network.
[0169] In step 802, a 3D pose estimation model is used to predict the 3D pose of the human body from the input video.
[0170] This step corresponds to Stage 1 in Figure 8B, and a schematic diagram of the network structure of the 3D pose estimation module in Stage 1 is shown in Figure 9. The implementation process of Step 802 will be described below with reference to Figure 9. First, a video (i.e., a time-series image) is input, then feature extraction is performed on each video frame in the video to obtain video features corresponding to each video frame, then pose prediction is performed on the video features using the GPU to obtain time-series features, and regression and integration are performed on the obtained time-series features to obtain shape parameters and pose parameters of the SMPL parameterized human body model corresponding to each frame image. These are used as initial values for the 3D pose correction module in Stage 2 in Figure 8B.
[0171] In step 803, the 3D pose is corrected using an iterative optimization method.
[0172] In some embodiments, the input video is first used to predict the 2D posture of the human body using a 2D posture estimation model, then the 2D posture is input to a foot contact prediction model to predict the predicted foot contact state, and finally, a 3D posture correction module further optimizes the human body's 3D posture based on the 2D posture and foot contact state to obtain a more accurate 3D posture.
[0173] This step corresponds to step 2 in Figure 8B, which includes a 2D posture estimation module 811, a foot contact prediction module 812, and a 3D posture correction module 813. Figure 10A is a schematic diagram of the network structure of the 2D posture estimation module provided in an embodiment of the present invention. As shown in Figure 10A, the input to the 2D posture estimation module is a 2D image. First, the 2D image is divided into multiple image regions. Next, the encoder 1001 performs encoding on the multiple image regions to obtain the encoded image result. Then, the decoder 1002 performs decoding on the encoded result to obtain 2D keypoints corresponding to the human body in the image. The 2D keypoints obtained in this step are used as supervisor signals for iterative optimization in the 3D posture correction module.
[0174] Figure 10B is a schematic diagram of the network structure of the foot contact prediction module provided in an embodiment of the present invention. As shown in Figure 10B, the foot contact prediction module includes multiple hidden layers (only two hidden layers are shown exemplarily in Figure 10B). These multiple hidden layers perform prediction processing on an input 2D keypoint sequence and output whether the left and right feet are in contact with the ground in each frame image. This output is used as a normalization decision signal during iterative optimization in the 3D posture correction module.
[0175] In this step, the initial 3D posture obtained in step 802 is not corrected. Instead, a motion simulation model is used to predict the change in posture from the current frame to the next frame based on the 3D posture of each video frame obtained in step 802. Then, the sum of the corrected posture of the current frame and this change is calculated to determine the corrected posture for the next frame. Finally, the corrected posture for the next frame is further adjusted using the 3D posture correction module based on the 2D posture of the human body and the foot contact state to obtain a more accurate posture.
[0176] Figure 10D is a schematic diagram of the network structure of the motion a priori model, which is a variational autoencoder trained using a large dataset of 3D human poses. The model predicts the amount of change from the current frame's pose to the next frame's pose based on the human pose of the current frame, and models the rationality of changes between human pose frames through learning with a large amount of data. As shown in Figure 10D, the motion a priori model includes an encoder 1011, a pose change determination module 1012, and a decoder 1013. First, the encoder 1011 in the motion a priori model is used to encode the initial pose parameters of two adjacent video frames of the target object to obtain the encoded result. Next, the pose change determination module 1012 in the motion a priori model is used to perform a prediction process on the encoded result to obtain each prediction result. Finally, the decoder 1013 in the motion a priori model is used to decode the prediction result to obtain pose change information between adjacent video frames.
[0177] One application method is as follows: The video to be optimized is divided into N segments, each segment having 45 frames. Traditionally, a set of posture parameters was optimized for each frame. However, after applying motion analysis, only the first frame of each segment needs to be optimized for a set of posture parameters, and for subsequent frames, only the change from the previous frame to the next frame needs to be optimized.
[0178] The 3D pose correction module is the core of the algorithm in the embodiment of this application. The total target function of iterative optimization is Defined as TIFF2026513435000016.tif6170, where β represents the human body shape parameter, θ represents the human body posture parameter, and K represents the camera parameter. TIFF2026513435000017.tif5170 represents the 2D keypoint coordinates predicted by the model. The role of the 3D pose correction module in the algorithm pipeline is as follows: Using the 3D human pose predicted in stage 1 as the initial value and the 2D keypoints predicted in stage 2 as the supervisor signal, the 3D human pose is iteratively optimized using gradient descent to minimize the overall target function.
[0179] The following is an explanation of each part of the target function.
[0180] 1) The reprojection error can be expressed by equation (2-1).
[0181] TIFF2026513435000018.tif13170 Here, TIFF2026513435000019.tif5170 represents the confidence level of the 2D keypoint predicted by the 2D keypoint model and is used as a weighting factor.
[0182] TIFF2026513435000020.tif5170 represents the coordinates of the i-th 2D keypoint in the image.
[0183] TIFF2026513435000021.tif5170 represents the coordinates in the camera coordinate system of the i-th joint point obtained after inputting the shape parameter β and posture parameter θ into the SMPL parameterized human body model.
[0184] TIFF2026513435000022.tif5170 represents the point coordinates [x,y] obtained by projecting the point coordinates [X,Y,Z] in the camera coordinate system onto the image coordinate system, and the specific calculation formula is: The filename is TIFF2026513435000023.tif9170, where f represents the camera's focal length, which is set to the default value of 1060 in the algorithm.
[0185] TIFF2026513435000024.tif8170 represents the Geman-McClure penalty function, where σ is set to 100. It is primarily used to punish outliers within the target function, making iterative optimization more stable.
[0186] 2) The prior term can be expressed by formula (2-2). TIFF2026513435000025.tif12170
[0187] a. The shape precedence can be expressed by formula (2-3). TIFF2026513435000026.tif12170 Here, β represents the shape parameter in the SMPL parameterized human body model and may be a 10-dimensional vector. The physical meaning of this shape prior term is to constrain the shape parameter of the human body so that it does not deviate significantly from the average human body.
[0188] Postural prioritization can be expressed by equation (2-4). TIFF2026513435000027.tif16170
[0189] Here, TIFF2026513435000028.tif5170 represents the encoding of posture parameters in an SMPL-parameterized human body model in latent space, and may be a 32-dimensional vector. The physical meaning of this posture prior term is to constrain the posture parameters of the human body so as not to deviate significantly from a normal posture.
[0190] The encoding of TIFF2026513435000029.tif5170 employs the variational autoencoder shown in Figure 10C. Since the operational variational autoencoder is trained using a large dataset of 3D human body poses, common human body poses are encoded in the latent space, preventing the appearance of poses that are impossible for the human body to perform.
[0191] 2) The normalization term can be expressed by equation (2-5). TIFF2026513435000030.tif7170
[0192] Here, TIFF2026513435000031.tif5170 shows that the positions of 3D joint points in adjacent frames are relatively close.
[0193] TIFF2026513435000032.tif6170 shows that the bone lengths are consistent across different frames.
[0194] TIFF2026513435000033.tif5170 indicates that the foot velocity error is 0 when both feet are in contact with the ground in two adjacent frames.
[0195] TIFF2026513435000034.tif6170 indicates that when the foot is in contact with the ground in a given frame, the height of the foot joint points from the ground is below a certain threshold, which the algorithm sets to 2 centimeters.
[0196] In step 804, the 3D pose on a standard human body is adapted to a different human character using motion retargeting.
[0197] This step corresponds to stage 3 in Figure 8B and is a motion retargeting process. The function of this step in the algorithm pipeline is to transfer the motion on the SMPL parameterized human body model calculated by the algorithm to another target character.
[0198] Figure 11 is a schematic diagram of the application interface in the virtual engine of the virtual object animation generation method provided in the embodiment of the present invention, where the virtual engine may be Unreal Engine 5 (UE5). As shown in Figure 11, in the application interface, select "Action Panel," select "Video" as the file type, click "Select File," and upload a video in a format that meets the requirements (such as mp4 format). Next, click the "Generate" button and wait for the algorithm to execute to acquire human body movements. Finally, check the acquired human body movement effect on the displayed character.
[0199] According to the virtual object animation generation method provided in the embodiment of this application, the optimized overall scheme ensures the accuracy of motion obtained by video motion capture, and in scenes such as dancing and walking, the movements of a person can be made to closely match the movements in the video. By introducing motion prior art in the 3D posture correction module, the rationality of human body movement in walking scenes can be ensured, and the effectiveness of video motion capture can be greatly improved.
[0200] To ensure understanding, the embodiments of this application involve relevant data such as video frame information, and if the embodiments of this application are applied to a specific product or technology, user permission or consent is required, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.
[0201] The following describes exemplary structures in which the virtual object animation generation device 455 provided in the embodiments of the present application is implemented as a software module. In some embodiments, as shown in Figure 2, the software module in the virtual object animation generation device 455 stored in memory 440 may include a first acquisition module 4551, a first determination module 4552, a second determination module 4553, a first processing module 4554, and a motion retargeting module 4555. The first acquisition module 4551 is configured to acquire a video to be processed and to determine the initial three-dimensional pose information, two-dimensional pose information, and foot contact information of the target object in the video to be processed. The first determination module 4552 is configured to determine pose change information between two adjacent video frames in the video to be processed based on the initial three-dimensional pose information of the target object. The second determination module 4553 is configured to determine the target object's awaiting correction pose information based on the target object's initial three-dimensional pose information and the pose change information. The first processing module 4554 is configured to perform a correction process on the target object's pending correction posture information based on the target object's two-dimensional posture information and foot contact information, thereby obtaining the corrected posture information of the target object. The motion retargeting module 4555 is configured to motion retarget the corrected posture information to a virtual object and generate a virtual object animation video corresponding to the video to be processed.
[0202] In some embodiments, the initial three-dimensional pose information includes initial pose parameters and initial shape parameters, and the first decision module 4551 is further configured to acquire a trained motion priori model, use the trained motion priori model to perform prediction processing on the initial pose parameters of the target object, and obtain pose change information between adjacent video frames.
[0203] In some embodiments, the first decision module 4551 further divides the video to be processed into a plurality of video segments, and if it is determined that the i-th video frame is the first frame in the video segment, it determines the initial three-dimensional pose information of the target object as the target object's awaited correction pose information, where i = 1, 2, ..., N, where N is the total number of video frames in the video to be processed, and the i-th video frame is the current video frame in which the target object is located. If it is determined that the i-th video frame is not the first frame in the video segment, it obtains the modified pose information of the target object in the (i-1)-th video frame, the modified pose information includes modified pose parameters and modified shape parameters, and determines the target object's awaited correction pose parameters based on the pose change information of the target object from the (i-1)-th video frame to the i-th video frame and the modified pose parameters of the target object in the (i-1)-th video frame. The system is configured to determine the initial shape parameters corresponding to the target object and the target object's awaiting modification orientation parameters as the target object's awaiting modification orientation information.
[0204] In some embodiments, the first processing module 4554 is configured to further determine the reprojection error of the target object based on the target object's awaiting correction posture information, two-dimensional posture information, and camera parameters; determine prior information of the target object based on the target object's awaiting correction posture information, the prior information including posture errors and shape errors; determine the normalization term error of the target object based on the target object's foot contact information and corrected posture information in the (i-1)th video frame and the target object's foot contact information and awaiting correction posture information in the ith video frame; determine the total error of the target object based on the target object's reprojection error value, prior information, and normalization term error; and iteratively optimize the target object's awaiting correction posture information using gradient descent until the total error of the target object reaches a minimum value, thereby obtaining the corrected posture information of the target object.
[0205] In some embodiments, the two-dimensional pose information includes reference coordinate information for each joint point of the target object and the confidence level of each joint point. The first processing module 4554 is further configured to determine camera coordinate information for each joint point of the target object based on the target object's awaiting correction pose information and the two-dimensional pose information, determine image coordinate information for each joint point based on the camera parameters and the camera coordinate information for each joint point of the target object, determine the coordinate error for each joint point based on the image coordinate information for each joint point and the reference coordinate information for each joint point, and determine the reprojection error of the target object based on the coordinate error for each joint point and the confidence level of each joint point.
[0206] In some embodiments, the awaiting posture information includes awaiting posture parameters and awaiting shape parameters, and the first processing module 4554 is configured to further determine the shape prior information of the target object based on the awaiting shape parameters of the target object, perform encoding on the awaiting posture parameters of the target object to obtain the encoded result of the awaiting posture parameters in latent space, determine the awaiting posture information of the target object based on the encoded result of the awaiting posture parameters in latent space, and determine the prior information of the target object based on the shape prior information and the awaiting posture information of the target object.
[0207] In some embodiments, the first processing module 4554 is further configured to determine the position difference value of each joint point of the target object and the length difference value of each skeleton of the target object based on the corrected posture information of the target object in the (i-1)th video frame and the pending posture information of the target object in the i-th video frame; to determine the foot velocity error of the target object based on the foot contact information of the target object in the (i-1)th video frame and the foot contact information of the target object in the i-th video frame; to determine the height difference between the foot joints of the target object and the ground; and to determine the normalization term error of the target object based on the position difference value of each joint point of the target object, the length difference value of each skeleton of the target object, the foot velocity error of the target object, and the height difference between the leg joint points of the target object and the ground.
[0208] In some embodiments, the virtual object animation generation device 455 further comprises a second acquisition module, a third decision module, a second processing module, a fourth decision module, and a training module. The second acquisition module is configured to acquire a training video and a predetermined motion prior model, where the training video includes a plurality of training video frames. The third decision module is configured to determine the k-th training video frame and the (k+1)-th training video frame as the k-th training data pair, where k = 1, 2, ..., M-1, and M is the total number of video frames in the training video. The second processing module is configured to perform encoding on the two training video frames in the k-th training data pair to obtain the encoded result of the two training video frames in the k-th training data pair in latent space. The fourth decision module is configured to determine predicted change information for the two training video frames in the k-th training data pair based on the encoded result of the two training video frames in the k-th training data pair in latent space. The training module is configured to train a predetermined motion priori model based on predicted change information of two training video frames within each training data pair, thereby obtaining a trained motion priori model.
[0209] In some embodiments, the first acquisition module 4551 is configured to further acquire a trained three-dimensional pose estimation model, a trained two-dimensional pose estimation model, and a trained foot contact detection model, to perform pose estimation on the video to be processed using the trained three-dimensional pose estimation model to obtain initial three-dimensional pose information of a target object in the video to be processed, to perform pose estimation on the video to be processed using the trained two-dimensional pose estimation model to obtain two-dimensional pose information of the target object, and to perform prediction processing on the two-dimensional pose information of the target object using the trained foot contact detection model to obtain foot contact information of the target object.
[0210] In some embodiments, the motion retargeting module 4555 is further configured to acquire limb mesh information of the virtual object to be driven, and to drive the virtual object to perform an action corresponding to the modified posture information by adapting the modified posture information to the virtual object through motion retargeting based on the limb mesh information of the virtual object.
[0211] The description of the embodiments of the virtual object animation generation apparatus described above is the same as the description of the method described above, and has the same beneficial effects as the embodiments of the method. For technical details not disclosed in the embodiments of the virtual object animation generation apparatus of this application, those skilled in the art should refer to the description of the embodiments of the method of this application for understanding.
[0212] Embodiments of the present application provide a computer program product, which includes a computer program or a computer executable instruction, and which is stored in a computer-readable storage medium. The processor of an electronic device reads the computer executable instruction from the computer-readable storage medium, and the processor executes the computer executable instruction, thereby causing the electronic device to execute the virtual object animation generation method of the embodiment of the present application.
[0213] Embodiments of the present invention provide a computer-readable storage medium in which computer executable instructions are stored, and when the computer executable instructions are executed by a processor, the processor is instructed to execute a virtual object animation generation method provided in embodiments of the present invention, for example, the virtual object animation generation method shown in Figures 3A and 5C.
[0214] In some embodiments, the computer-readable storage medium may be memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM, or it may be a variety of devices that include any one or any combination of the above memories.
[0215] In some embodiments, computer executable instructions can take the form of programs, software, software modules, scripts, or code, and can be written in any form of programming language (including compiled or interpreted languages, declarative or procedural languages), and can be arranged in any form, such as as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.
[0216] For example, computer executable instructions do not necessarily correspond to files in the file system; they may be stored as part of files that store other programs or data. For instance, they may be stored in one or more scripts within a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program being discussed, or in multiple collaborative files (such as files that store one or more modules, subroutines, or code sections).
[0217] For example, executable instructions may be configured to run on a single electronic device, on multiple electronic devices located in one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.
[0218] The above are merely examples of the present application and do not limit the scope of protection. Any modifications, equivalent substitutions, or improvements made in the spirit and scope of the present application shall also be included within the scope of protection.
Claims
1. A method for generating virtual object animations, The process involves acquiring a video to be processed and determining the initial three-dimensional pose information, two-dimensional pose information, and foot contact information of the target object in the video to be processed. The steps include determining the orientation change information between two adjacent video frames in the video to be processed, based on the initial three-dimensional orientation information of the target object, The steps include determining the target object's awaiting correction posture information based on the target object's initial three-dimensional posture information and posture change information, The steps include: performing a correction process on the target object's awaiting correction posture information based on the target object's two-dimensional posture information and foot contact information to obtain the corrected posture information of the target object; A method for generating virtual object animation, comprising the steps of: motion retargeting the modified posture information to a virtual object; and generating a virtual object animation video corresponding to the video to be processed.
2. The initial three-dimensional pose information includes initial pose parameters and initial shape parameters, and the step of determining pose change information between two adjacent video frames in the video to be processed based on the initial three-dimensional pose information of the target object is: Steps to obtain a trained motor a priori model, The process includes the steps of using the pre-trained motion priori model to perform a prediction process on the initial pose parameters of the target object and obtaining pose change information between two adjacent video frames in the video to be processed, The method for generating virtual object animation according to claim 1.
3. The step of determining the target object's awaiting correction pose information based on the initial three-dimensional pose information and the pose change information of the target object in each of the aforementioned video frames is: The steps include dividing the video to be processed into multiple video segments, If it is determined that the i-th video frame is the first frame in the video segment, the initial three-dimensional pose information of the target object is determined as the pose information of the target object awaiting correction, wherein i = 1, 2, ..., N, where N is the total number of video frames in the video to be processed, and the i-th video frame is the current video frame in which the target object is located. If it is determined that the i-th video frame is not the first frame in the video segment, the step is to obtain the modified pose information of the target object in the (i-1)-th video frame, wherein the modified pose information includes modified pose parameters and modified shape parameters. A step of determining the target object's awaiting correction pose parameter based on the pose change information of the target object from the (i-1)th video frame to the i-th video frame and the corrected pose parameter of the target object in the (i-1)th video frame, The process includes the step of determining the initial shape parameters corresponding to the target object and the target object's awaiting modification orientation parameters as the awaiting modification orientation information of the target object. The method for generating virtual object animation according to claim 2.
4. The step of performing a correction process on the target object's awaiting correction posture information based on the target object's two-dimensional posture information and foot contact information, and obtaining the corrected posture information of the target object, is as follows: The steps include determining the reprojection error of the target object based on the target object's awaiting correction orientation information, two-dimensional orientation information, and camera parameters, A step of determining prior information of the target object based on the target object's awaiting correction posture information, The steps include determining the normalization term error of the target object based on the foot contact information and corrected posture information of the target object in the (i-1)th video frame, and the foot contact information and awaiting correction posture information of the target object in the i-th video frame, A step of determining the total error of the target object based on the reprojection error value of the target object, prior information, and normalization term error, The process includes the step of iteratively optimizing the target object's awaiting correction orientation information using gradient descent until the total error of the target object reaches its minimum, thereby obtaining the corrected orientation information of the target object. The method for generating virtual object animation according to claim 1.
5. The two-dimensional pose information includes reference coordinate information for each joint point of the target object and the confidence level of each joint point. The step of determining the reprojection error of the target object based on the target object's pending pose information, two-dimensional pose information, and camera parameters is as follows: The steps include determining the camera coordinate information of each joint point of the target object based on the target object's awaiting correction pose information and two-dimensional pose information, The steps include determining the image coordinate information of each joint point based on the camera parameters and the camera coordinate information of each joint point of the target object, The steps include determining the coordinate error of each joint point based on the image coordinate information of each joint point and the reference coordinate information of each joint point, The process includes the step of determining the reprojection error of the target object based on the coordinate error of each joint point and the reliability of each joint point. The method for generating virtual object animation according to claim 4.
6. The awaiting correction posture information includes awaiting correction posture parameters and awaiting correction shape parameters, the prior information includes posture prior information and shape prior information, and the step of determining the prior information of the target object based on the awaiting correction posture information of the target object is: The steps include determining the shape prior information of the target object based on the shape parameters of the target object awaiting modification, The steps include: performing an encoding process on the target object's pending posture parameters to obtain the encoded result of the pending posture parameters in latent space; A step of determining prior posture information of the target object based on the encoding result of the posture parameters awaiting correction in latent space, The process includes the step of determining prior information of the target object based on prior information of the shape of the target object and prior information of the posture of the target object. The method for generating virtual object animation according to claim 4.
7. The step of determining the normalization term error of the target object based on the foot contact information and corrected posture information of the target object in the (i-1)th video frame and the foot contact information and awaiting correction posture information of the target object in the i-th video frame is: A step of determining the positional difference value of each joint point of the target object and the length difference value of each skeleton of the target object based on the modified pose information of the target object in the (i-1)th video frame and the pending pose information of the target object in the i-th video frame, A step of determining the foot velocity error of the target object based on the foot contact information of the target object in the (i-1)th video frame and the foot contact information of the target object in the i-th video frame, The steps include determining the height difference between the foot joint of the target object and the ground, The process includes determining the normalized term error of the target object based on the positional difference values of each joint point of the target object, the length difference values of each skeleton of the target object, the foot velocity error of the target object, and the height difference between the leg joint points of the target object and the ground. The method for generating virtual object animation according to claim 4.
8. The virtual object animation generation method described above is: A step of obtaining a training video and a predetermined motor priori model, wherein the training video includes a plurality of training video frames, A step in which the k-th training video frame and the (k+1)-th training video frame are determined as the k-th training data pair, where k = 1, 2, ..., M-1, and M is the total number of video frames in the training video. The steps include: performing an encoding process on two training video frames in the k-th training data pair to obtain the encoding result of the two training video frames in the k-th training data pair in the latent space; A step of determining predicted change information for the two training video frames in the k-th training data pair based on the encoding results of the two training video frames in the k-th training data pair in the latent space, The further step includes adjusting the parameters of the predetermined motor a priori model based on posture change information of two video frames within each training data pair to obtain a trained motor a priori model. A method for generating virtual object animation according to any one of claims 2 to 7.
9. The step of determining the initial three-dimensional pose information, two-dimensional pose information, and foot contact information of the target object in the video to be processed is: The steps include obtaining a trained three-dimensional pose estimation model, a trained two-dimensional pose estimation model, and a trained foot contact detection model, The steps include: performing pose estimation on the video to be processed using the trained three-dimensional pose estimation model to obtain initial three-dimensional pose information of the target object in the video to be processed; The steps include: performing pose estimation on the video to be processed using the trained two-dimensional pose estimation model to obtain two-dimensional pose information of the target object; The process includes the step of using the trained foot contact detection model to perform a prediction process on the two-dimensional posture information of the target object and obtaining the foot contact information of the target object. A method for generating virtual object animation according to any one of claims 1 to 7.
10. The steps of motion retargeting the modified posture information to a virtual object and generating a virtual object animation video corresponding to the video to be processed are: The steps include obtaining limb mesh information of the virtual object to be driven, The process includes the steps of: driving the virtual object to perform an action corresponding to the modified posture information by fitting the modified posture information to the virtual object using motion retargeting based on the limb mesh information of the virtual object, and generating a virtual object animation video corresponding to the video to be processed. A method for generating virtual object animation according to any one of claims 1 to 7.
11. A virtual object animation generation device, A first acquisition module is configured to acquire a video to be processed and to determine the initial three-dimensional pose information, two-dimensional pose information, and foot contact information of the target object in the video to be processed. A first determination module is configured to determine pose change information between two adjacent video frames in the video to be processed, based on the initial three-dimensional pose information of the target object. A second determination module is configured to determine the target object's awaiting correction posture information based on the target object's initial three-dimensional posture information and posture change information. A first processing module is configured to perform a correction process on the target object's awaiting correction posture information based on the target object's two-dimensional posture information and foot contact information, and to obtain the target object's corrected posture information. A virtual object animation generation device comprising: a motion retargeting module configured to motion retarget the modified posture information to a virtual object and generate a virtual object animation video corresponding to the video to be processed.
12. It is an electronic device, Memory configured to store computer executable instructions, An electronic device comprising: a processor configured to realize the virtual object animation generation method described in any one of claims 1 to 10 by executing computer-executable instructions stored in the memory.
13. A computer-readable storage medium storing computer-executable instructions for a processor to perform the virtual object animation generation method described in any one of claims 1 to 10.
14. A computer program product comprising a computer program or computer executable instruction for causing a processor to perform the virtual object animation generation method described in any one of claims 1 to 10.
Citation Information
Patent Citations
Three-dimensional animation attitude prediction method and system
CN111311714A
Data processing method and device, electronic device and computer readable storage medium
CN113706699A
Video generation method and device, electronic equipment and readable storage medium
CN113961746A
Human-shaped agent posture generation method based on physical simulation
CN115018963A
Apparatus and method for estimating posture
JP2007310707A